Abstract
Background:
Assessing walking impairment in those with multiple sclerosis (MS) is common, however little is known about the reliability, precision and clinically important change of walking outcomes.
Objective:
The purpose of this study was to determine the reliability, precision and clinically important change of the Timed 25-Foot Walk (T25FW), Six-Minute Walk (6MW), Multiple Sclerosis Walking Scale-12 (MSWS-12) and accelerometry.
Methods:
Data were collected from 82 persons with MS at two time points, six months apart. Analyses were undertaken for the whole sample and stratified based on disability level and usage of walking aids. Intraclass correlation coefficient (ICC) analyses established reliability: standard error of measurement (SEM) and coefficient of variation (CV) determined precision; and minimal detectable change (MDC) defined clinically important change.
Results:
All outcome measures were reliable with precision and MDC varying between measures in the whole sample: T25FW: ICC=0.991; SEM=1 s; CV=6.2%; MDC=2.7 s (36%), 6MW: ICC=0.959; SEM=32 m; CV=6.2%; MDC=88 m (20%), MSWS-12: ICC=0.927; SEM=8; CV=27%; MDC=22 (53%), accelerometry counts/day: ICC=0.883; SEM=28450; CV=17%; MDC=78860 (52%), accelerometry steps/day: ICC=0.907; SEM=726; CV=16%; MDC=2011 (45%). Variation in these estimates was seen based on disability level and walking aid.
Conclusion:
The reliability of these outcomes is good and falls within acceptable ranges. Precision and clinically important change estimates provide guidelines for interpreting these outcomes in clinical and research settings.
Introduction
Multiple sclerosis (MS) is a prevalent immune-mediated, progressive and disabling neurological disease. Walking impairment is an almost ubiquitous feature of MS, 1 affecting approximately 80% of cases 2 and represents a valued function across the disability spectrum. 3 Thus, walking is a common metric for monitoring disease progression and a central outcome of pharmacotherapy and rehabilitation interventions.4,5
Multiple outcomes exist for measuring walking in MS, including the Timed 25-Foot Walk (T25FW), 6 Six-Minute Walk (6MW), 7 Multiple Sclerosis Walking Scale-12 (MSWS-12) 8 and free-living accelerometry.9,10 It is acknowledged within the literature that different walking outcome assessments used in MS research capture different aspects of walking, 11 and to better inform users of these scales there is a strong need to acknowledge validity and reliability to enhance decision making when choosing outcome assessments. There is evidence to support the validity of inferences from scores on all four outcomes as measures of walking impairment.12–14 The T25FW and 6MW have differentiated between MS and controls12,13 and between persons with MS who vary in disability based on the Expanded Disability Status Scale (EDSS).12,14,15 MSWS-12 scores have correlated with T25FW, 6MW and gait kinematics in MS 16 and differed between persons with MS who vary in disability. 8 Accelerometer output of counts/day and steps/day have differed between MS and controls and disability level within those with MS10,17 and correlated with T25FW, 6MW and MSWS-12 in MS. 18
Less is known about the reliability of T25FW, 6MW, MSWS-12 and accelerometry metrics over time.19,20 Reliability provides an indication of a measure’s consistency and precision over time in the absence of a change based on the intraclass correlation (ICC) coefficient, standard error of measurement (SEM) and coefficient of variation (CV). The clinical importance of change can be captured using the minimal detectable change (MDC). 21 T25FW scores have been reliable over a one-week time period in those with MS who had EDSS scores between 5–6.5 based on an ICC of 0.94 with a SEM of 4.56 s and an MDC of 12.6 s. 19 6MW distance has been reliable over one week based on an ICC of 0.96, SEM of 30 m and MDC of 76 m. Nilsagård et al. 22 has demonstrated that reliability of outcomes assessing walking may be affected by level of disability in MS, when assessed over one week.
There are limitations of previous research and reasons for further examining the reliability, precision and clinically detectable change of walking measures in MS. One limitation is the lack of research reporting ICC, SEM, CV and MDC values for MSWS-12 scores or accelerometry metrics. Another limitation is that reliability has only been established over a one-week period and primarily in homogeneous groups. These psychometric characteristics should be examined over a longer period and between MS patient characteristics such as disability status or assistive device use. Indeed, clinicians monitor disease progression based on walking outcomes during clinical visits that typically occur bi-annually. Researchers who conduct clinical trials often include walking outcomes to monitor patients over longer intervals.23–27 We further note that ICC estimates are important components of power analyses for planning longitudinal and clinical trials. The default ICC value in some power analysis software is 0.50 28 but this value may be incorrect, thereby threatening the veracity of the power analysis and sample size estimates.
This study examined the reliability, precision and clinically detectable change of the T25FW, 6MW, MSWS-12 and accelerometry over a six-month period of time. We further examined those psychometric parameters between benchmarks of disability and use of a walking aid.
Method
Ethical approval was granted by a University Institutional Review Board, with all participants providing written informed consent. The protocol comprised two testing sessions, separated by six months with no intervention provided during this time.
Recruitment and participants
The data were secondary outcomes from a non-intervention, six-month observational period of an ongoing intervention targeted toward changing physical activity in MS. Participants were contacted either by a flyer that was distributed amongst patients in the North American Research Committee on Multiple Sclerosis (NARCOMS) registry or by e-mail through a flyer that was distributed amongst participants in a database from previous studies conducted. There were 511 participants who initially expressed interest and who were contacted by the project-coordinator. After explaining the study protocol, the project-coordinator undertook screening for inclusion with 230 individuals who remained interested. The inclusion criteria involved: (a) diagnosis of MS; (b) relapse-free for the past 30 days; (c) ability to walk with or without an assistive device; (d) age between 18–64 years; (e) willingness and ability to travel to the research site, complete the walking assessments and wear an accelerometer for one week; and (f) physician’s approval for participation in the study. For this study 106 individuals did not meet one or more inclusion criteria with their primary reasons being too physically active and unwilling to travel, 39 people did not provide physician’s approval and three people cancelled the testing session due to scheduling conflicts. The final sample included 82 persons with MS.
Procedures
Disability was ascertained using the Self-Report Expanded Disability Status Scale (SR-EDSS). 29 This SR-EDSS has demonstrated validity based on good overall agreement (ICC=0.90) and correlation (r=0.90) with the clinician-administered EDSS and a non-significant difference in overall mean scores and distributions between EDSS versions. 29 Four walking measures were administered during this study, the T25FW, 6MW, MSWS-12 and accelerometry and each is described in Table 1. The protocols have previously been described.7,8,12,30,31 The same procedure was repeated six months later.
Description of each mobility outcome measure assessed.
6MW: Six-Minute Walk; MS: multiple sclerosis; MSWS-12: Multiple Sclerosis Walking Scale-12; T25FW: Timed 25-Foot Walk.
Statistical analysis
Data were analysed using PASW v18 (SPSS Inc., Chicago, Illinois, USA). ICC analyses (2,1 mixed model) were performed to assess test-retest reliability over time based on average group performance. 32 Estimates close to 1.0 indicate strong reliability, scores of 0.8 suggest good reliability and ICC estimates of 0.6 suggest moderate, but still acceptable, reliability. 28 The SEM was calculated by multiplying the baseline standard deviation (SD) of each outcome measure by the square root of one minus the reliability coefficient; SEM = SDbaseline × √(1–ICC).28,32,33 This value indicates the amount of variability inherent in the measurement attributable to measurement error.20,33 CV was calculated by dividing the SD of the difference between the two time-points for the whole sample, by the mean difference between the two time points multiplied by 100.20,33 The presented CV is a mean of this calculation which provides an estimate of consistency over time and reflects the percentage of change within the measure due to measurement error. 33 MDC was calculated by multiplying 1.96 (derived from the 95% confidence interval (CI)) by the square root of two (as two measurements were taken) times the SEM; MDC=1.96×√(2)×SEM. 34 The MDC reflects the amount of change not due to measurement error and therefore, may provide an indication of real change.21,35 We further expressed the MDC as a percentage of the overall mean from both time-points. Pairwise analyses were undertaken for all reliability, precision and MDC analyses as missing data varied dependant on the outcome.
Results
Participant characteristics
Participant demographic characteristics are provided in Table 2. Participants appeared to be stable over the six-month study, based on no significant change in disability status over the two time points (SR-EDSS, p=0.871, PDDS (Patient Determined Disease Steps), p=0.118). Furthermore, no participant reported symptom relapse or clinically diagnosed relapse during the study.
Demographic data and multiple sclerosis (MS)-related characteristics of participants at baseline.
IQR: interquartile range; SD: standard deviation; SR-EDSS: Self-Report-Expanded Disability Status Scale. PDDS = Patient Determined Disease Steps scale.
Overall sample
Table 3 contains estimates of reliability (ICC), precision (SEM and CV) and clinically important change (MDC) for the whole sample. The ICC values for the test-retest reliability over six months ranged between 0.883–0.991. This indicates strong (ICC>0.8) reliability. The SEM and CV estimates provided an indicator of precision (i.e. error of the measurement) and should be considered alongside the mean scores. The T25FW and the 6MW had the best estimates of precision (CV and SEM), whereas the MSWS-12 and accelerometer metrics had less, but still acceptable, precision. The SEM for the 6MW was 32 m (where the mean score was 439 m), indicating that a change of 32 m or less may be due to measurement noise. The CV for the 6MW indicated that a change in score of 6.2% or less over six months may be expected with the 6MW and thus interpreted as no change. By comparison, the SEM for the MSWS-12 was eight points (where the mean score was 41), indicating that a change of eight points or less may be due to measurement noise. The CV for the MSWS-12 indicated that a change of 27% or less over six months may be interpreted as normal.
Showing mean score, ICC score, SEM, CV and MDC results for mobility outcomes in 82 persons with multiple sclerosis (MS).
CI: confidence interval; CV: coefficient of variation; ICC: intraclass correlation coefficient (2,1); MDC95: minimal detectable change (at 95% CI); SD: standard deviation; SE: standard error of the mean; SEM: standard error of the measurement. a4 missing cases, b7 missing cases.
The MDC provides an indication of a clinically important change and is based on the SEM: it too should be considered alongside the overall mean. The 6MW and T25FW had the smallest estimates of clinical importance, whereas the MSWS-12 and accelerometer metrics had the largest estimates. The MDC for the 6MW was 88 m and a 20% change would indicate a clinical change. The MDC for MSWS-12 scores was 22 points and a 53% change would indicate a clinically important change.
Subsamples of disability and aid use
Tables 3 and 4 provided estimates of reliability, precision and clinically important change for subsamples based on disability categories (i.e. mild disability=SR-EDSS<4; moderate disability=SR-EDSS≥4) and mobility use (i.e. no walking aid; unilateral or bilateral support), respectively. When the results were dichotomised based on SR-EDSS status, distribution of scores for the MSWS-12 were similar in both groups (i.e. SD and SE results) but for the timed walking outcomes (T25FW, 6MW) and accelerometer metrics, scores were distributed wider in the more disabled group. ICC estimates indicated that all walking measures were stable over time for the group with mild disability (ICC range=0.875–0.981). The ICC estimates were strong for T25FW, 6MW and accelerometer steps/day for the group with moderate disability, but weaker, although still acceptable, for accelerometer counts/day and MSWS-12 score (i.e. ICC<0.8). Precision estimates for the T25FW and 6MW were better for those who were less disabled. By comparison, precision estimates for accelerometer metrics were comparable between disability groups and precision for the MSWS-12 was better in the group with moderate disability (Table 4). MDC was higher in those who had moderate disability for the 6MW, T25FW and accelerometer metrics, but lower for the MSWS-12, compared with those who had mild disability.
Showing mean score, ICC score, SEM, CV and MDC results for mobility outcomes in 82 persons with multiple sclerosis (MS) based on Self-Report Expanded Disability Status (SR-EDSS) score.
CI: confidence interval; CV: coefficient of variation; ICC: intraclass correlation coefficient (2,1); MDC95: minimal detectable change (at 95% CI); SD: standard deviation; SE: standard error of the mean; SEM: standard error of the measurement.
4 missing cases, b7 missing cases.
Reliability was strong (i.e. ICC≥0.8) for all measures in those who did not use walking aids (ICC range=0.878–0.989) (Table 5). This was not the case for accelerometer metrics among those who used a walking aid, whereby the ICC estimates were less than 0.80 for those who used walking aids (results were still within acceptable ranges of reliability). 32 The ICCs for the other measures were strong in the subgroup that walked with assistance (ICC≥0.8). Precision and MDC estimates mimicked the pattern seen in the SR-EDSS group comparisons.
Showing mean score, ICC score, SEM, CV and MDC results for mobility outcomes in 82 persons with multiple sclerosis (MS) based on walking assistive device.
CI: confidence interval; CV: coefficient of variation; ICC: intraclass correlation coefficient (2,1); MDC95: minimal detectable change (at 95% CI); SD: standard deviation; SE: standard error of the mean; SEM: standard error of the measurement.
4 missing cases, b7 missing cases.
Discussion
Walking impairment is an important consequence of MS that impacts function and quality of life 36 and has become a primary end-point in clinical research and practice. There is evidence for the validity of the T25FW, 6MW, MSWS-12 and accelerometry scores as measures of walking impairment in MS, 37 but less evidence exists for other psychometric properties. This study assessed the test-retest reliability, precision and minimal detectable change for the T25FW, 6MW, MSWS-12 and accelerometry over a six-month period based on group averages from 82 seemingly stable persons with MS.
Reliability
Overall, the outcome measures were highly reliable across six months in the entire sample, with ICC estimates exceeding 0.80. This complements and extends previous studies19,20 which indicate that the T25FW and 6MW were reliable over a one-week period in homogeneous samples of 19 people who had an EDSS score of 6.5 or less 20 and 24 people who had an EDSS score of 5 to 6.5. 19 We extend that evidence by demonstrating strong reliability of T25FW, 6MW, MSWS-12 and accelerometer metrics (i.e. multiple outcome measures) in a larger, heterogeneous sample across a longer time period that is consistent with observations performed in clinical trials and practice. This indicates that these outcomes are generally reliable for monitoring walking status over time among persons with MS.
The reliability of measures varied between groups who differed in disability and ambulatory device usage, a finding similar to the results of Nilsagard et al. 22 The 6MW and T25FW were highly reliable across six months in both subsample groupings (i.e. disability and device usage) with all ICC estimates exceeding 0.80. By comparison, the accelerometer metrics and MSWS-12 scores had reliability estimates less than 0.80 in those with moderate disability and the accelerometer metrics further had reliability estimates less than 0.80 in those who used an aid for ambulation. Importantly, the reliability estimates, even when below 0.80, still satisfied criteria for acceptable reliability in the subsamples. 32 Previous research has acknowledged that different walking scales perform differently across the MS disability range. 22 For example the MSWS-12 is thought to provide an adequate assessment of walking performance across the disability range, whilst the T25FW may be less sensitive to change in those with milder disability, 11 The distribution of scores in this study imply that this may be the case. Further differences attributable to the difference in reliability between the scales may be that: the T25FW and 6MW were carried out under laboratory conditions, following strict protocols that regulate behavioral variation; accelerometer data were collected during walking undertaken in one’s daily life under free-living conditions (i.e. daily walking in one’s own environment); the MSWS-12 similarly reflects perceived real-world ambulatory impairment. This may impact the reliability of behavior captured by the accelerometer metrics. Nevertheless, collecting data in the natural environment is of high importance as it provides an indication of walking under real world conditions (i.e. ecological validity). Accelerometry should be included in MS research and practice and, considering the results of this study, it should be complemented with other measures of walking.
Precision and clinical important change
The precision varied amongst the outcome measures in the overall sample. The T25FW and 6MW were the most precise based on SEM and CV estimates followed by accelerometer metrics. The MSWS-12 was less precise based on SEM and CV estimates. The clinically important change estimates varied amongst the outcomes such that the 6MW and T25FW required the least change: larger changes were required in MSWS-12 and accelerometer metrics to indicate a clinically important change. This difference again makes sense, as discussed with ICC estimates, considering the conditions and context under which data were collected for the T25FW and 6MW compared with accelerometry and the MSWS-12.
Interesting results emerged when we considered precision and clinical important change for the outcomes between groups who differed in disability and ambulatory device usage. The precision and MDC for the 6MW, T25FW and accelerometer metrics were better among those who were less disabled and those who did not use walking aids. The exact opposite pattern was observed for the MSWS-12 where the precision and MDC were better among persons with moderate disability and who used walking aids. This is probably associated with the greater variability compared with mean MSWS-12 scores for those with less disability, for example, than those with moderate disability: the exact opposite is observed for the T25FW, 6MW and accelerometer metrics.
Implications for research and practice
The precision and clinically important change scores established in this study may guide clinicians and researchers in interpreting if change in participant performance is meaningful (MDC scores) or associated with noise in the measurement (SEM and CV scores). Importantly, previous research has indicated that a change between 20%6,38 (based on the variability of mean scores over consecutive walks) and 70% 19 (based on MDC calculations) in T25FW performance represents a clinically significant change. The present study established a 36% change in T25FW performance to be clinically meaningful in the overall sample, but this value was substantially lower in those with mild disability and those who walked without an aid (12% and 15% respectively) compared to the respective comparative subgroup. This highlights variability in estimates for interpreting the meaningfulness of a change in performance on walking outcomes and this might be associated with the characteristics of the sample and the conditions and design of data collection. Researchers and clinicians should be aware that the characteristics of the sample are important when selecting criteria for interpreting meaningful change in walking outcomes.
There were differences amongst the outcome measures in reliability, precision and clinically important change by disability level and use of walking aids. Interestingly, the T25FW, 6MW and accelerometry were less reliable in those who were more disabled (SR-EDSS≥4) and furthermore had less precision and a larger clinically important change score in the subgroups. The MSWS-12 appeared to be less reliable in the more disabled group, but had better precision and a smaller clinically important change score in the more impaired subgroups. There are many possible reasons for the differential results with the MSWS-12, but one implication is that this scale might not be an ideal choice for measuring walking impairment in those with mild disability or those who walk without a device. Yet, our results indicate that other walking outcome measures are more precise and capable of detecting clinical change in those who have less disability or mobility problems.
The study might have implications for sample size estimates in future clinical trials. To that end, we provide sample size estimates using reliability parameters identified herein for detecting a small interaction in a typical randomised controlled trial using G*Power Version 3.1.2. We assumed a mixed model analysis of variance (ANOVA) with group as a between-subjects factor (i.e. intervention vs control) and time as a within-subject factor (i.e. pre-post assessments). The parameters for α and β were 0.05 and 0.80, respectively, the effect size was small (f-value=0.10) and non-sphericity (ϵ) was 1.0. Using the lowest reliability estimate of 0.88 for the accelerometer output in the entire sample, the power analysis indicated a total sample size of 50 persons is necessary for detecting a small interaction effect. Using the lowest reliability estimate of 0.67 for the accelerometer output in those who walk with an aid, the power analysis indicated that a total sample size of 132 persons is necessary for detecting a small interaction effect. Using the default reliability estimate of 0.50 in G*Power, the power analysis indicated that a total sample size of 200 persons is necessary for detecting a small interaction effect. This illustrates the importance of empirically defining the reliability parameter for power analyses.
Limitations
This paper provides worthwhile information to guide clinicians and researchers. The strengths include: the inclusion of multiple outcome measures, clinical and real-world measures; analyses across a wide range of disability levels in MS and between groups such as those deemed to be mildly and moderately disabled; multiple statistical methods; and a six-month timeframe, similar to clinical trials and practice. Furthermore statistical methods to report reliability were chosen based on previous research, and can be easily replicated. However it is acknowledged that other statistical methods are available (such as Rasch analysis) which may complement the findings. In addition, other limitations within this work include the following: only walking outcomes were assessed in this study; comparison with other studies is limited by the use of the SR-EDSS scale, such that disability level was self-reported, rather than administered by a clinician. Lastly, we report data based on an apparently stable sample (i.e. no significant change in SR-EDSS or PDDS score with no reported symptom or clinically diagnosed relapse, and acknowledge that in the absence of objective clinical data it cannot be confirmed if individual participants were experiencing episodes of disease.
Summary
The reliability of the T25FW, 6MW, MSWS-12 and accelerometry were acceptable in the entire sample and subsamples. There is variation in the reliability, precision and clinically important changes amongst walking measures based on level of disability and use of walking aid. Researchers and clinicians should be aware of the subtle, yet important, differences in reliability, precision and clinically important changes when administering and interpreting walking measures in MS.
Footnotes
Acknowledgements
RM acquired funding. DD, LP, BS and RM initiated the overall study and oversaw all data collection. YL and RM undertook statistical analysis. YL wrote the draft manuscript, DD, LP, BS and RM revised the draft manuscript. All authors gave approval of the final submitted version. The authors wish to thank all participants in this research and all staff and students involved in data collection.
Declaration of conflicting interests
None declared.
Funding
This research was funded by a grant from the National Multiple Sclerosis Society (PP1695). YL was the recipient of a Du Prẻ Award from the Multiple Sclerosis International Federation.
