Abstract
Background:
The Expanded Disability Status Scale (EDSS) has low sensitivity and reliability for detecting sustained disability progression (SDP) in multiple sclerosis (MS) trials.
Objective:
This study evaluated composite disability end points as alternatives to EDSS alone.
Methods:
SDP rates were determined using 96-week data from the Olympus trial (rituximab in patients with primary progressive MS). SDP was analyzed using composite disability end points: SDP in EDSS, timed 25-foot walk test (T25FWT), or 9-hole peg test (9HPT) (composite A); SDP in T25FWT or 9HPT (composite B); SDP in EDSS and (T25FWT or 9HPT) (composite C); and SDP in any two (EDSS, T25FWT, and 9HPT) (composite D).
Results:
Overall agreements between EDSS and other disability measures in defining SDP were 66%−73%. Composite A showed similar treatment effect estimate versus EDSS alone with much higher SDP rates. Composite B, C, and D all showed larger treatment effect estimate with different or similar SDP rates versus EDSS alone. Using composite A (24-week confirmation only), B, C, or D could reduce sample sizes needed for MS trials.
Conclusion:
Composite end points including multiple accepted disability measures could be superior to EDSS alone in analyzing disability progression and should be considered in future MS trials.
Introduction
The long-term clinical outcome of multiple sclerosis (MS) is largely determined by disability progression and, to a lesser extent, by the number of clinical relapses. 1 With multiple MS drugs now approved for reducing relapse rate, the focus of clinical trials is shifting to the prominent unmet medical need of limiting disability progression.
The Expanded Disability Status Scale (EDSS) has been the only accepted outcome measure of disability progression for registration MS trials in the past, despite its poor inter-test reproducibility (reliability) and insensitivity to longitudinal change (responsiveness). 2 Approaches have been proposed to reduce its overall variability, such as confirming an initial progression event at a later time point (confirmed disability progression or sustained disability progression (SDP)), using a more reliable baseline for defining progression, 3 keeping the same EDSS raters (if possible) throughout a trial, and using a tablet system with automatic edit checks. However, given the complexity of the EDSS, these approaches show limited effect at reducing variability, with no impact on the poor responsiveness of the EDSS to short-term (2–3 years) changes in disability progression. Thus, large sample sizes remain necessary to obtain sufficient statistical power to demonstrate a drug’s clinical efficacy in slowing disability progression, especially in superiority trials with an active control arm.
The timed 25-foot walk test (T25FWT) and the 9-hole peg test (9HPT) are alternative disability measures that have been accepted in MS clinical trials. Although they overlap with some EDSS domains (i.e. pyramidal, cerebellar, cerebral scores) typically affected by MS, they also represent more focused and quantifiable neurologic functional systems. The T25FWT measures peak performance of walking in a distance-locked approach, whereas its EDSS counterpart quantifies walking distance as a measure of gait endurance and need for support. The 9HPT measures a combination of cognitive, motor and coordinative functions of upper extremities, which have no direct counterpart in the EDSS test battery. Recent research showed that a 20% worsening from baseline in T25FWT and 9HPT indicate a consistent clinical impact on disability,4–6 enabling disability progression events to be easily identified by these two measures. We hypothesized that individual SDP results from these three tests overlap only partially; thus, their combination may represent a more sensitive and reliable disability measure.
In this study, composite end points for measuring sustained disability were assessed by reanalysis of data from a phase III clinical trial of rituximab in primary progressive (PP) MS (Olympus). 7 We first evaluated the overlap of SDP events in the placebo arm between the EDSS, T25FWT, and 9HPT at both 12- and 24-week confirmation points. We then evaluated how composite end points combining SDP results from the single tests could alter treatment effect estimates and SDP event rates. Furthermore, using these composite end points, we assessed the various sample sizes needed to detect significant treatment effect in future trials of patients with similar projected disability progression rates. Based on these results, along with limited information from other progressive and relapsing MS cohorts, we discussed the potential implementation of the composite end points in future MS trials.
Methods
Olympus is a randomized, double-blind, placebo-controlled, multicenter trial in PPMS (n = 439). The protocol was approved by the institutional review board and the ethics committee of each institution. Written informed consent was obtained from each patient or the patient’s legal guardian. Patients with EDSS scores of 2.0–6.5 inclusive at screening were randomly assigned (2:1) to dual intravenous infusions of rituximab 1000 mg or placebo every 24 weeks for four treatment cycles over 96 weeks. Details of the specific patient inclusion and exclusion criteria, study design and primary results were reported previously. 7 The EDSS evaluation, T25FWT and 9HPT (dominant and non-dominant hand) were administered at baseline and every 12 weeks during the study. T25FWT was done twice at each visit and the average outcome of the two T25FWT trials was used in the analysis. 9HPT was done twice for each hand at each visit. The average outcome of the four 9HPT trials was used in the analysis.
For EDSS, initial disability progression was defined as an increase of ≥1 point from baseline if the baseline score was ≤5.5 or an increase of ≥0.5 points if the baseline score was >5.5. For T25FWT and 9HPT, initial disability progression was defined as 20% worsening from baseline. SDP was assessed by the same test up to 12 and 24 weeks (12- and 24-week confirmation point) after initial disability progression was noted. All available outcomes from the same test between initial disability progression and the confirmation point needed to meet the disability progression criteria to confirm a SDP event. Time to SDP was defined as time from randomization to the initial time of SDP (i.e. not the confirmation time point). For patients who did not have SDP by their last scheduled visit, time to SDP was censored at the last assessment.
Placebo data were used to assess agreement among EDSS, T25WT, and 9HPT SDP outcomes. The overlap of SDP events, as captured by the three single tests, was quantified and presented by Venn diagrams for 12- and 24-week confirmation points.
Composite end points were generated from prime sets (EDSS, T25FWT, 9HPT SDP results); by combination of unions (set 1 ∪ set 2, ‘OR’); and intersections (set 1 ∩ set 2, ‘AND’, i.e. overlaps between prime sets). If multiple SDP events were shown from the components of the composite end point, SDP event time for the composite end point was defined as the time of the first SDP event. SDP rates and treatment effects in Olympus, as well as sample sizes needed to achieve adequate power for the future trial, were examined by the following composite end points: composite A, EDSS ∪ T25FWT ∪ 9HPT (SDP from any one of the three tests); composite B, T25FWT ∪ 9HPT (SDP from T25FWT or 9HPT); composite C (EDSS ∩ T25FWT) ∪ (EDSS ∩ 9HPT) (EDSS SDP confirmed by T25FWT or 9HPT SDP); and composite D (EDSS ∩ T25FWT) ∪ (EDSS ∩ 9HPT) ∪ (T25FWT ∩ 9HPT) (SDP confirmed by any two of the three tests).
Statistical analysis
The SDP rates for each treatment group were calculated and plotted using Kaplan–Meier (KM) lifetime table methods. The percentage of agreed outcomes (total number of patients with agreed outcomes (SDP or non-SDP) divided by total number of patients) and Cohen’s kappa coefficient were used to assess the agreement of EDSS and T25WT/9HPT in defining SDP. Log-rank test and Cox proportional hazards model were used to estimate treatment effects in a single test or in composite end points. A two-group test of equal exponential survival with exponential dropout was used to determine the sample size/power for detecting significant treatment effects in delaying time to SDP.
Results
SDP rates and agreement between EDSS, T25FWT, and 9HPT
In the Olympus placebo arm, the overall agreement in defining SDP (12- or 24-week confirmation) between EDSS and T25WT and between EDSS and 9HPT was 66%−73% (with kappa coefficients ranging from 0.18–0.36), indicating at most only fair agreements (Table 1). Individual SDP rates were highest with T25FWT, followed by EDSS and 9HPT at both confirmation points. For all three measurements, the vast majority of 12-week SDP events (88%−91%) were confirmed at 24 weeks, when only those events with initial progression events occurring 24 weeks before end of study were compared (Table 1).
Agreement and 96-week SDP rates in Olympus placebo group (n = 147).
9HPT: nine-hole peg test; CI: confidence interval; EDSS: Expanded Disability Status Scale; SDP: sustained disability progression; T25FWT: timed 25-foot walk test.
12-week confirmation events of SDP occurring after week 72 excluded.
The overlap between 12-week SDP measured by EDSS (SDPEDSS), T25FWT (SDPT25FWT), and 9HPT (SDP9HPT) is shown in Figure 1 and the online supplementary table. About 75% of 12-week SDPEDSS are overlapped by SDPT25FWT or SDP9HPT (74% of SDPEDSS overlapped with SDPT25FWT, 30% of SDPEDSS overlapped with SDP9HPT). About 54% of 12-week SDPT25FWT and 55% SDP9HPT events are overlapped by SDPEDSS. The respective overlaps for 24-week SDP were similar (Figure 2).

Venn diagram of 12-week confirmed sustained disability progression (SDP) of prime sets and their intersections. Areas of prime sets for 12-week sustained progression in Expanded Disability Status Scale (EDSS), timed 25-foot walk test (T25FWT), nine-hole peg test (9HPT), and intersects come from the placebo arm.

Comparison of 12-week/24-week placebo sustained disability progression (SDP) event overlapping pattern. 9HPT = nine-hole peg test; EDSS = Expanded Disability Status Scale; T25FWT = timed 25-foot walk test.
Treatment effects assessed by SDP of single versus composite end points
Table 2 shows the four composite end points evaluated in this study. The estimates of rituximab treatment effects by different end points are summarized in Table 3 and Figure 3. Using SDPT25FWT alone or SDP9HPT alone led to a larger treatment effect (smaller hazard ratio) than did using SDPEDSS alone, at both 12- and 24-week confirmation. Composite A (EDSS or T25FWT or 9HPT) had much higher event rates with little change in the treatment effect estimate (Table 3) than SDPEDSS alone. Composite B (T25FWT or 9HPT) led to a larger treatment effect estimate and higher event rates than SDPEDSS alone. Composite C identified a subgroup of SDPEDSS events (75% of all SDPEDSS events) that was confirmed by T25FWT or 9HPT; thus, it had lower SDP rates than the prime set of EDSS alone. However, it led to a much larger treatment effect. In composite D, the inclusion of intersections of SDPT25FWT + SDP9HPT events, not captured by EDSS, increased both event rates and treatment effect estimate over composite C. Results using 12-week or 24-week confirmation points were in general comparable.
Composite end points.
9HPT: nine-hole peg test; EDSS: Expanded Disability Status Scale; SDP: sustained disability progression; T25FWT: timed 25-foot walk test.
Per Figure 1.
Treatment effect assessed by different SDP measures during treatment period.

Selected Kaplan–Meier curves of treatments using different composite end points. 9HPT = nine-hole peg test; CI = confidence interval; EDSS = Expanded Disability Status Scale; SDP = sustained disability progression; T25FWT = timed 25-foot walk test.
Impact of composite end points on cohort size
Table 4 shows the power to achieve statistical significance (at 0.05 significance level) using different end points for a sample size of 439 (Olympus sample size) as well as the sample size needed to keep 80% power based on the observed treatment effects and SDP rates in Olympus for respective end points. The power of a sample size of 439 to reach statistical significance in SDPEDSS alone was only 52% and 11% for 12- and 24-week confirmation points, respectively. If composite A was used, the power was 51% and 59% for 12- and 24-week confirmation points, respectively. However, the power of composites B, C and D based on a sample size of 439 were all over 80% for both 12- and 24-week confirmation points, translating to a 20%−75% reduction of the sample size needed to detect statistically significant treatment effect in delaying SDP compared with the sample size needed for 12-week SDPEDSS, the primary end point of Olympus.
Impact of choice of end points on power and sample size* calculation.
EDSS: Expanded Disability Status Scale; SDP: sustained disability progression.
Total sample size needed for a trial. 2:1 randomization, alpha level = 0.05, dropout rate = 15% based on the observed treatment effect estimates and event rate in Olympus, as shown in Table 3.
Discussion
Although Olympus failed its primary end point in 12-week SDPEDSS (with p = 0.14 for a 22% relative reduction in 96-week SDPEDSS rate), 7 the sample size of Olympus (n = 439) was originally chosen to detect a statistically significant (p < 0.05) 50% relative reduction in 96-week SDPEDSS rates and was, therefore, underpowered to detect a statistically significant 20%−30% relative reduction with this measure alone. Moreover, the directions of treatment effect estimates using SDPEDSS, SDPT25FWT and SDP9HPT were consistent in Olympus, indicating it is unlikely that the observed treatment effects in Olympus were caused by random noise. Therefore using Olympus data to evaluate different composite end points in PPMS population is valid despite the non-significant result for its primary end point. With Olympus data we showed that composite end points based on SDPEDSS, SDPT25FWT, and SDP9HPT could be superior to SDPEDSS alone in demonstrating the effects of test drug in delaying disability progression.
Our study is the first to visually display the overlapping SDPEDSS, SDPT25FWT and SDP9HPT events with both 12- and 24-week confirmation in the PPMS cohort. We limited this evaluation to placebo data to exclude a potential impact of active treatment on the different neurofunctional systems and to allow cross-comparison of results with other trials. Based on a fair degree of ‘agreement’ among the tests, we conclude that these tests measure different but overlapping functional features within ‘disability progression’ in patients with PPMS, consistent with earlier observations in relapsing MS.4,8 Thus, combining these measures into a composite is methodologically sound and would be expected to improve the sensitivity and reliability of detecting disability progression.
Although EDSS itself is a composite measure, alternative composite end points are considered as potential solutions to overcome the operational and conceptual limitations of the EDSS in detecting MS disability progression. 2 The MS Functional Composite (MSFC) was the first major alternative to the EDSS proposed for MS trials. 9 It combines changes in the paced auditory serial addition test (PASAT), T25FWT, and 9HPT into a composite score, but it has not gained broad acceptance, primarily because of concerns about the clinical meaningfulness of its dimensionless reduction to a Z score.2,10 One way to improve on this aspect is to define sustained disability progression within each test first and then to combine SDP results from each test to make sure the progression has truly occurred. We took this approach in this analysis with a cut-off of 20% to define disease progression in T25FWT and 9HPT. This cut-off has been shown to reliably indicate a consistent clinical impact on disability and corresponds to a change in functional ability discernible to patients.4–6 With this cut-off along with standard cut-off used for EDSS, the SDP results from individual EDSS, T25FWT, and 9HPT as well as composite end points including these SDP results are considered clinical meaningful. An important future research topic will be correlating these clinical neurofunctional outcomes to other disability measures reported by patients to assess direct impact on patient daily life.
We limited the components of composite end point to EDSS, T25WT, and 9HPT in this study because PASAT results are strongly impacted by training effects.11–13 Such problems do not occur with the T25FWT and are less pronounced with the 9HPT.12,13 This may partly explain why the MSFC was less responsive than EDSS to detect disability progression in a PPMS cohort. 14
The union of basic tests (composite A, only ‘OR’ criteria) increased the sensitivity for detecting disability progression by including SDP events not captured by EDSS. However, its reliability was inherently not increased because all SDPEDSS events were taken and all additional events added were not confirmed by EDSS. With higher SDP rates and comparable treatment effect estimate, composite A can potentially increase power of detecting treatment effect than EDSS alone. Goodkin et al. showed similar results for composite A at the 24-week confirmation point in patients with relapsing MS using data from a phase III study of intramuscular recombinant interferon beta-1a for disease progression. 8 The treatment effect of intramuscular recombinant interferon beta-1a in delaying disability progression was significant by both SDPEDSS and SDPT25FWT alone. 8 SDP in EDSS, 9HPT, T25FWT, or Box-and-Block Test (with 24-week confirmation; only 2% SDP were captured by Box-and-Block Test alone) resulted in a gain in power (driven by higher event rate) with relatively similar treatment effect estimates as that from SDPEDSS alone. Some ongoing progressive MS trials have also adapted the concept of composite A by adding 20% SDP on T25FWT or 9HPT as components of their primary composite end points (ASCEND: natalizumab in SPMS; INFORMS: fingolimod in PPMS), and should provide more data on this end point in the near future.
Although composite B did not include EDSS outcome directly, it still covered around 75% of SDPEDSS events and led to a larger treatment estimate and higher event rate than SDPEDSS alone for both 12- and 24-week confirmation points in OLYMPUS. This composite end point is similar to MSFC progression −20 defined by Rudick et al. except that SDPPASAT was not included. 4 Rudick showed that in RRMS cohorts (Affirm and Sentinel, pivotal studies of natalizumab), MSFC progression −20 with 12-week confirmation covered about 45% of SDPEDSS and led to a comparable treatment effect estimate and higher event rate than that from SDPEDSS alone. These results indicate composite B might cover a higher percentage of SDPEDSS events in PPMS than RRMS population and could be an alternative to SDPEDSS alone in both RRMS and PPMS populations.
Composite C used ‘AND’ criteria to increase the reliability of the SDPEDSS event, but had a lower SDP event rate because it only included SDPEDSS confirmed by at least one of the other two tests. Regardless of lower SDP rate, it detected a much larger treatment effect using a subgroup (75%) of SDPEDSS events. Composite D further increased sensitivity of composite C by including events captured by both SDPT25FWT and SDP9HPT events but not captured by EDSS (through ‘OR’ criteria). Composite D detected the largest treatment effect with a SDP rate similar to that captured by SDPEDSS alone. In the past, improving the sensitivity of disability end points by combining end points using ‘OR’ criteria, lowering the detection level for disability progression, was the primary focus.4,15 Our results show that in PPMS cohorts increasing reliability using ‘AND’ criteria has comparable or even larger effects, with the largest increase in power achieved by using the combination of ‘OR’ and ‘AND’ criteria.
The sample size needed (or power for a fixed sample size) for achieving a statistically significant result is determined by both the magnitude of treatment effect and SDP event rates. Based on the SDP event rates and treatment effect estimates observed for these composite end points in Olympus, the sample size needed to achieve significant result with composite B, C or D may be one half or one third of that needed for SDPEDSS alone in the PPMS cohorts. For composite A, substantial reduction in sample size than that using SDPEDSS alone could only be achieved by using 24-week confirmation in Olympus, which is consistent from the results from a RRMS cohort. 5 Although 24-week SDP usually lead to larger treatment effect estimates but lower SDP event rates compared with 12-week SDP, the increase in treatment effect estimate is more substantial for composite end points that only use ‘OR’ criteria (such as composite A and B), mainly because these composite end points have relatively high variability. Therefore, even with a lower event rate, the 24-week SDP can bring higher power for composite end points that only use the ‘OR’ criteria compared with the 12-week SDP for the same composite end point.
In conclusion, together with results from relapsing MS trial,4,8 our study confirmed that composite A with 24-week confirmation and composite B could be a more sensitive measure than SDPEDSS alone for analysis disability progression in MS trials. Composites C and D could substantially increase the efficiency of a clinical trial in PPMS. However, using composites C and D in relapsing MS cohorts should be studied first with existing data from relapsing MS trials. With more supportive data from other trials, these composites may become the new standard disability end points for future MS clinical trials. Furthermore, when additional clinically meaningful disability measures are validated (e.g. 6 minute walk, accelerometer, cognitive, visual measures), these could be added to or replace one or more of the composite components we tested using the same algorithm.
Footnotes
Acknowledgements
We thank each of the Olympus investigators, including Drs Virender Bhan and Michael P Biber; Staley A Brod; Peter Calabresi; Denise Campagnolo; Mark Cascione; Nelson (Kanter) Cooke; Joanne A Cooper; John Corboy; Bruce Cree; Anne Cross; Pierre Duquette; Stranton Elias; Corey Ford; Mark S Freedman; Mitchell S Freedman; Suzanne K Gazda; Francois Grand’Maison; Michael (Lava) Gruenthal; Stuart Hoffman; William Honeycutt; John Huddlestone; Bruce Hughes; George J Hutton; Daniel Jacobs; Francois H Jacques; Lloyd H Kasper; Lorne Kastrukoff; Michael D Kaufman; Omar Khan; Bhupendrea O Khatri; Mariko Kita; Yves La Pierre; Sharon Lynch; Clyde Markowitz; Michele Mass; David Mattson; Aaron Miller; Harold Moses; Paul O’Connor; Hillel Panitch; Mary Ann Picone; Kottil Rammohan; Anthony T Reder; Peter Riskind; Syed W Rizvi; Howard Rossman; Steven R Schwid; Stuart Shafer; Virginia Simnad; Michael Stein; William H Stuart; Olaf Stuve; Ben Thrower; Stephen E Thurston; Michael (Mihai) Vertino; Leslie P Weiner; Dean Wingerchuk; and MC Yeung. Support for third-party editorial assistance for this manuscript was provided by Genentech Inc.
Conflict of interest
Dr Zhang is an employee of Genentech, a member of the Roche Group of companies.
Dr Waubant has received honoraria from Teva, Questcor, and Genzyme for three educational lectures. She is on an advisory board for a trial of Novartis.
Dr Cutter has participated in DMC sponsored by Sanofi-Aventis, Cleveland Clinic, Daiichi-Sankyo, GlaxoSmithKline Pharmaceuticals, Genmab Biopharmaceuticals, Eli Lilly, Medivation, Modigenetech, Ono Pharmaceuticals, PTC Therapeutics, Teva, Vivus, University of Pennsylvania, NHLBI, NINDS, and NMSS. He has received consulting and speaking fees from Alexion, Bayhill, Bayer, Novartis, Consortium of MS Centers (grant), Genzyme, Klein-Buendel Incorporated, Nuron Biotech, Peptimmune, Somnus Pharmaceuticals, Sandoz, Teva Pharmaceuticals, UT Southwestern, and Visioneering Technologies, Inc.
Dr Wolinsky has received compensation for service on steering committees or data monitoring boards for Eli Lilly, Novartis Pharmaceuticals, Sanofi, and Teva Pharmaceuticals; as a consultant to Acetilon, Athersys, Inc., Bayer HealthCare, Celgene, Genentech, Genzyme, Novartis, F. Hoffmann-La Roche, Ltd., Jansen RND, Teva and Teva Neurosciences, and XenoPort; has received honoraria from Biogen Idec, the Consortium of Multiple Sclerosis Centers, Medscape CME, Prime, Serono Symposia International Foundation, Teva Pharmaceuticals and Teva Neuroscience; and has received or receives research support from Genzyme, Sanofi, the National Institutes of Health, the Clayton Foundation for Research, and the National Multiple Sclerosis Society through the University of Texas Health Science Center at Houston (UTHSCH) and royalties for monoclonal antibodies out-licensed to Chemicon International through UTHSCH.
Funding
This was supported by F. Hoffmann-La Roche Ltd and Biogen Idec.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
