Abstract
Summary
To enable an assessment of the costs and benefits of a new health technology one should use a range of outcome measures, including medical, psychosocial and economic. Therefore, unless a patient-reported outcome as well as clinical outcome is assessed in a study, the effect of a health technology on the patient will remain unknown as two therapies may have similar clinical consequences but different impacts upon the quality of the life of the patients. An important issue when designing a study with a new patient-reported outcome is the quantification of an effect size. Through a case study we highlight how simple calculations can enable the estimation of the effect sizes if there is information on established outcomes. This is done by mapping changes on the new scale to clinically relevant and important changes on established scales. We recommend the approaches described in this paper be considered for the quantification of important treatment effects when designing a clinical trial with a new patient-reported outcome measure.
1 Introduction
To enable an assessment of the costs and benefits of a new health technology one should use a range of outcome measures including clinical, psychosocial, patient-reported and economic. In this context, unless patient-reported outcomes (PROs) as well as clinical outcomes are assessed, the effect of the new health technology on the patient may remain unknown as the two therapies (new and existing) may have similar clinical consequences but different impacts upon the quality of life (QoL) of the patients. A treatment that enables a patient to feel better (or even not worse) while also treating the clinical condition is important as it could, for example, affect compliance with a medication. Indeed, a patient may trade off some clinical benefit for an improvement in QoL.
There are clinical conditions, for example osteoarthritis, where a demonstration of efficacy may require statistically significant improvement in a number of outcomes including a clinical assessment of the condition: patient-reported pain and patient-reported physical function. This is because to show efficacy, it may be that one endpoint or outcome in isolation is not sufficient and therefore to show a benefit evidence must be provided on (here) three endpoints, which represent different aspects of the condition–clinical and patient reported. Having co-primaries such as this is referred to as “multiple must-win” and impacts on sample size calculations. 1
Researchers have used a variety of names to describe QoL measurement scales. Some prefer to use the term health-related quality of life (HRQoL or HRQL) to stress that we are only concerned with health aspects. Others have used the terms health status or self-reported health. The United States Food & Drug Administration has adopted the term Patient-Reported Outcome (PRO) in its guidance to the pharmaceutical industry for supporting labelling claims for medical product development. 2 However, not all people who complete such outcomes are ill and patients and hence PRO could legitimately stand for person-reported outcome. Mostly, we shall assume that the QoL instrument or outcome is self-reported, by the person whose experience we are interested in, but it could be completed by another person or proxy. The term health outcome assessment has been put forward as an alternative which avoids specifying the respondent.
In truth QoL is a complex concept with multiple aspects. 3 These aspects (usually referred to as domains or dimensions) can include: cognitive functioning; emotional functioning; psychological well-being; general health; physical functioning; physical symptoms and toxicity; role functioning; sexual functioning; social well-being and functioning; spiritual/existential issues and many more. This paper will assume a wide definition for QoL and will describe the design and interpretation of responses in context with responses observed on other patient assessments–clinical or patient-reported.
2 Designing based on a HRQoL endpoint
An important step in the design of a clinical trial is the estimation of the sample size. The sample size should always be large enough to yield a reliable conclusion and be the smallest necessary to reach this conclusion and the intended aim of the trial. If the sample size is too small the usefulness of the results will be reduced because of the inadequate number of patients and clinically important changes may be concealed. However, if the sample size is too large, the trial will incur a waste of resources and patients may continue to receive a less effective or even dangerous treatment.
An informed calculation of the sample size is an essential prerequisite for any trial. When an investigator is designing a study to compare the outcomes of an intervention, an essential step is the calculation of sample sizes that will allow a reasonable chance (power) of detecting a predetermined difference (effect size) in the outcome variable if there is truly a difference, at a given level of statistical significance.
Sample size is critically dependent on the purpose of the study, the outcome measure and how it is summarised, the proposed effect size and the method of calculating the test statistic. Unlike power and the level of statistical significance, for which precedents dictate values, the “target” or anticipated effect size under the alternative hypothesis must be determined for each trial based on experience, published data or pilot studies. The target or anticipated effect size under the alternative hypothesis is usually the smallest difference or minimum important difference (MID) that is considered clinically important–the MCID (minimum clinically important difference).
The definition of a clinically important difference is based on such things as experience of the outcome measure of interest and the population in which the outcome measure is being applied with the mean effect size of interest at a population level typically being appreciably less than the smallest change that is important to an individual patient.
The quantification of an effect size is not straightforward for an established endpoint or outcome measure. However, for a novel endpoint it is particularly difficult as the clinical experience with using the endpoint has not been established to evaluate what a clinically meaningful difference is.
An accurate quantification of an effect size with PROs is important as many QoL instruments are able to detect statistically significant mean changes or differences that are very small. Therefore it is important to consider whether such changes are clinically meaningful. To help interpret the scores of QoL instruments it is important to identify (and specify) the smallest difference or change in QoL scores between individuals or groups that is clinically and practically important. This benchmark value is usually called the minimum important difference (MID).
An MID is usually specific to the population under study. The FDA describes a variety of methods for determining minimum important difference.
2
Using a clinical or non-clinical anchor-based approach (mapping changes in QoL scores to clinically relevant and important changes in non-QoL measures of treatment outcome in the condition of interest). Mapping changes in QoL scores to other QoL scores to arrive at an MID that is appreciable to patients (e.g. when multi-item QoL scores are mapped to a single question asking the patient to rate their global impression of change since the start of treatment). Using a distribution-based approach (e.g. definition of the MID as 0.5 times the standard deviation of the QoL scores). Using an empirical rule (e.g. 8% of the theoretical or empirical range of scores).
It is an anchor-based approach that is being considered in this paper where we are looking to establish a clinically meaningful difference for a new QoL outcome by using the association of the new QoL outcome with an established clinical endpoint or outcome. Thus, if a clinically meaningful difference is known for these clinical endpoints then this association may be used to quantify a clinically meaningful effect for the new QoL outcome.4,5
Note there may be circumstances where it is more reasonable to characterise the meaningfulness of an individual’s response to treatment rather than a group’s average response. This leads to the definition of an individual responder to treatment. This individual response to treatment definition should be based upon pre-specified criteria. Examples include categorising a patient as a responder based upon a pre-specified change from baseline in one or more scales, a change in score of certain size or greater (e.g. a 10-point change on a 100-point scale) or a percentage change from baseline. The definition of a responder should be backed by empirical evidence that such a change, in QoL scores, truly benefits individual patients. The case study describes the calculations for quantifying an effect size in a stroke trial using a 16-question version of stroke impact scale (SIS-16). 6
2.1 Case study
The SIS-16 is a 16-item self-completed questionnaire, assessing physical functioning following a stroke. The 16 items ask questions on a number of aspects of a patient’s self-reported health including hand function, strength and mobility. The 16 items are combined to generate an overall score ranging from 0 (poor) to 100 (good).
For the purpose of this paper, it is being supposed that the stroke impact scale (SIS-16) is to be assessed at 3 months and will be the primary endpoint for a trial in stroke patients to determine the superiority or effectiveness of a new technology (active treatment) compared to an existing treatment or placebo.
We shall assume that data is available on the SIS-16 alongside data on a well-established stroke outcome measure. The Rankin Scale is a commonly used 7-point ordinal scale for measuring the degree of disability or dependence in the daily activities of people who have suffered a stroke, and it has become the most widely used clinical outcome measure for stroke clinical trials. The scale runs from 0 to 6, running from perfect health without symptoms (score 0) to 6 (dead).
Range of scores and how efficacy is determined for SIS-16 and Rankin.
2.2 The methodology
The estimated treatment effect for each established health outcome was taken as an odds ratio. A previous empirical distribution for the health outcome scores at 3 months was assumed to be the prospective placebo response. The health outcome score distribution on active was estimated from the placebo response under the assumption of proportional odds. 8 In the case study, we will describe how the odds ratio can be estimated.
To estimate an effect size for the SIS-16, the following four steps were applied.
For each Rankin category the mean SIS-16 response was calculated. The expected proportions on active and placebo were estimated in each category. Multiplying the mean SIS-16 response for each Rankin category by the expected proportion and then summing an expected mean response, on the new SIS-16 scale, was obtained for both active and placebo. The difference in the mean overall responses on active and placebo were taken as an estimate of treatment effect for SIS-16 equivalent to the effect of interest on Rankin scale.
To highlight the calculations, the worked example will be repeated twice. Once assuming the scales are dichotomised and once assuming the scale is ordinal.
Figure 1 illustrates how the association between Rankin and SIS-16 will be used in the context of this paper. We undertook a simulation of 1000 studies each of size 1000. Each study had two Normally distributed outcomes with various correlations (ranging from 0.00 to 0.95) between the two outcomes. For each simulation we then uniformly categorised one outcome to 20 categories (so there was an equal sample size in each category) and calculated the mean of the second outcome for each category. We then estimated the correlation between this new 20-point categorised scale and the means on the other scale. We then repeated this 1000 times (and took the average) for each correlation. We then further repeated the simulation for 10 categories and 5 categories.
Correlations between a categorised outcome and the means for each outcome for different numbers of categories.
What the simulations show is that the correlation between the means and the categories is very high–higher in fact than the correlation between the original data. It is this feature which we will use as the basis of our work to estimate effect sizes.
2.3 Worked example–dichotomised response
For expository purposes to highlight the calculations the worked example will be undertaken first by dichotomising each of the scales around the clinically meaningful cut-offs for the established outcome.
Different clinically meaningful cut-offs were investigated on the Rankin to determine the associated effects with SIS-16. For the Rankin scale for expository purposes, a 10% increase in the proportions in the sample scoring 0 or 1 or 0, 1 or 2 is regarded as clinically meaningful difference or improvement in functioning.
Worked example of the effect size estimation through associating SIS-16 with Rankin–dichotomised scale.
Treatment effects for SIS-16 associated with effects on the Rankin–dichotomised Scale.
If increasing the proportion of subjects on the Rankin scale who score either 0 or 1 or 0 to 2 by 10% on the active compared to placebo treatment is regarded as a clinically important effect, then these simple calculations show that a mean effect of around 4 points on the 100-point SIS-16 scale would be associated with this effect on the Rankin scale.
2.4 Quantifying the odds ratio
The effect size when expressed as an odds ratio is defined as:
As an odds ratio implies, it is a ratio of two odds. For example, on a dichotomised scale an odds of 2:1 would mean that for every three subjects on a control regimen we would expect one event, i.e. non-events are twice as numerous as events. Odds of 4:1 on the investigative regimen would mean non-events are four times as numerous.
For ordinal data the interpretation is more complicated as what one is assuming is proportional odds between the treatments across the full scale of the health outcome parameter. This implies that the odds ratios are identical for each pair of adjacent categories throughout the scale. What this means practically can be highlighted by example.
Responses anticipated on active under the alternative hypothesis for different odds ratios and the response on placebo (π P ).
The assumption of proportional odds implies that if instead of using a Rankin score of ≤ 1 as a cut-off we had used a Rankin score of ≤2, we would still obtain OR2 = 0.67 and so on for OR3 = OR4 =…etc
Although the actual observed odds ratios might differ from each other across the scale, the calculations of sample size are robust to departures from this ideal, provided all the odds ratios indicate an advantage to the same treatment. 5
Given the anticipated response in the control arm (π
P
) and a given odds ratio (OR >1), the response on active (π
A
) can be estimated from
2.5 Worked example ordinal response
Worked example of the effect size estimation through associating SIS-16 with Rankin–ordinal scale assuming a odds ratio of 1.5 a clinically meaningful effect size on the Rankin.
The same effect size is used for the ordinal response as for the dichotomised scale, but here the absolute difference is converted to an odds ratio which is then applied across the full scale to estimate the anticipated active response as in Table 5.
Treatment effects for SIS-16 associated with effects on the Rankin - ordinal scale.
It is reassuring that the ordinal calculations are consistent with those for the more simple dichotomised approach. However, due to the additional information used in the calculations, it is recommended that the ordinal approach be used for effect size estimation.
Odds ratios on Rankin and corresponding mean differences on SIS-16.
3 Conclusions
The results in this case study highlight how simple calculations can enable us to estimate a clinically meaningful difference or effect size for a new PRO (or indeed any outcome) if there is information on a clinically important difference for an established outcome and the two outcomes have been used together. This is done by mapping changes in QoL scores to clinically relevant and important changes in non-QoL or other QoL measures of treatment outcome in the condition of interest. We have also shown how bounded estimates (confidence intervals) for the clinical important effect on the new QoL measure can also be derived to show the uncertainty or imprecision of our estimate of the clinical important effect on the new QoL outcome.
In the worked example, we anchored the QoL improvements to improvements in a clinical scale–we have assumed a positive association. This may not always be appropriate and in fact a patient may trade off some clinical benefit of a treatment for an improvement in QoL–negative association.
The method requires a dataset where both the new outcome has been used alongside an existing “gold standard” outcome, which already has a minimum clinically important difference (MCID) defined. This could be considered a strong assumption/or requirement. Upon saying this part of the validation of a new QoL scale should include co-administration with a standard scale and/or with a clinical outcome assessment. So in reality one is likely to have data on both outcome measures.
We recommend the approaches described in this paper be considered for the quantification of treatment effects when designing a clinical trial.
