Abstract
The treatments under comparison in a randomised trial should ideally have equal value and acceptability – a position of equipoise – to study participants. However, it is unlikely that true equipoise exists in practice, because at least some participants may have preferences for one treatment or the other, for a variety of reasons. These preferences may be related to study outcomes, and hence affect the estimation of the treatment effect. Furthermore, the effects of preferences can sometimes be substantial, and may even be larger than the direct effect of treatment. Preference effects are of interest in their own right, but they cannot be assessed in the standard parallel group design for a randomised trial. In this paper, we describe a model to represent the impact of preferences on trial outcomes, in addition to the usual treatment effect. In particular, we describe how outcomes might differ between participants who would choose one treatment or the other, if they were free to do so. Additionally, we investigate the difference in outcomes depending on whether or not a participant receives his or her preferred treatment, which we characterise through a so-called preference effect. We then discuss several study designs that have been proposed to measure and exploit data on preferences, and which constitute alternatives to the conventional parallel group design. Based on the model framework, we determine which of the various preference effects can or cannot be estimated with each design. We also illustrate these ideas with some examples of preference designs from the literature.
1 Introduction
A primary objective of most randomised trials is to estimate the effect of treatment (experimental vs. control) on patient outcomes. In theory, the two treatments under comparison should have equal value and acceptability – a position of equipoise – to study participants (and also their clinicians in the case of a randomised clinical trial). However, it is unlikely that true equipoise exists in practice. For instance, even if there is equipoise overall, some patients might prefer to accept short-term risks associated with surgery in exchange for better outcomes in the longer term, while other patients might be risk averse in the short term, and would therefore prefer a nonsurgical treatment for their disease. Importantly, the preferences that patients have for one treatment or the other may be related to study outcomes, and hence affect the estimation of the treatment effect (TE) in a trial. Additionally, the interaction between preference and treatment is of intrinsic interest. For example, patients who would opt for an exercise- and diet-based approach to the treatment of their hypertension and elevated cholesterol may have different expected outcomes from patients who would select a medication-based strategy, these effects being independent of any actual differences in effectiveness of these two treatments.
Although the estimated TE will be unbiased by preference effects (PEs) in randomised, blinded studies, evaluating the impact of preferences on trial outcomes is often a reasonable scientific objective. Unfortunately, the impact of participant preferences cannot be assessed in the standard parallel group design for a randomised trial, because these effects are confounded with the TE. Furthermore, in some circumstances, the effects of preferences can be substantial, and may even be larger than the direct effect of treatment. If so, investigators would certainly want to know that. Estimating how preference and treatment might interact has direct implications for shared decision making.
In this paper, we describe a model to represent the impact of preferences on study outcomes, in addition to the TE. In particular, we describe how study outcomes might differ (even if treatment were ineffective), in particular between participants who would select one treatment or the other, if they were free to do so; this is characterised by a so-called selection effect (SE). Additionally, we investigate the difference in outcomes depending on whether a participant receives his or her selected treatment or the alternative treatment, which we characterise through a PE. We then discuss several study designs that have been proposed to measure and exploit data on preferences, and which constitute alternatives to the conventional parallel group design. (In our later discussion, we consider how the measurement of preferences relates to whether a study design might remain blinded or not.) Based on the model framework, we determine which of these various effects can or cannot be estimated with each design. We also illustrate these ideas with some examples of preference designs from the literature.
2 Methods
We will focus on preferences arising from study participants. (Some related aspects of clinician preferences in clinical trials are mentioned in Section 4.) We begin with a simple linear model, which includes terms representing the effects on a continuous outcome variable Y of the treatment i actually received by participant k, of the treatment j that a participant would select (if allowed to do so), of the combination (i, j) of the actual and selected treatments, as well as the overall mean μ. Specifically
the treatment effect (denoted by TE), which is usually a direct effect of treatment. In terms of the model parameters, TE = τ1 − τ2; the SE, which is the expected difference in outcomes between those who would choose A (if allowed to do so) and those who would choose B (if allowed to do so). Under the model, SE = ν1 − ν2; the PE, given by PE = π11 the additional difference between outcomes on A versus B for those who receive their selected treatment and those who do not. We will refer to this as the concordance effect (CE), because it represents the difference in TE for people whose preferred treatment is concordant with the actual treatment, versus those where they are discordant. In the model notation, CE = (π11
The term ɛijk arises from random error, with expectation zero. We will limit attention to the common situation where there are only two treatments (A and B) under comparison, labelled i = 1 and i = 2, respectively. As usual, constraints on parameters are required in order to avoid redundancy. This model’s three constraints are:
We now describe four designs for randomised trials, three of which incorporate preference data, and then we identify which of the effects described above are estimable from them. Figure 1 shows the basic architecture of these four designs schematically.
Main design features of four randomised trial designs. (a) Conventional parallel group design. (b) Two-stage randomised design. (c) Fully randomised preference design and (d) Partially randomised preference design.
2.1 Conventional parallel group randomised trial
The relevant observable outcomes in this design are the mean responses in the two treatment groups. From equation (1), their expectations are given by
2.2 Two-stage randomised design
In this design, a randomly selected subgroup of participants is allowed to choose their treatment, while the remaining participants are randomised1–5 (see Figure 1). The treatment preferences of those in the random arm are not known. This design is sometimes referred to as the ‘preference vs. conventional’ design. There are outcomes observable in four participant groups, corresponding to the two treatments in each of the choice versus random arms of the study. Again using model (1), their means are
In order to estimate SE and PE, we first note that the preference rate α for treatment A is an unknown parameter that must be estimated from the empirical distribution of treatments adopted in the choice arm. Second, we see that direct estimates of μ12 and μ21, the mean outcomes for persons who receive their less desirable treatment, are not observable, but estimates of these quantities are nevertheless required. The solution to this challenge is to infer from the first stage of randomisation (into the choice and random arms) that the expected distributions of preferences in the two arms are equivalent, and hence one can impute suitable estimates of μ12 and μ21. As described in detail elsewhere,
5
estimators of SE and PE are then given by
This approach to the analysis exploits the information from the treatments actually taken by participants in the choice group, and it is this feature that provides the estimability of SE and PE, as well as the usual TE. Because the two stages of randomisation for participants in the random arm produce the same expected result as a single randomisation (‘random + random = random’), an alternative analytic approach is to make comparisons between three randomised groups (randomised to A, randomised to B, or randomised to choice), but without using the additional information on treatments chosen. Comparisons of the A or B groups with the choice arm as a whole then provide valid estimates of TEs associated with those interventions, but they do not allow exploration of how those effects are influenced by participant preferences.
2.3 Fully randomised preference design
In this design, the stated preferences of all participants are identified at baseline, but they are nevertheless all randomised to treatment (see Figure 1). This is in contrast to the situation in the two-stage design, where participants in the choice arm are permitted to actually exercise their preference. We discuss possible differences between stated and exercised preferences later. In this design, outcomes are observable in four participant groups, with means derived from equation (1) as follows
Note that the preference subgroup estimators of TE are each individually ‘valid’ in the sense that the randomisation within subgroups protects, on average, against confounding by known or unknown variables. However, these estimates pertain to subgroups of individuals with different preferences, and hence the associated TEs may be different in those subgroups, as shown by the addition of πij terms in equations (6) and (7).
We remark that an estimate of TE is also available by ignoring the baseline information on preferences, and simply comparing outcomes in the two treatment groups. This would be equivalent to the analysis of the conventional parallel group design; the estimate of TE is unconditionally unbiased, but conditional on the empirical distribution of preferences not being exactly equal to its expectation (as will happen in almost all trials), the TE estimate will be potentially confounded by preference effects. Note that a weighted average of (6) and (7) (using empirical estimates of α and 1− α) will yield the same estimate of TE as in the parallel group design.
We may also consider whether SE or PE is estimable with this design. An estimator of PE is directly available as the difference in TEs within the A-preferer and B-preferer subgroups, i.e.
The best available estimator of SE can be constructed as the mean difference between outcomes for the A preferers and B preferers, i.e.
To summarise, the fully randomised preference design yields three possible estimators of TE, all of which are biased by πij terms, but which nevertheless are valid estimators of the TE within stated preference subgroups, or averaged over those subgroups; an unbiased estimator of PE is available, but an unbiased estimator of SE is not possible in general.
2.4 Partially randomised preference design
In this design, participants are first asked whether they have a preferred treatment. Those expressing no preference (i.e. they are undecided) are randomly assigned to treatment; whereas, those with a preference are allowed to receive it. As in the two-stage and fully randomised preference designs, there are four observable groups, with means from equation (1) as follows
One possible estimator of TE comes from the group of undecided participants (which has the advantage that they are all randomised to treatment), which we denote by
The comparison of outcomes between patients who choose their treatment is observational, and hence is confounded by SEs and by a term π11 − π22. Finally, a combination of these two approaches gives an estimator with the expectation
An estimator of SE is available through
PE is not estimable at all, even if the previous assumptions are made, because none of the observed means involve π12 and π21 (which are components of PE), and unlike the situation in the two-stage design, imputed estimates of these parameters are not possible.
A closely related design is the comprehensive cohort study, in which nonrandomised participants are able to choose their treatment, or receive their treatment through the choice of their clinician. However, the precise ways in which this design has been implemented have been quite variable in practice, and the estimation of SE or PE is correspondingly confused, as we will see in the illustrative examples below.
To summarise, the partially randomised preference design gives three possible estimators of TE, all of which are biased by PE terms. An estimate of SE is available, which is unbiased if certain assumptions about the PEs are correct; however, none of these assumptions are testable. PE is not estimable.
Summary of estimable effects with various randomised trial designs.
Valid in preference subgroups.
Potentially biased.
3 Illustrative examples
We now illustrate the three main designs (two-stage, fully randomised and partially randomised) using selected examples that have utilised preference data. We indicate the diversity of settings in which the designs have been applied and the different approaches that have been used for statistical analysis. Some studies have estimated preference and selection effects as described above. For some other studies, we have been able to reanalyse the reported data to estimate these effects, even if the original authors had not done so, and comparisons with the published results are then possible.
3.1 Two-stage design
Although less commonly used than the parallel group design, the two-stage design has been adopted in a wide variety of research areas. Examples include randomised trials of a decision aid describing management strategies for women with atypical cells detected during routine cervical screening, in which informed choice between two alternative management strategies using a decision aid was compared to no choice; 3 medical versus surgical interventions for heavy menstrual bleeding;7,8 the impact on student examination scores of background music while studying; 9 medication versus cognitive behavioural therapy 10 and behavioural versus cognitive therapy in the treatment of depression; 11 self-directed versus group behavioural programs for women with cardiac disease;12,13 two placebos (and control) in the minimisation of pain perception; 14 motivated reasoning; 15 alternative educational curricula for diabetics 16 and dietary interventions on weight loss.17–19
The authors of these studies have reported their results in various ways, not all of which are sufficiently detailed to permit an analysis that takes full advantage of the preference data. Some of these studies also involve some deviations from the standard design, as we will mention later. In one of our own studies, we have previously estimated selection and preference effects and showed that they can sometimes be important, even if the TE is small. 4 For the present paper, we have chosen a weight loss trial to further illustrate this type of analysis.
To illustrate the calculation of preference-related effects in a two-stage design, we use data from the PREFER randomised trial reported by Burke et al.17–19 This study was intended to evaluate the impacts of weight loss interventions in sedentary and overweight adults. Two diets were compared: a ‘standard’ weight loss diet and a lacto-ovo-vegetarian diet. In the first stage, participants were randomised either to be allowed to choose their own diet or to be assigned a diet at random in the second stage. There were also dietary and exercise goals conveyed during behavioural therapy sessions during the first 12 months, followed by a 6-month maintenance phase with no further contact with participants until the final assessment at 18 months.
Estimated treatment, selection and preference effects from a randomised trial of weight loss interventions.
Note: M: mean; SD: standard deviation; BMI: body mass index; LDL: low-density lipoprotein; HDL: high-density lipoprotein.
All outcomes are the percentage change from baseline to the final analysis after 18 months of follow-up.
p Values are two-sided.
There were 200 participants, randomised to the choice or random arms in a 3:2 ratio. This was because the investigators anticipated that the LOV diet would be less popular in the choice arm, based on pilot data that indicated about two-third of the participants would select the STD diet. At the second stage of randomisation, all participants who chose LOV received it; however, only 48 of the 63 persons who chose STD were retained, while the remainder were not involved further in the study. This was to avoid the STD subgroup being ‘excessively larger’ than the LOV subgroup. We remark that an analysis of the preference data in such trials can accommodate unequal sized groups without difficulty; hence, the main justification for discarding participants in this way would be on the basis of study efficiency, to achieve approximately equal numbers of participants in each study subgroup. 5 Individuals randomised to the random arm were assigned randomly in equal numbers to the two diets.
Table 2 shows sample sizes, means and standard deviations for seven study outcomes. After the exclusion of persons who no longer met eligibility criteria after enrolment, there were 176 participants available for analysis (48 and 45 assigned randomly to STD and LOV in the random arm and 48 and 35 who chose these diets in the choice arm). Among these, 132 completed the intervention program, but all 176 are included in the analyses we report here.
Table 2 also shows the estimated treatment, selection and PEs for each of the seven outcomes, together with their z-test statistics and p values.1,5 Each outcome is shown as the percentage change from baseline up to the 18-month analysis. First, considering the body weight outcome, we see that the TE is very small and nonsignificant, indicating little overall difference between the STD and LOV diets. The SE has a positive but not significant value of 2.61, indicating that the expected weight change is 2.61% larger in participants who would prefer STD than in those who would prefer LOV. The PE is the largest of the three measures, at 7.11, expressing the difference in the expected percentage difference in weight change between treatments for participants who received their preferred diet versus those who did not. Given that negative weight changes are the goal of the study, the results indicate that better outcomes were observed among those who did not receive their preferred treatment. The results for BMI are qualitatively similar.
The remaining outcomes show a variety of patterns for the relative magnitudes and statistical significance of TE, SE and PE. For the waist circumference, PE is again the largest effect, but here is not significant. For low-density lipoprotein, all three effects are moderately large (SE being the largest), while for HDL, PE is the largest; however, none of these effects is significant. Glucose shows a statistically significant TE, with the LOV having the larger percentage decreases; its PE has a value of similar magnitude but is not significant; its SE is the largest effect, with its positive and significant values indicating larger decreases in the LOV group. None of the effects is significant for the insulin outcome.
The investigators’ own analysis of these data cannot be directly compared with the results above, because they use individual participant values measured repeatedly over time (and these are not accessible from the published report), whereas we have used the 18-month values only. Furthermore, they used an analysis of variance (ANOVA) for outcomes in the four observable study subgroups (STD and LOV, by choice or random arm), with factors preference, diet and time. This approach examines the effect of preferences only through the interaction of preference group and diet group, but it does not lead to estimates or inferences concerning the preferences of individuals, as does the analysis we have shown here. The interpretation of the preference effects is thus different in the two analyses: in the investigators’ ANOVA, we can identify the impact of being in the choice or random method of assigning their dietary intervention, while our analysis identifies differences in outcomes for individual participants who would prefer one diet or the other. Note that the latter utilises data from all participants for this inference, even though preferences are only actually realised for participants in the choice arm.
The authors of the PREFER study had hypothesised that the freedom to choose a diet would achieve greater weight loss than being randomly assigned to a diet. In fact, the reverse was observed. The PREFER authors present several possible explanations for this surprising result; first that those in the random arm were more determined to succeed despite their assigned treatment. They also suggest that these people may have forgotten their original preferences, which would therefore become irrelevant to the 18-month outcome; but, in contrast, our own analysis provides evidence of persistent selection and preference effects for some outcomes. Another possible explanation is that participants who actually received their preferred treatment may have expected more from it, particularly bearing in mind that they had been able to choose, and were then disappointed in the results. Finally, we remark that participants in the choice arm may have been more likely to adopt the study diet that was more similar to their prestudy diet, which would then imply less potential for weight loss through the dietary intervention in the study.
The findings in PREFER contrast with the IMAP study 4 where choice of management (between either human papilloma virus testing or repeat Pap testing) for a screen-detected mildly abnormal Pap smear was associated with a positive PE on quality of life. Mental Component SF36 scores showed evidence of a PE through a 6-point relative improvement (out of a possible 100) among participants who received their preferred treatment (and with a moderate standardised effect size 0.61, p = 0.07); this indicates that there was a larger TE among women who received their preferred treatment compared to those who did not. These two examples illustrate that it is possible for the pattern of PEs to be in the either direction.
The analysis we have given here shows the additional insight it provides on important determinants of the outcome beyond the TE itself. In the PREFER example, several outcomes had selection or preference effects that were at least as large as the direct TE, for instance, where the PE dominated the TE for the body weight outcome.
Others have approached the analysis of two-stage designs somewhat differently. Using the Women Take Pride study of group versus self-directed management of heart disease as an example, Janevic et al. 13 compare the random and choice arms at baseline, and also estimate the effect of choosing treatment versus being randomised, for each treatment separately. 13 There is, however, no direct estimation of SE and PE. In a reanalysis of the same trial, Long et al. 12 estimate the TE in those who preferred group management, and in those who preferred self-directed treatment. The difference between these quantities is actually equivalent to PE, but Long et al. did not consider the estimation of SE.
3.2 Fully randomised preference design
The Preference Collaborative Review Group 6 carried out a review of fully randomised preference studies, in which baseline preferences for treatment were measured for all participants and accessed for subsequent analysis. A total of 17 trials were identified, and individual level data were available for 11 of them. The analysis examined the differences in outcomes and in attrition rates between three study groups: patients randomly allocated to their preferred treatment; patients with a preference who were randomly assigned to their nonpreferred treatment and patients with no preference. Note that this approach to the analysis involves only one possible comparison between the two groups of patients who had a preference, but the empirical outcomes in the four combinations of preferred and actual treatments are not exploited. However, this review did allow for the possibility of undecided participants, who were analysed as a single group, but again ignoring the actual treatment which had been randomly assigned to them.
In brief, the results of this review showed that outcomes were typically better for patients who were on their preferred treatment. There was less attrition in patients with their nonpreferred treatment than in undecided patients, but there was no significant difference between patients on a preferred treatment and undecided patients. The percentage of patients with a preference was highly variable between studies, with a median of 56%. The relative preference rate for the experimental treatment was also highly variable, ranging from 14% to 100%.
Other examples of this design include trials of low back pain 20 and shoulder pain. 21 Finally, in a trial of weight management strategies, 22 baseline preferences were recorded among three alternative methods of keeping a diary on foods and exercise. Participants were then randomly assigned to one of these methods. The analysis characterised outcomes in those participants whose assigned method was or was not concordant with their preference. As noted earlier, this is equivalent to estimating PE in the context of this design. This last example illustrates the possibility of having more than the usual two treatment options. One should note that attrition from the study was very high, but its authors claim that there was no consequent bias.
3.3 Partially randomised preference design
Examples of this design can be found in various domains, including trials of the addition of X-rays to usual care for the diagnosis and treatment of low back pain; 23 alternative types of hip surgery (THA vs. hip resurfacing arthroplasty); 24 topical versus oral ibuprofen for knee pain; 25 and surgery versus medical treatment for spinal stenosis. 26 Several of these studies allowed participants to choose their treatment if they had first refused to be in the randomised part of the study. In a further variant illustrated in the study by Underwood et al., 25 patients were asked initially if they wished to be in a randomised controlled trial or in a preference study (where they could choose their treatment).
In a related work, King et al. 27 provide a systematic review of 27 trials that used the somewhat similar comprehensive cohort design. In this design, patients who were approached to participate in the randomised trial but refused are nevertheless followed for outcomes; the follow-up is typically achieved through monitoring of an administrative health care database. King’s review excluded studies for which no preference information was available from the individuals who refused randomisation. These studies had been carried out on a wide variety of topics, including pregnancy termination, mental disorders, diabetes, cancer, pain relief, drug abuse and infections. Patient preferences were identified as an important reason to refuse participation in a randomised trial, but there was no evidence of a loss of external validity as a result. There was some indication of PEs on outcomes, particularly in some small trials, but the direction of these effects was inconsistent between studies. There was no evidence of preference effects on participant attrition. The authors concluded that preferences affect whether people participate in studies, but that there was little evidence that they affect study validity. On the other hand, in an earlier review of the comprehensive cohort design, Schmoor et al. 28 concluded that external validity was doubtful in a set of cancer trials.
The trials in the review by King et al. 27 employed several variants of how the study was explained to participants, particularly with respect to the method of treatment assignment. The most frequent methods were: (1) patients who refused randomisation were allowed to choose their treatment; (2) patients were given a choice between being randomised and being able to select their treatment; (3) patients who could not make a choice between the treatments being offered were randomised; (4) all patients except those with a ‘strong’ preference were randomised (5) randomisation preconsent, in which prospective participants were either initially informed (or not) about their randomly assigned treatment, and either were told (or not) about the existence of the choice group. The number of patients who accepted to be randomised ranged from 26% to 88%. These various approaches clearly differ rather substantially in their ability to identify undecided participants, and the consequent interpretation of their estimated preference effects. Possibly because of the diversity of designs used, King et al. 27 calculated what they refer to as treatment-specific preference effects by comparing the choice and random groups for each treatment separately.
A variety of other approaches to the analysis of these designs have been adopted. Some23,24 make separate comparisons of outcomes between participants being randomised to and choosing the experimental treatment, and similarly for the control treatment. Others 25 have carried out separate analyses of the randomised and preference arms of the study, but with only qualitative, nonstatistical comparisons between them.
Weinstein et al. 26 use a preliminary analysis in which the TE sizes were found to be ‘comparable’ in the trial and preference groups, and these groups were therefore combined for the final analysis; in effect, this amounts to the individual data on preferences being ignored in the final analysis. One should note, however, that there were substantial differences in the preference patterns between the random and choice arms.
Gemmell and Dunn 29 carried out a theoretical simulation to evaluate potential bias and inferential error rates in this design, arising because of an unmeasured confounder, and they concluded that the design was not recommended. However, they used relatively extreme simulation scenarios, corresponding to cases where the partially randomised preference design would not be expected to do well. For instance, they assumed there would be 87.5% of participants who would prefer the experimental treatment in the choice arm, which is a relatively extreme value. They also assumed a TE of 2 points on the Beck Depression Inventory, and an unknown confounder effect of same magnitude but in the opposite direction. The percentage of participants with the confounder present was 85% and 45% in the control and experimental groups, respectively, in the choice arm, which all amounts to rather strong confounding. In situations where less severe confounding is in effect, the partially randomised trial design will be more attractive, and some investigators have adopted it, as illustrated by the trials we have reviewed above.
4 Discussion
We have described and compared three main designs that involve the use of preference data, with each being presented in its simplest form. We have outlined a framework for the estimation and interpretation of effects with these designs, fully utilising the available data on preferences, and thus gaining greater insight into the determinants of the study outcome. In this section, we more briefly review some further methodologic issues that require exploration, describe some variants on the three basic designs and mention some other alternative designs.
One area that needs further methodologic development is that of optimisation and sample size determination in preference designs. We have previously examined the optimisation of an important feature of the two-stage design that is under the investigators’ control, that of deciding the relative numbers of patients who should be assigned to the random or choice arms. 5 In brief, we found that the optimum fraction in the choice arm could be above or below 50% in general, but was most often just slightly below 50%; the range 40–55% includes the optimum for most practical cases. The design is more efficient when each of the two treatments being compared is preferred by approximately equal numbers of participants.
Some investigators using the two-stage design have elected to randomise approximately one-third of participants into the choice arm,2,3,9 thinking of the situation as one of the simple randomisation between the three options (A, B or choice). In contrast, in another example,7,8 57% were randomised into the choice arm. These studies will have experienced some loss in precision in estimating preference effects, compared to the more optimal design where 50% are assigned to the choice arm.
The optimum allocation to the choice arm also depends on the relative level of interest of the investigators in the treatment, preference and selection effects. In terms of sample size requirements, it is clear that preference designs will require larger samples if the same level of precision in the estimated TE is required, in comparison to a parallel group design. If the supply of patients is limited, then there is a trade-off between estimating the TE and the selection and preference effects. We are currently working on an approach to this problem in the two-stage design, to be published elsewhere.
Except in the context of the partially randomised design, we assumed that all study participants actually have a preferred treatment. In fact, as originally proposed, 1 the two-stage design can additionally accommodate undecided participants without a preference, by randomising them to treatment, but with the rest of the design unchanged. One then has another randomised comparison that provides an estimate of the TE among undecided participants, but in general that estimate could of course be different from the TE among patients with a preference. Various assumptions are possible (and testable) about the differences between undecided and decided participants, and how their data should fit into the analysis of the entire trial.
Undecided patients may not be uncommon in practice. For instance, in a two-stage randomised trial of medical versus surgical management of heavy menstrual bleeding,7,8 69% of patients were undecided. However, in other cases, the undecided rate may be zero, 9 which is obviously an advantage if the two-stage design is adopted. Another feature that affects the desirability of preference designs is the relative preference rate between the two treatments being compared; as mentioned earlier, having approximately equally preferred treatments is desirable, as happened in the first example just cited.7,8 However, in the other, 9 the relative preference for one treatment over the other was only 34%, which leads to some loss of precision in the estimated selection and preference effects. Further work is needed to consider optimisation, sample size requirements and analytic approach for the fully and partially randomised designs, as well as the two-stage design, when the possibility of undecided participants is taken into account.
A variant design involving preferences was used in a trial of complementary therapies for acute low back pain. 30 One-third of the patients were randomised to usual care, and the remaining two-third were given usual care and additionally allowed to choose one of the several complementary therapies available. The analysis was a standard randomised comparison of outcomes in the usual care versus choice groups. Nothing else seems possible here, because the complementary treatment options available in the choice group were not used in the usual care group. So only the overall effect of being allowed a choice of an additional therapy can be estimated, but the effects of individual therapies are not estimable.
In comparing alternative designs, it has been noted that parallel group trials involve arms that are not comparable psychologically immediately after randomisation, because some patients will be pleased to get their preferred treatment, but others are not; 31 in the partially randomised design, one cannot distinguish the effect of treatment preferences from the confounding of preference with prognosis. This is in accord with one of our general findings on this design, that PE is not estimable, and that getting an unbiased estimate of SE depends on making an untestable assumption that certain contrasts of the PEs are null.
An interesting use of a treatment choice options occurs in a factorial designed trial of radiotherapy and tamoxifen for the treatment of ductal carcinoma in situ, in which one factor was randomised and the other was chosen by ‘patients and clinicians’.32,33 This study is therefore a hybrid between a conventional parallel group and a preference design. However, without some additional randomisation, the effects of preference cannot be uniquely identified for the observationally assigned component of the treatment. Long et al. also discuss certain other more complex ‘hybrid’ designs, with various combinations of who is asked about treatment preferences, and who is randomised or allowed to choose.
Other reviews of the literature on preference designs have been given with an emphasis on psychiatry trials, 34 and in the context of surgery. 35 Janevic et al. 13 have also discussed single and double consent designs, as originally proposed by Zelen,36,37 in comparison to more recently proposed designs, and they describe limitations and examples of each. The single and double consent designs have been used only infrequently, probably because of ethical concerns, but a fairly recent example is a comparison of surveillance versus a choice between surveillance and colposcopy for abnormal cervical smears. 38
Given that preferences usually exist, the question arises about the ethics of randomisation; specifically, it may be unethical to randomise people who hold a strong preference for one particular treatment.39,40 With this in mind, Dumville et al. 41 discuss the advantages of unbalanced (but fixed) randomisation ratios that could reflect greater relative preferences for one of the two treatments, and one can also consider using unbalanced randomisation from the perspective of equipoise. 42 Schultz 43 suggests that subversion of randomisation might occur because of the investigator preferences for one treatment over another. In this context, one should note here that patient and doctor preferences may be different, in which case there will be a conflict between the patient and his medical team about if or how to participate in a trial.40,44 In another area entirely, that of criminology, similar considerations have led to proposals to randomise convicted individuals to alternative sentencing options in various ratios (e.g. 30:70, 50:50 or 70:30), 45 depending on the preferences of the presiding judge; outcomes here could include variables such as the recidivism rate.
It may be difficult to measure preferences reliably, 39 and it has been argued that there may therefore be only a weak relation between stated and actual preferences. 10 Bowling and Rowe 40 suggest that substantial measurement error in preferences may be why they have shown limited effects on outcomes. As we saw earlier for comprehensive cohort designs, there are a variety of ways in which information that informs patients about available treatments has been presented, and this may influence preferences, a topic that has been discussed more fully elsewhere.46–49
We have noted that there are some differences between the designs with respect to the interpretation of their preference data. In particular, recall that the fully randomised preference design elicits a stated preference, whereas in the two-stage and partially randomised preference designs, selected patients are allowed to actually exercise their choice of treatment. It is entirely possible that the treatment chosen by a patient might be different in these two circumstances. Furthermore, it is quite plausible that the fact of being asked about one’s preference but then not being allowed to exercise it could affect the attitude of the patient towards the study, and that in turn might have a bearing on outcomes.
In some studies, there is evidence that treatment choices may evolve as patients move through the study recruitment process. For instance, in a two-stage randomised trial on the treatment of depression, 10 treatment preferences were measured before randomisation, but they did not always correspond to the chosen treatment for people in the choice arm. This is because before randomisation, all patients received written information about all of the various treatment arms. Patients who were randomised to the choice arm were then asked by a physician which treatment they wanted to get, and at this point they had the opportunity to inform themselves about more details of the interventions. In some cases, this new information led to a change of mind, and hence to the discrepancies observed between the initial treatment preference and the actual treatment adopted.
The determination of treatment preferences also raises the question of treatment blinding, which is generally a desirable feature of clinical trials, if possible. Clearly, in the situations where trial participants are allowed to actually choose the treatment they will receive, they become unblinded: this is the case for patients in the choice arm of the two-stage design, and for patients with a definite preference for a certain treatment in the partially randomised design. However, all the other participants in the four trial designs that we have reviewed could potentially be blinded, depending on the circumstances of the particular clinical interventions involved. (In some studies, for instance, in a comparison of medical vs. surgical treatments, blinding the treatment identities would clearly be impossible.) However, even for patients who have chosen their actual treatment, it may still be possible to blind the evaluation of their outcomes, for example, if causes of death are adjudicated by a panel of independent adjudicators who are blinded with respect to the treatment received. Thus, bias may be avoided even if some patients or their clinicians are unblinded.
As a trial progresses, patients may develop a like or dislike of the treatment they are receiving, depending on their response to treatment, or perhaps because of the side effects. This type of effect can only occur after randomisation and initiation of treatment, and they would clearly be related to the ultimate evaluation of treatment benefit or harm. To be clear, note that we have limited our attention to treatment preferences that are identified before randomisation, and which are therefore not biased by effects of the treatment experience at the individual level.
We have also not considered the possibility of covariates that would predict treatment preferences (for instance, perhaps the proportion of patients who select surgical interventions is age related), but this topic would be worthy of further research. Another possibility would be to regard the selection and preference effects as random, rather than fixed effects, as in the models we have explored. In our analysis, we have expressed treatment preferences as binary (for situations with a distinct choice between two treatments), or as a three-category variable (when being undecided about treatments is admitted as a possibility). If a reliable instrument could be developed, preferences could instead be measured on a continuous scale, which would then permit a more detailed evaluation of the impact of preferences, for instance, by examining the relationship of preference strength to study outcomes. However there is also an argument for retaining the categorical representation of preferences, because the acceptance or otherwise of a given treatment is, in practical terms, a binary decision.
We have concentrated on the impact of patient or participant preferences in randomised trials, but their clinicians may also have preferences. Surgeons, for instance, will tend to prefer operative approaches that are more familiar to them, based on their training and experience, as opposed to a more innovative and less familiar techniques. Such clinician preferences may also be associated with differences in study outcomes. Newer treatment methods may thus be disadvantaged because of lower levels of skill and familiarity among clinicians participating in a trial, compared to their greater comfort and competence associated with a more established and standard treatment; this leads to so-called expertise bias. 50 Furthermore, there tend to be more co-interventions and protocol violations (such as crossovers to the alternative treatment arm) in patients assigned to the treatment with which the clinician is less familiar or skilled, 51 and this will also potentially affect study outcomes. More generally, clinicians may prefer to administer one or other treatment for specific patients, because of individual features of their case presentations.
Clinician preferences can be accommodated in the expertise-based design, in which clinicians administer only their treatment of choice, but where patients are still randomised to treatment.50,52 We are otherwise unaware of studies where clinician preferences are incorporated directly into the protocol of the preference designs we have discussed in this paper. However, clinician preferences might still have some influence, for example, in decision aid studies (which are structurally equivalent to the two-stage design); here clinicians and their patients may discuss treatment options, and the clinician’s input may carry considerable weight for patients, especially those without strong opinions of their own. In this context, the selection and preference effects in the analysis of a decision aid design should be interpreted as representing some blend of the clinician and patient beliefs and opinions, rather than those of the patient alone.
In summary, we have seen that the conventional parallel group trial design does not permit investigators to examine the impact of preferences on treatment outcomes, and also that these preferences may be a reason for some people to refuse participation in a randomised trial. In contrast, the alternative designs that we have considered variously permit examination of selection and preference effects. In some instances, these effects may be important, even if the main TE is small; if so, diligent investigators would want to know about them.
The two-stage randomised design seems to be the method of choice in situations where both selection and preference effects are of interest, as well as the usual TE. The fully and partially randomised designs offer some scope in this direction, although neither provides an estimate of both the selection and preference effects with unbiasedness assured. The partially randomised design may be easier to implement in situations where patients who refuse to be randomised can be followed for outcomes without difficulty. The randomised component of this design is then limited to those without a treatment preference, which potentially limits its generalizability; but allowing patients with a preference to receive their desired treatment adds ethical credibility, and those nonrandomised patients also enhance the analysis of the randomised portion of the trial. The fully randomised design suffers from the difficulty of ascertaining preferences of study participants in an unbiased way, while also randomising those same individuals, but not necessarily respecting their preferences in the treatment assignment.
The case for using a preference design can be a strong one, in terms of elucidating important determinants of study outcomes, ethical considerations and practical implementation. Given that treatment preferences have been shown to be a common reason for patients to refuse to be randomised, accommodating those preferences within a rigorous and practical study design seems very desirable. In particular, the two-stage randomised design permits at least some participants to exercise a choice of treatment, and the partially randomised preference design allows all patients with a preference to do so. This feature should make the study more attractive to participating patients; similarly, because there is evidence that physician preferences are sometimes given as a reason not to randomise patients, 27 the potential to recognise preferences will potentially improve recruitment of clinicians and clinical centres to join the group of study investigators.
Footnotes
Funding
This paper was partially funded by NSERC (Natural Sciences and Engineering Research Council, Canada).
