Abstract
Preferences for health outcomes, such as quality-of-life or health status, have not been widely studied in children (i.e., those under the age of 18 years). These preferences are usually measured using direct methods, which use stimuli with either certain or uncertain outcomes and ask the respondent to perform either a scaling or choice task. 1 Common direct method approaches include the standard gamble, time trade-off, and visual analogue scale. In general, these methods present respondents with one or more health states—that is, a description of physical, social, and/or emotional functionality— and ask that they rate the health state based on its desirability under the circumstance presented. 2
The use of these direct methods to elicit preferences for health states from children is debatable 3 . For example, Prosser argued that direct methods are too cognitively challenging for children, particularly relatively younger children. 4 Some of these methods ask respondents to make tradeoffs between quality-of-life and length of life. However, if the child-respondent has not experienced a range of symptoms or impairments, they cannot be reasonably expected to hold an informed preference for these health states. 5 Moreover, children’s ongoing cognitive development limit their ability to contemplate living with these health states, particularly death (commonly used in the standard gamble and time trade-off) and, therefore, threaten the validity of using direct methods. 6 Previous studies have observed that even adult respondents are inconsistent when their preferences are elicited through these methods;7,8 it is likely that children provide even less reliable data.
Consequently, proxy preferences are commonly collected from parents, care providers, and adults from the general public. 9 When preferences for health states are collected for the purposes of conducting resource allocation or prioritization (e.g., cost utility analysis), this perspective may actually be preferred. As tax-paying and voting members of society, adult preferences are, arguably, the only ones that should “count”. 10 When preferences for health states are collected for the purposes of informing health care decisions, proxy preferences may also be preferred. Children are often inexperienced in making medical decisions for themselves, with these decisions often involving the one ultimately responsible for caring for the child. 6 The care giver’s preferences should be accounted for in the decision making process.
However, significant differences between proxy and children preferences for health states have been observed.11–14 Though parent proxies are able to fairly accurately report on their children’s physical functioning and symptoms, they are less accurate in reporting their mental health (e.g., cognitive abilities, social and emotional well-being).12,15 Though these observations—and their possible explanations—are heterogeneous, 16 they draw into question the appropriateness of using proxy preferences to guide either clinical care or policy and program development aimed at children.
Discordant preferences for health outcomes drawn from proxies carry risks for both the delivery of care and comparative evaluations. Children’s attitudes towards prescribed interventions are closely linked to their compliance and adherence.10,17 Delivering high-quality care, particularly in preference-sensitive contexts, should include shared decision making,12,18,19 which accounts for the children’s own preferences for outcomes. Likewise, from a policy perspective, these preferences are necessary for evaluating interventions and programs. For example, in comparative effectiveness research, preferences act as a critical input into models comparing alternative interventions 20 and, in cost-utility analyses, they are required for calculating quality-adjusted life years.3,5,9 Given that the Institute of Medicine has identified pediatrics as one of the top ten fields of medicine in need of comparative effectiveness research, 21 these missing preference data represent a considerable gap in our ability to move the field forward.
Thus, given the shortcomings of proxy preferences, there is a motivation for measuring preferences for health states from children. This leaves the question: if we are to go about measuring preferences for health states from children, which direct method should we use? To answer this question, we conducted a systematic review of the literature. Specifically, the purpose of this study was to conduct a systematic review of the validity, reliability, and feasibility of direct preference elicitation methods for health states from children
Methods
Search Methods
Common databases of published, peer-reviewed medical research were searched: Ovid Medline, PsycINFO, Scopus, and EconLit. Medical Subject Headings (MeSH) were used as search terms or keywords. Where appropriate, common acronyms (e.g., “QOL” for quality-of-life) were also used in addition to keywords (see Appendix A–C for the full search). The search was restricted to studies published in English between January 1990 and May 2015. The search was conducted by a trained medical librarian (RS).
Inclusion criteria for this systematic review consisted of any study that elicited preferences from children, whether they were patients or members of the general population. For this study, preferences were defined as individuals’ assessments of the relative desirability v. undesirability of a particular health state. These preferences had to be elicited using a direct method, such as the standard gamble or time tradeoff (full list of methods provided in Appendix A). Studies were excluded if they elicited preferences using indirect health-related quality-of-life measures [e.g., health assessment instruments like EQ-5D(Y)] or if they elicited preferences from adults as proxies.
Abstracts were reviewed for relevance. Specifically, abstracts had to identify that its respective study was based on a sample of children (i.e., individuals under the age of 18 years) and that preferences for health states (i.e., either generic or condition-specific) were elicited. Those studies that clearly did not meet the criteria were excluded from further review, including those that used preferences from indirect methods or one of the methods to measure the current level of symptom severity (e.g., using a visual analogue scale to measure current level of pain). If it was not clear from the abstract that the study met the criteria, it was included for a more detailed review.
Those studies that remained were accessed and read in-full to determine their relevance. The references of these studies were also reviewed and crossed-checked against the list of studies included in the initial search. Studies that were not included in the initial search were retrieved and also reviewed.
Review Methods
Two of the authors (RTC and LMB) split the pool of studies and systematically reviewed the content for the following properties: validity, reliability, or feasibility of using the preference elicitation method. At least one of these properties had to be explicitly reported in the study for it to be included in the review. We based our definitions of these psychometric properties on Froberg and Kane’s 1989 work.
22
Data were systematically collected for the following: 1a) whether validity of the preference elicitation method was reported, 1b) if so, the type of validity reported, and 1c) the statistical outcome of the validity; 2a) whether reliability of the preference elicitation method was reported, 2b) if so, the type of reliability reported, and 2c) the statistical outcome of the reliability; and 3a) whether feasibility of the preference elicitation method was reported, 3b) if so, the type of feasibility reported, and 3c) the statistical outcome of the feasibility.
An assessment of the risk of bias was conducted for both the individual studies included in the review and the aggregated results. For individual studies, three potential sources of bias were considered: selection bias, loss to follow-up bias, interviewer bias. Collectively, these potential biases were weighed when considering the aggregated results and the quality of the evidence. The aggregated results were also considered in light of potential publication bias.
At the end of the initial review, the two authors switched the pool of studies they reviewed in order to assess the extent of agreement. If elements within a study were not clear—such as whether the sample included children, or elicited preferences for health states—the issues were discussed between the reviewers. If the elements remained unclear after the discussion, that study’s authors were contacted via e-mail.
Analysis
In addition to examining the validity, reliability, and feasibility reported by the studies, there were two areas that were specifically examined for comparison. The first area of comparison was between the psychometric properties reported from samples drawn from a pediatric patient population and those drawn from a population of general children. The second area was comparing the psychometric properties reported for older v. younger children. We did not define older or younger a priori, instead let these definitions emerge from the evidence.
Results
The search results are presented in Figure 1. A total of 943 studies were retrieved from the searched databases. Of these, 903 were immediately excluded after a review of their respective abstracts. A review of the references used by those studies that remained identified an additional 34 studies. In total, 74 studies underwent full review. Ultimately, 26 studies were included in the analysis and are summarized in Table 1. Some of the studies used more than one preference elicitation method: the time trade-off was reported in 14 studies, the standard gamble was reported in 11 studies, the visual analogue scale was reported in eight studies, and seven studies reported other methods (e.g., willingness to pay, discrete choice experiment). Several of the excluded studies sampled children as well as adults, but aggregated the results in such a way that made isolating the pediatric sub-sample impossible.

Flow diagram of search and review result.
Characteristics of Studies Included in the Review
Sample size and age ranges have been adjusted to meet our criteria for pediatric sample. Abbreviations: DCE, discrete-choice experiment; NR, not reported; SG, standard gamble; TTO, time trade off; WTP, willingness to pay; VAS, visual analogue scale.
Conventionally, both time tradeoff and standard gamble used “death” as a possible health state. 1 However, the concept of death can be challenging cognitively for children to grasp, 6 and some ethics review committees consider it to be potentially distressing to study participants. 10 Of the studies that used either of these methods, four modified the method to use an outcome other than death as the worst case scenario.10,23–25
Validity
Of the 26 studies included in this review, 7 (27%) reported some form of validity (see Table 2). Construct validity was reported most often, usually operationalized by measuring the correlation of preferences elicited using a direct method to those elicited using an indirect method (i.e., a preference-based, health-related, quality-of-life instrument). Studies that used the standard gamble were more likely to report validity; it was reported in 4 (36%) of the 11 studies that used the method. Comparatively, 2 (14%) out of the 14 studies that used the time tradeoff reported some form of validity, and there were single studies that reported validity for paired comparisons, willingness to pay, and the visual analogue scale.
Reported Validity of Preference Elicitation Methods
Abbreviations: NR, not reported; SG, standard gamble; TTO, time trade off; WTP, willingness to pay; VAS, visual analog scale.
Overall, the direct methods have the strongest construct validity with condition-specific quality-of-life measures. The standard gamble, for example, demonstrates the strongest correlations with the Pediatric Quality of Life Inventory Rheumatology Module (PedsQL-RM). By contrast, the direct methods do not demonstrate strong construct validity with generic measures of quality-of-life. Again, using the standard gamble as an example, its correlations with the Pediatric Quality of Life Inventory Generic Core Module (PedsQL-GC), Health Utilities Index, and Childhood Health Assessment Questionnaire are relatively poor.
None of the studies assessing validity reported their response rates. The majority (n = 5) of studies used either a convenience sample or a mix of convenience and consecutive patients. This kind of sampling strategy can result in self-selection bias. The two remaining studies used more rigorous sampling strategies (Bos and others 17 recruited all children from select Dutch primary schools, and Brunner and others 2003 6 recruited consecutive patients from the Sick Children’s Hospital in Toronto, Canada) and both report poor validity statistics. Five of the studies used in-person interviews, of these, only Brunner and others 2003 6 and 2004 23 reported the number of interviewers that were used (2 and 5, respectively). In both cases, interviewers were trained to elicit preferences using the standard gamble. Rhodes and others 12 reported that the interviews were audio recorded for quality control, but did not mention whether any quality issues were raised.
Reliability
Four (15%) of the 26 studies included in the review reported reliability (see Table 3). These four studies tested reliability using a measure of either 1) between-subject correlation (i.e., Kendall’s coefficient of consistency or intra-class correlation), or 2) test-retest (i.e., Kappa or Spearman’s rank correlation coefficient). Two studies reported the reliability of the time tradeoff, and another two reported the reliability of the visual analogue scale. The standard gamble and paired comparisons were both reported in single studies.
Reported Reliability of Preference Elicitation Methods
Abbreviations: ICC, intra-class correlation; NS, not significant – no other information provided by the author; SG, standard gamble; TTO, time trade off; VAS, visual analog scale.
Juniper and others 26 reported the between-subject reliability for both the standard gamble and the feeling thermometer (a form of visual analogue scale) using a set of cut-offs for age (8 years for the feeling thermometer, 12 years for the standard gamble), school grade (3 for the feeling thermometer, 6 for the standard gamble), and comprehension level (school grade 2 for the feeling thermometer, 6 for the standard gamble). By doing so, they observed an increase in the intra-class correlation for both methods into “acceptable” levels of performance.
For test-retest reliability, a threshold of 0.70 is considered acceptable. 13 Using this threshold, the time tradeoff appears to perform well, coming close to or exceeding the threshold in the two studies that report it. The test-retest reliability of the visual analogue scale did not perform as well. However, of these two studies, only Fox and others 24 reported the number of participants that were retested (10 out of the original 90) and they did not report on what basis these participants were selected for retesting. This could be a source of potential bias.
None of the studies assessing reliability reported their response rates. The sampling strategy used by these studies were of mixed quality. Two studies used either consecutive patients (Bos and others 17 ) or quasi-randomization (Fox and others 24 ). However, Bos and others 17 do not report their reliability statistics and, as previously noted, Fox and others 24 only retested 11% of their original sample. The other two studies recruited using comparatively less rigorous convenience samples from clinics or clinical trials. Three of the four studies used in-person interviews; however, none of these reported how many interviewers were used or if any steps were taken to assure quality or consistency across interviews.
Feasibility
Nine (35%) studies included in the review reported a measure of feasibility, whether explicitly or implicitly by reporting the completion rates of the direct methods used (see Table 4). The majority of these studies administered the direct method as part of an in-person interview.
Reported Feasibility of Preference Elicitation Methods
Abbreviations: DCE, discrete-choice experiment; SG, standard gamble; TTO, time trade off.
The two most common direct methods reported on were the time tradeoff and standard gamble. Cumulatively, 548 out of 584 participants completed the time trade-off (completion rate of = 94%), and 250 of 270 participants completed the standard gamble (completion rate = 93%). Individual study completion rates ranged from 12.5% to 100% for the time tradeoff, and 12.5% to 99% for the standard gamble. For both methods, the lower response rates were observed in studies with smaller samples sizes. The remaining direct methods reported relatively high response rates.
Three of the nine studies assessing feasibility reported their response rates. Moodie and others 25 achieved a 100% response rate; Ratcliffe and others 2011, 10 36%, and Tong and others 2011, 27 84%. These studies varied in terms of their reported completion rates. The majority (n = 7) of studies used a convenience sample as the basis for their respective studies. Bos and others 17 and Moodie and others 25 were the only two studies to recruit all children from local schools, and both of these studies reported high completion rates. Seven of the studies were conducted with interviewers; two of which explicitly reported the number of interviewers that were used and that they were trained to administer the survey (Moodie and others 25 and Yi and others 2009 19 ). While Ratcliffe and others 10 did not explicitly state the number of interviewers, it can be inferred from the methods that two were used for the study. Two studies reported that the interviews were tape recorded and reviewed (Ratcliffe and others 10 and Tong and others 27 ).
Sample Type and Age Comparisons
Recall that there were two areas that we specifically wanted to compare: the nature of the sample used and the age of the participants. In terms of the sample, we compared whether any differences were observed between those studies that drew their sample from a pediatric population (i.e., patients with a specific condition) compared with those drawn from a population of general children. Only two studies made such comparisons: 1) Bos and others, 17 which used both a sample of pediatric orthodontic patients and general children, and 2) Yi and others 2009, 19 which used a sample of pediatric patients with inflammatory bowel disease and general children. Neither study observed any differences in terms of validity or feasibility.
In terms of age, we compared whether any differences were observed between relatively younger and relatively older study participants. There were four studies that specifically made these comparisons.6,17,26,27 Two studies involved the standard gamble and both concluded that, while feasible, the method may not be appropriate for younger children.6,26 Tong and others, 27 in using the time tradeoff method, concluded that it may not be feasible to use with “younger” children. Studies involving the visual analogue scale 26 and paired comparison method 17 reported they could be used with children as young as eight and nine years old, respectively.
Discussion
This systematic review of the literature regarding the use of direct preference elicitation methods with children resulted in 26 published studies. Thirteen studies reported at least one of the psychometric properties of interest—validity, reliability, or feasibility. Seven studies reported some measure of validity, four reported some measure of reliability, and nine reported feasibility. Three studies reported all of these measurement properties.
The standard gamble and time tradeoff were the elicitation methods that were most frequently reported, which is understandable given their reputation as the “gold standard” for preference elicitation methods. 1 The standard gamble demonstrated moderate construct validity in terms of its correlation with condition-specific measures of health status. Much lower construct validity was observed with measure of generic health status. It should be highlighted, however, that these were reported in relatively small studies, with all but one with a sample size over 50 participants. A single—again, relatively small—study reported the reliability of the standard gamble; it was very dependent on age and education level, with younger and less educated children not performing as well as their older peers. Despite the reported challenges associated with the standard gamble, it appears to be feasible, with completion rates greater than 75% in all but a single study.
These results should be considered in light of the risk of bias. Three of the six studies that reported on the standard gamble used convenience samples, and an additional two used a mix of convenience and consecutive sampling. This opens the chance of self-selection bias, where only those children who are engaged or confident enough in their abilities are participating in studies of this nature. Four studies explicitly reported that steps were taken to address the potential for interviewer bias but two of these studies were relatively small (<50 participants).
In terms of the time tradeoff, there were only two studies that reported its validity: one compared it to a condition-specific measure of health status, the other to a generic measure of health status (i.e., the Health Utilities Index). Much like the standard gamble, the time tradeoff appears to have had more construct validity with the former than it did with the latter. However, unlike the standard gamble, these two studies had much larger sample sizes, each involving over 100 participants. The time tradeoff demonstrated acceptable levels of test-retest reliability in two studies. However, only one study reported how many participants were re-tested, and this was a minority of the original sample, indicating that the chance for loss to follow-up bias is high. The completion rates for the time trade-off were mixed; although, the studies reporting completion rates involved much larger sample sizes. Five studies employed sampling strategies with a high probability of self-selection bias, using convenience samples from clinics or local schools or soliciting interested participants in the local media.
Summarizing the two most commonly used direct preference elicitation methods, there is some preliminary evidence that these methods appear to be valid, reliable, and feasible. For the standard gamble, those that have reported these properties have been mostly small studies, with only two having samples over 50 participants. Studies reporting the time trade-off, on the other hand, have had much larger samples. This should be considered when weighing the results reported in this review. While both methods appear to have some validity when eliciting preference regarding condition-specific health states, reliability when used with older children, and reasonable completion rates, there are simply too few studies upon which a definitive conclusion can be drawn. Neither method stands out as being clearly superior to the other in terms of its ability to elicit preferences for health states from children.
Of the other direct preference elicitation methods that have been reported—discrete choice experiment, paired comparison, visual analogue scale, and willingness to pay—there is insufficient evidence to make any firm conclusions. Either the methods were only reported in a single study, or a single psychometric property was reported for the method.
In general, there is a dearth of evidence with respect to the use of direct preference elicitation methods with children, particularly when compared to the adult population; a recent systematic review identified 344 studies using these methods in adults. 28 One reason for this may be the methodological challenges associated with studying children. Prosser and colleagues 9 have previously described the numerous underlying issues that researchers need to overcome when eliciting preferences from children using direct methods. These include developing health states that are relevant to children living as a dependent on family members and caregivers, accounting for numeracy levels and understanding of risk, and articulating time horizons in a way that can be comprehended by children. Indeed, our review observed that several direct methods performed better with relatively older children.
The results of this study are limited by several of its methodological approaches. First, the search was limited to four databases and studies published after 1990. However, these databases represent a significant portion of studies published in peer-reviewed journals in medicine, psychology, and the social sciences, and it is unlikely that expanding the search to other databases would yield very different results from those presented in this study. We scanned the references of the 26 studies included in this review, none of which identified any additional studies published prior to the 1990 cut-off. A second limitation is that the search was restricted to English studies using MeSH terms or commonly used keywords. Some of the studies that may have been missed because of these restrictions were identified by reviewing the references of the identified studies. Third, there is a chance of publication bias; i.e., studies using methods that demonstrated poor validity, reliability, or feasibility, post hoc, may not have been published, or not reported.
This systematic review demonstrates shortcomings in the published literature regarding the use of direct preference elicitation methods with children in health care. It is clear that there is considerable work needed to establish the fundamental properties of using these methods with children. This work is critical if value in pediatric medicine is going to be defined and operationalized through comparative effectiveness research, 29 or if advances in the delivery of patient-centered pediatric medicine are to be made.
In conclusion, relatively few studies have published evidence regarding the validity, reliability, and feasibility of using direct preference elicitation methods to value health states from children. The standard gamble and time tradeoff were the methods that had the most evidence reported. Other methods had less evidence supporting their validity, reliability, or feasibility. Further work is needed in this area to establish whether such methods can be applied appropriately to this population.
Footnotes
Acknowledgements
We acknowledge the careful review of draft manuscripts of this study by Drs. Julie Panepinto and Gillian Currie.
This work was conducted at the Department of Pediatrics, Medical College of Wisconsin. Financial support for this study was provided entirely by a grant from the Advancing a Healthier Wisconsin Endowment at the Medical College of Wisconsin. The funding agreement ensured the authors’ independence in designing the study, interpreting the data, writing, and publishing the report.
