Abstract
We test for the presence of differential item functioning (DIF) in the EQ-5D health-related quality-of-life instrument, using data from a large clinical trial in acute stroke (ISRCTN 99414122). DIF occurs when subjects in different subsets of a sample respond differently to items in a measurement instrument, despite possessing the same latent traits. The data comprised 1462 patient records. We analyzed DIF specifically with respect to responses obtained from different geographical regions and responses obtained from proxies as opposed to the patients themselves. We mapped clinical outcome measures (scores from the modified Rankin Scale, the Barthel Index, and the Zung Depression scale) onto EQ-5D index scores and included dummy variables for proxy responses and for region of treatment (United Kingdom, Asia, rest of world). We predicted the level of problem severity reported on each of the EQ-5D’s five constituent dimensions from the clinical measures and the dummy variables. For given clinical characteristics, proxies were more likely to report health problems than were the patients themselves, although the divergences were not sufficiently large to result in any significant difference in mean index scores between patient and proxy reports. However, the distributions of reported levels of problems for similar clinical states diverged significantly by region, and these translated into different index scores. The mean index score for UK responses was significantly higher than the mean index scores from Asia and the rest of the world.
Keywords
The EQ-5D has become the most widely used health-related quality-of-life instrument for obtaining health state utilities. 1 The questionnaire 2 comprises 2 components, a health state classification and a visual analog scale (VAS). To describe his or her own health circumstances, the subject assigns 1 of 3 levels of health problem (none, moderate, severe, coded 1–3, respectively) to each of 5 health-related domains or dimensions (mobility, self-care, usual activities, pain and discomfort, and anxiety and depression, presented in that order). The subject’s state of health is thereafter described as a 5-digit vector; for example, 11223 implies no problems with mobility and self-care, moderate limitations on usual activities, moderate pain and discomfort, and severe anxiety and depression. Each of the 243 possible states so classified can be associated with a specific social value weight, or an “index score” of health state utility. Value weights have been derived from samples of general populations in various countries. Using the UK value set, 3 the state 11223 translates to an index score of 0.26, relative to utilities of 1 and 0 for perfect health (no problems in any domain) and death, respectively. The VAS allows the subject to assess his or her current health subjectively. The extremes of the scale are “worst imaginable” and “best imaginable” health state, represented as end points of 0 and 100, respectively.
The popularity of the EQ-5D in clinical studies of new health technologies is understandable. The questionnaire is simple to complete and has been endorsed by official technology assessment bodies, such as the UK’s National Institute for Health and Clinical Excellence and, in the United States, the Washington Panel on Cost Effectiveness in Health and Medicine. It has been translated into more than 150 languages, making it usable in almost any country. There currently exist more than a dozen different value sets, that is, country-specific index scores corresponding to the health state descriptions. 4 For a patient with cognitive or dexterity deficits, the questionnaire can be completed by an amanuensis or proxy. An extensive body of studies using the EQ-5D has become established, providing contexts for, and comparisons with, new results. 2 Despite the instrument’s widespread recognition, the testing of its properties continues to be of interest. This article investigates differential item functioning (DIF). DIF is said to occur when subjects in different subsets of a sample respond differently to items within an instrument, despite possessing the same latent trait or characteristic that the instrument is supposedly measuring.
Data from a large multinational trial of the management of acute stroke allowed 2 specific DIF issues to be explored. First, multinational clinical trials are becoming increasingly commonplace because, in comparison with a single-country study, they offer the advantages of accelerated recruitment, cost economies, and increased external validity. 5 The logic behind the multinational trial is that, although clinical outcomes might vary by country, the basis for measuring those outcomes does not. Clinical results are expected to be relevant in all participating countries, not only to those countries in which the particular observations have been made, and this presumption applies equally to EQ-5D patient descriptions. More explicitly, although the EQ-5D approach to international transferability allows any given health state vector to generate different utility values in different countries, it makes the strong assumption that patients in different countries use the same vector to describe the same clinical circumstances. Thus, the health condition that a patient in, say, Spain describes by the vector 11223 would also be described as 11223 by a patient experiencing the identical condition in the United Kingdom or in any other country. If this were found not to be the case, then the pooling of EQ-5D outcomes in multinational trials would become questionable, despite the existence of translations and national value sets.
Second, the debilitating nature of many conditions, including cerebrovascular disease, means that patients might not be able to complete a questionnaire without assistance. The EQ-5D instrument has been designed for either self-completion or proxy completion, yet it is not evident that the results obtained must be commensurable. Experiments involving comparisons of paired patient-and-proxy responses from stroke patients and their carers6,7 have already revealed differences in health state descriptions and in EQ-5D index scores. These findings, however, have yet to be replicated under field conditions, and there are grounds for disputing the validity of the experiments. For example, a patient in a patient-and-proxy pairing who was able to contribute a result was clearly insufficiently incapacitated to require a proxy and, therefore, was not representative of the sort of patient for whom proxy results would normally be collected in practice. Furthermore, with subjects conscious that parallel data were being gathered, discrepancies in judgment might have resulted from gaming between patient and proxy.
Method
Data for this study were derived from the ongoing Efficacy of Nitric Oxide in Stroke (ENOS) trial, investigating the use of glyceryl trinitrate therapy following acute stroke. 8 ENOS has recruited patients in more than 150 hospitals in 16 different countries. At around 90 days following the stroke event, trial personnel record a range of clinical data for each patient. At the same time, the EQ-5D questionnaire is completed either by the patients themselves or by their carers (proxies), depending on the former’s capabilities.
Each patient record included scores for the modified Rankin Scale (mRS) and for the Barthel Index (BI). These instruments have been validated extensively as clinically relevant measures in cerebrovascular disease.9,10 The mRS is a 6-point ordinal classification of increasing disability and dependence, ranging from “no symptoms” (mRS0) through “severe” (mRS5), the latter implying the patient is bedridden, is incontinent, and requires constant attention. In the trial context, a further class (mRS6) is assigned to “dead,” although no mRS6 records were included in the present analysis. The BI assesses patient performance in 10 activities of daily living related to self-care and mobility. Scores vary continuously between 0 and 100, a higher score implying greater independence in physical functioning. All patients in the ENOS trial also completed the Short Zung Depression Scale, which possesses a score range of 25 to 100. A higher score indicates a more severe problem with mood, although Zung scores are known to be positively associated with extent of physical illness also. 11
In a recent study of stroke patients, 12 EQ-5D index scores were mapped from observed mRS scores using regression analysis. We replicated this mapping using the ENOS data, although data availability offered the possibility of a superior model specification. Compared with the Rankin classification, the Barthel Index assesses particular, as opposed to generalized, aspects of physical limitation. BI scores have been used previously as a predictor of EQ-5D index scores in stroke studies, either alone13,14 or alongside Rankin scores. 15 Arguing that both the BI and the mRS neglected psychological aspects of quality of life in stroke patients, a German study 16 added a specific depression measure as a predictor of index scores. In line with these approaches, we added the BI and Zung Depression variables to our mapping model.
In the earlier study, 12 all patients were located in and around Oxford, United Kingdom, and no proxy responses were included. We recognized that the inclusion of the relevant dummy variables in the mapping procedure for the ENOS data could establish the existence, or otherwise, of DIF with respect to region-specific and proxy responses. As the assignment of index scores to health states follows standard EQ-5D algorithms, any differences in mean scores observed as a result of including the dummy variables would be attributable to differences in the distributions of reported dimension levels. In principle, the regression models could be estimated using any of the national value sets of index scores, although we chose to employ the UK set. The ENOS trial is being coordinated from the United Kingdom, and the majority of patients have been recruited by UK sites. The use of the UK set, moreover, facilitates a direct comparison with the previous analysis. 12 The index scores in the UK value set range from 1 (vector 11111, no health problems in any domain) to −0.59 (vector 33333, severe problems in all health domains).
The EQ-5D index score is determined by severity levels on 5 separate dimensions and is, in effect, a summary measure. A particular score is not unique to one configuration of health problems only and might be associated with several different health state vectors. For example, in the UK value set, vectors 13232, 23231, and 32322 all translate to an index score of −0.06 at 2 decimal places, whereas both 12321 and 13212 translate to 0.33. The mapping analysis, therefore, might succeed in signaling the presence of DIF in the instrument, but it would be unable to identify which of the dimensions had constituted the source. To illuminate this question, we constructed logistic regression models, as employed routinely in assessing DIF. 17 These predicted the likelihood of a particular problem severity level being chosen on each of the 5 dimensions, given the latent characteristics as measured by the mRS, BI, and Zung scores, plus the proxy and region dummy variables. All analyses were conducted using SPSS 16.0 (SPSS, Inc., an IBM Company, Chicago, IL).
Ideally, each country contributing records would have been assigned a dummy variable, but this was precluded by the skewed distribution of responses. Most records in our sample (58.3%) originated from the United Kingdom, and the numbers of patients recruited in each of the remaining countries were relatively small. We therefore divided the non-UK data into 2 further regions. The continent of Asia provided 25.1% of the responses, including those from sites in Singapore and Malaysia (10.5% of all responses), India and Sri Lanka (7.0%), and China, including Hong Kong (6.9%). The final region was a residual, the “rest of world.” This provided the remaining 16.5% of responses, with most coming from Poland (5.8%), Romania (4.5%), and Australia/New Zealand (3.0%).
Results
The sample comprised records for 1462 patients. Of these, 24.4% were in the least severe mRS categories (mRS0 and mRS1), 40.7% were slightly or moderately disabled (mRS2 and mRS3), and 34.8% were severely disabled, requiring assistance in attending to bodily needs (mRS4 and mRS5). Patients’ Rankin scores were closely associated with sex and age. Proportionately fewer females than males were assessed as occupying the 2 least severe categories (18.2% v. 28.8%, respectively), and more females than males were classified as being severely disabled (42.7% v. 29.2%, χ2 = 49.6, P < 0.01). Patients classified as mRS5 were, on average, 6.2, 8.4, 11.0, and 13.2 years older than those classified as mRS4, mRS3, mRS2, and mRS1, respectively (one-way analysis of variance, all differences significant at 5%). The mean (SD) BI and Zung Depression scores for the sample were 80.3 (24.9) and 51.0 (16.6).
The mean (SD) EQ-5D index score for the full sample was 0.53 (0.37). A skew in the data was evident, the median being 0.64 (interquartile range [IQR], 0.26−0.81). The mean (SD) VAS score was 65.63 (22.25), the median being 70 (IQR, 50–80). Table 1 displays the distribution of EQ-5D descriptions. These data suggest that, at 3 months, limitations on usual activities were being experienced by around three-quarters of the sample, with around two-fifths reporting no mobility or self-care problems. Severe problems with pain and discomfort and with anxiety and depression were relatively rare. In total, 103 different health state vectors were represented in the data set, including both of the extreme states. Only 11.0% of the sample recorded no health problems (level 1) on all 5 dimensions, whereas 35.8% of records each included at least one level 3 response.
Distribution of EQ-5D Responses, by Dimension (%)
EQ-5D questionnaires were completed by carers or proxies, rather than by the patients themselves, in 436 cases (29.8%). Understandably, the proportion of EQ-5D descriptions recorded by proxy was strongly associated with mRS score (χ2 = 310.9, P < 0.01). Although 13.5% of records for all patients below mRS3 (mRS0 to mRS2) were completed by proxies, 40.1% and 87.1% of mRS4 and mRS5 patient records, respectively, were proxy reported. The mean BI score of patients reporting by proxy was significantly lower than that of patients recording themselves (50.0 v. 83.6, t = 21.9, P < 0.01), and the mean Zung Depression score was significantly higher (56.6 vs 49.7, t = 5.9, P < 0.01).
Model 1 in Table 2 uses ENOS data to replicate the structure of the model constructed by the previous researchers using Oxford data. 12 The modified Rankin Scale is an ordinal classification, and scores are represented as binary dummy variables. As with the Oxford model, mRS1 is the reference category in model 1, and the constant term therefore corresponds to the mean index score for mRS1. The beta coefficient for each of the mRS scores in model 1 is significantly different from zero. Given the coefficients in model 1, the mean EQ-5D index scores for mRS states 0 through 4 amount to 0.93, 0.85, 0.71, 0.55, and 0.28, respectively. In all cases, these lie within 0.03 of the index scores that had been estimated in the earlier Oxford study. The ENOS and Oxford results differ in 2 respects, however. First, the adjusted coefficient of determination for the Oxford model was, at 0.45, considerably lower than the 0.66 reported for model 1. Second, the mean index score for mRS5 using model 1 amounts to −0.15, considerably lower than the −0.06 reported for mRS5 for Oxford. It should be noted that the authors of the Oxford study reported underprediction of the poorest health states, which they considered explicable in part by the small number of records from severely disabled patients in their own sample (7.9% assessed at mRS4 or mRS5 v. 34.8% of ENOS). By contrast, the ENOS predictions were more accurate, in the sense that the 6 mean index scores by mRS category for the ENOS sample corresponded to the mean scores estimated using model 1 to within 2 decimal places.
Linear Regressions Predicting EQ-5D Index and Visual Analog Scale (VAS) Scores
SE, standard error.
Model 2 in Table 2 adds the BI and Zung Depression scores to the model 1 structure. The coefficients for both variables are significant and of the expected sign. The higher coefficient of determination and the continued significance of the mRS coefficients support the view that the new variables augment rather replace the mRS classifications. Model 3 adds the dummy variables for region and proxy recording to model 2. This model is suggestive of DIF with respect to region but not to proxy reporting. Non-UK patients evidently offered EQ-5D descriptions that translated to lower index scores at any given mRS score level. Table 2 also displays a replication of the model 3 formulation using the VAS score, rather than index score, as the dependent variable. The VAS model paralleled the index score model, in terms of signs on coefficients and significance, although the coefficient for the proxy dummy came closer to statistical significance.
The extent to which differences in distributions of descriptors correspond to differences in index scores depends on the algorithms used to construct the value set. In an international trial setting, the choice of value set is essentially arbitrary, although only a limited number of national value sets have been verified to date. By way of experiment, we replicated the Table 2 regression models using value sets from the United States 18 and Japan. 19 These yielded essentially similar results (not shown)—namely, insignificant coefficients for the proxy variable and negative coefficients for the regional dummies (i.e., lower mean scores for the non-UK regions). The adjusted coefficients of determination were marginally higher, however. In fact, these results are unsurprising, as the full sets of the 243 EQ-5D index scores valued for each country are highly correlated (r > 0.95 for each pair of sets). This having been said, the range of national value sets is steadily increasing, and differential responsiveness on the part of some European sets has already been identified. 20
Table 3 presents binary logistic regression models estimated using backward stepwise elimination. These predict whether either moderate or severe health problems (level 2 or 3), as opposed to no problems (level 1), were reported for each of the EQ-5D’s dimensions. The independent variables were those for the Table 2 models, although the mRS4 and mRS5 variables were excluded because no patients with the highest mRS scores gave level 1 responses. The odds ratios (ORs) are easily derived in a logistic regression because the beta coefficient equals the log of the OR. Two indications of model quality are provided. Nagelkerke’s generalized coefficient of determination 21 is a measure of the effect size of the model, 22 and the area under the receiver operating characteristic curve assesses capacity to discriminate. The area under the curve (AUC) is equivalent to the probability that the model will rank a randomly chosen positive instance higher than a randomly chosen negative instance, assuming positive indeed ranks higher than negative. The statistic takes the range 1 (certainty) to 0.5 (i.e., equivalent to random ranking). 23
Stepwise Logistic Regressions Predicting Problems More Than Level 1 on 5 EQ-5D Dimensions
OR, odds ratio; SE, standard error.
Table 3 indicates that higher BI scores decreased the likelihood of problems being reported for the first 4 dimensions, although not for anxiety and depression. Reporting problems was less likely at Rankin scores signaling low disability and dependence, especially mRS0 and mRS1. As with the BI, Rankin scores failed to predict responses on the anxiety and depression dimension. A higher Zung Depression score made reporting problems more probable, for all dimensions except self-care. Proxy responses were more likely to indicate problems than patient responses on the usual activities, pain and discomfort, and anxiety and depression dimensions. Non-UK respondents were more likely than UK patients to indicate problems on the pain and discomfort and anxiety and depression dimensions. Respondents in Asia were also more likely to report mobility problems but were less likely to report problems with self-care and with usual activities. Compared with the first 3 dimensions, the models were less successful in predicting problems for the pain and discomfort and anxiety and depression dimensions. This having been said, problems in these dimensions were less prevalent in the sample than those in the other dimensions (Table 1), and none of the clinical instruments contained an explicit pain measure.
Having discriminated between those reporting problems and those reporting no problems (Table 3), we constructed corresponding binary logistic models to discriminate between those with extreme or severe problems (level 3), as opposed to moderate problems (level 2). Estimation for each dimension entailed excluding those in the sample reporting no problems (level 1), as per Table 1. Again, these models resulted from backward stepwise elimination, and Table 4 presents the results. Only higher-level mRS scores were included as an independent variable, because patients with the lower mRS scores rarely reported problems. Higher BI scores reduced the likelihood that severe problems would be reported, for all dimensions except pain and discomfort, although a high Rankin score was influential in predicting severity on the mobility and usual activities dimensions only. Higher Zung Depression scores predicted severe problems for pain and discomfort and anxiety and depression. Proxy respondents were more likely to indicate moderate than severe problems on the usual activities dimension. When problems in self-care were reported, non-UK subjects were more inclined to rate them as severe. Asian subjects were less inclined to rate pain-discomfort problems as severe. As with those of Table 3, the models were less successful in predicting problems for the pain and discomfort and anxiety and depression dimensions.
Stepwise Logistic Regressions Predicting Severe (Level 3) Problems on 5 EQ-5D Dimensions
OR, odds ratio; SE, standard error.
Summarizing the results in Tables 3 and 4, proxy respondents were more likely than patients to report problems on the usual activities, pain and discomfort, and anxiety and depression dimensions. They were no more likely to report severe problems on the latter 2 dimensions but, with respect to usual activities, were less likely to report the problems as severe. Respondents completing questionnaires in Asian sites were more likely than UK respondents to report problems for mobility and anxiety and depression but less likely for usual activities. In addition, they were less likely to record problems in self-care but, when such problems were reported, were more likely to rate them as severe. The situation with respect to pain and discomfort was the reverse: Asian respondents were more likely than UK respondents to record problems but, when such problems were reported, were less likely rate them as severe. Finally, respondents from the rest of the world were more likely to report problems for pain and discomfort and anxiety and depression only, although, when reported, problems in self-care were more likely to be assessed as severe. These variations in the impact of the dummy variables by level suggest that the DIF identified was nonuniform.
Discussion
Our mapping of EQ-5D index scores on stroke patients’ modified Rankin scores produced accurate predictions and a better fit than had a previous analysis. 12 This fit was improved further, albeit by a modest amount, by including additional clinical measures in the model. Building on this mapping, we were able to test whether the existence of 2 forms of DIF were sufficient to affect either the index scores or the distribution of dimension levels.
With respect to patient/proxy reporting, disparities in judgment in level of health problem were observed for 3 dimensions—namely, limitations on usual activities, pain and discomfort, and anxiety and depression. Proxies were more likely to report health problems than were patients, although less severe ones in the case of usual activities. The divergences were not sufficiently large to result in significant differences in index scores between patient and proxy reports, presumably because, first, moderate or severe problems with pain and discomfort and anxiety and depression were reported the least frequently (Table 1) and, second, the different opinions on severity with respect to limitations on usual activities were, in some degree, compensating. The significance of proxy overstatement on the VAS was marginal (Table 2). Our findings are broadly consistent with those of most other quality-of-life studies of stroke patients, which conclude that proxy respondents overestimate patients’ functional impairment and dependence compared with the estimates of the patients themselves. 24 The timing of the ENOS data collection might well have been important in this context: a proxy-patient pairing study using the EQ-5D 7 noted that differences between responses decreased with time after the stroke event, especially beyond 1 month.
DIF across regions seems more influential than DIF by proxy. Distributions of reported levels of problems for similar clinical states differed by region (Tables 3 and 4), and these translated into significantly different mean index and VAS scores (Table 2). Both were higher for UK responses than for those of responses from Asia and the rest of the world. To the best of our knowledge, regional variation in EQ-5D response has been observed previously in 2 multinational trials only. First, a Rasch analysis was conducted on records from a large (n = 10,205) trial of treatment for schizophrenia, which recruited across 10 countries in Western Europe. 25 Rasch analysis tests for differences in item response patterns, based on the assumption of an underlying uniformity in latent characteristics. The study concluded that national response patterns were congruent across the dimensions with a notable exception: patients were more likely to indicate no problems on the mobility dimension in 9 countries but less likely to do so in the tenth. The authors speculated that the behavior of the outlier might have been accounted for by translation imprecision, “cultural habits,” or small sample size.
The second and more recent study 26 examined data from a large randomized controlled trial (n = 11,118) of treatment for diabetes. Records were available from 20 countries grouped into 3 regions: Asia, Eastern Europe, and established market economies, including the United Kingdom. Logistic regression models were fitted to predict the reporting of health problems on EQ-5D dimensions from clinical histories, risk factors, and sociodemographic characteristics, plus the regional dummy variables. This second study has obvious parallels with the present examination of ENOS data: it employed the same statistical method to detect DIF and constructed broadly similar regional classifications (around two-thirds of ENOS’ “rest-of-world” records originated from Eastern Europe). However, the subjects in the 2 trials differed in their EQ-5D profiles. Only 4% of the diabetes patients recorded a severe health problem in at least 1 domain, and 38% reported no problems on any, compared with 35.8% and 11.8%, respectively, for the ENOS sample. The diabetes study was therefore able to produce models analogous to Table 3 but not to Table 4, while the low prevalence of severe health problems made proxy reporting unnecessary. It was discovered that, allowing for the control variables, diabetes patients in Eastern Europe were significantly more likely to report health problems on all 5 dimensions, relative to patients from established market economies. In contrast, and also relative to market economy respondents, diabetes patients from Asia were less likely to report a problem on the mobility, self-care, and usual activities dimensions but more likely to report one on the anxiety and depression dimension.
In view of the differences in the diseases and countries being considered in these 2 studies and our own, it is unsurprising that, although region-related DIF is evident, no consistent structure emerges. In none of the studies can the explanation for DIF be found in the existing data. However, we note that, whatever that explanation, DIF in quality-of-life instruments used in multinational settings is not confined to the EQ-5D. Interregional variations in quality of life after stroke as measured by the SF-36 instrument have been observed, even following adjustment for prognostic case mix and care quality. 27 In the cancer field, the QLQ-C30 instrument 28 has been translated extensively and used frequently in international studies. An analysis of a large sample of previously collected QLQ-C30 results revealed degrees of DIF by region relative to the response patterns of UK subjects. 29 The variation in response styles was ascribed to sociocultural differences, for example, whether or not the society in which respondents lived embodied an ethic of conformity or of individuality, was tolerant of assertiveness and criticism, or fostered aversion to ambiguity or uncertainty. 30 DIF was most pronounced between responses from the United Kingdom, on one hand, and Eastern Europe and Asia, on the other, a result that bears a strong resemblance to our own findings. Cultural differences might well lie at the root of the explanation for DIF within the EQ-5D also; indeed, they have been offered as an explanation for differences in population health state utilities across the national value sets. 31
We conclude that, according to this study of stroke patients, differential item functioning with respect to patient-v.-proxy responses to the EQ-5D is detectable, although the impact on the mean index scores is insignificant. DIF with respect to multinational responses is detectable, but in this case, the impact on index scores is significant. Whether such a bias constitutes a problem, from the point of view of using trial results to inform decisions, remains debatable. 17 Suppose, for example, that we elected to use our DIF analysis to disinfect the trial results. Table 2 (model 3) indicates that, for the same latent health characteristics, non-UK patients provided EQ-5D descriptions that translate to index scores 0.04 lower than those provided by UK patients. By implication, had identical patients been responding in the United Kingdom rather than elsewhere, their scores would, on average, have been 0.04 higher. By weighting the index scores by the proportions of patients by region, we might calculate a “true UK” mean index score for the entire sample, using both UK utility values and UK EQ-5D distribution of levels. This mean would be 0.55, rather than the present 0.53. Alternatively, we could simply ignore DIF and take the results at face value. We might accept that all measurement entails error margins and that the error introduced by DIF is acceptable. We could question the identification itself, by arguing that health state utility is more than a few latent clinical characteristics, although this argument does weaken the case for the popular practice of mapping utilities onto clinical outcomes. 32 Either way, it would be foolish to make recommendations until further multinational EQ-5D studies of stroke, as well as other medical conditions and circumstances, have been undertaken.
Footnotes
Acknowledgements
We thank all the many patients and investigators who are participating in ENOS (
). ENOS is funded by the UK Medical Research Council and has also received support from the BUPA Foundation, Singapore A*STAR, and The Hypertension Trust. We are grateful for supportive and enlightening comments from the referees.
