Abstract
Objective
To evaluate the reproducibility, longitudinal validity, and interpretability of the disease burden morbidity assessment in people with chronic conditions including multimorbidity.
Methods
The study was conducted using a longitudinal cohort design. A large consecutive sample of adult patients at an Australian community-based rehabilitation service was included with testing at baseline and three-month follow-up (testing longitudinal validity and interpretability). A smaller subsample of patients completed a one-week test–retest (testing reproducibility). Outcome measures included the Disease Burden Morbidity Assessment and 36-item Short-Form Health Survey. Participants in the study received tailored, interdisciplinary intervention between baseline, and three-month follow-up but did not typically receive intervention between baseline and retest.
Results
The longitudinal validity and interpretability sample included 351 participants and the reproducibility sample included 56 participants. Longitudinal validity and interpretability were generally supported with hypotheses supported or partly supported and a small percentage of lowest total scores for impact on daily activities (0.6% at baseline, 1.3% at three-month follow-up). Reproducibility parameters were acceptable for the total score measuring impact on daily activities (e.g. ICC = 0.76).
Discussion
Reproducibility, longitudinal validity, and interpretability of the disease burden morbidity assessment were generally supported for community-based chronic disease patients.
Introduction
The patient-reported Disease Burden Morbidity Assessment (DBMA), 1 by Bayliss et al., was designed to measure the disease burden of chronic conditions by rating the impact of chronic conditions on daily activities. This is important as patients have identified daily activities as an essential domain of quality of life, 2 that may influence a patient’s capacity to participate in healthcare interventions that ultimately reduce the impact of their illness.3,4 Patient-reported comorbidity questionnaires are appropriate for population3 and community-based research5, and are the only method available to obtain information on disease burden over time. 6 Further, the potential confounding effects of disease burden may need to be adjusted when an index condition is of interest 1 or weighted to account for the severity of the impact of comorbidity.7–9
Testing of comorbidity measures has predominantly focused on predicting mortality or HRQOL10,11 or for diagnosis. 11 Monitoring the impact of disease burden is also important in healthcare settings as this information may influence the intervention provided and has the potential to reduce further comorbidity. Central to psychometric testing of the DBMA for an evaluative purpose is the conceptualisation of disease burden and multimorbidity (both complex constructs which can be measured using the DBMA). The design of the DBMA is consistent with multimorbidity defined as the coexistance of two or more conditions within a person without reference to an index condition12,13 as no index condition needs to be listed. Measurement in settings where care focuses on the patient as a whole rather than prioritising one condition are also consistent with this conceptualisation. 12 Of importance in testing for an evaluative purpose is considering prior research that examines temporal patterns of disease burden and multimorbidity. This includes work that indicates congestive heart failure, diabetes, and chronic respiratory conditions predict clinically significant decline in SF-36 physical components scores. 14 As few studies of temporal patterns exist, other work may be useful to inform this testing. For example, comorbidity measures that incorporate disease burden have been found to be stronger predictors of health-related quality of life than measures that incorporate a simple count of comorbidities.1,3,15–17
Testing of the psychometric properties of the DBMA has primarily been conducted by the original authors using people receiving primary, specialist or hospital care 1 and by independent researchers using a French version of the measure in people attending a general practice clinic. 18 The DBMA has previously been found to correlate more strongly with overall health status, physical functioning, depression and self-efficacy than the Charlson Index, the Risk Adjustment (RxRisk) score or counts of individual conditions, 1 supporting validity of the measure. However, no known study exists comparing the DBMA with other patient-report comorbidity questionnaires or determining the quality of the original measure when used for an evaluative purpose (measuring change over time). Test–retest reliability of the total score of the French version of the measure – the DBMA-Fv has been found to be acceptable (Intraclass Correlation Coefficient (ICC), 0.86; Confidence Interval (CI), 0.79–0.92) in patients aged 18 years or older attending a general practice clinic. 18 The aim of this study was to evaluate the reproducibility, longitudinal validity (responsiveness), and interpretability of the DBMA in adults with chronic disease (predominantly with multimorbidity).
Methods
Study design and participants
A prospective consecutively sampled cohort of people with chronic conditions was used with data collected as part of a larger longitudinal study of people attending a community-based, subacute, interdisciplinary rehabilitation clinic in Australia. 5 The study setting and characteristics of participants are described in Table 1. Participants with baseline (first attendance at the clinic) and three-month follow-up data from the larger study were used to investigate longitudinal validity and interpretability (Figure 1). Longitudinal validity was defined as correlation between changes in the DBMA and related and non-related constructs (i.e. the ability to detect real change in the construct being measured) rather than treatment effect. 19 To investigate the reproducibility of the DBMA, a smaller consecutively sampled subsample of the study cohort were used. This subsample completed the DBMA at baseline and again one week later with no clinical interventions typically received between these assessments at the interdisciplinary rehabilitation clinic. However, these participants, as well as the remainder of the study cohort typically received tailored, interdisciplinary intervention between the baseline assessment and three-month follow-up. This intervention included individual sessions with allied health professionals as well as programs targeting presenting conditions (e.g. gym-based exercise, back care). A sample size estimate of a minimum of 50 participants was required to detect an ICC of .80 with a 95% confidence interval (CI) from .70 to .90 20 and has been reported as an acceptable sample size for testing reproducibility, responsiveness, and construct validity. 21 Ethical approval was granted by the Central Queensland Human Research Ethics Committee (HREC 11/QCQ/14). Written informed consent to participate was obtained from each participant. The study design and reporting follow the COnsensus-based Standards for the selection of health status Measurement INstruments (COSMIN) checklist. 22
Participant inclusion and exclusion criteria and description of the study setting and processes.

Flowchart of the participants included in the study.
Procedure
The reproducibility subsample self-completed the DBMA (written in English) and SF-36 at baseline in the clinic environment and at home one week later, in paper form, between April and July 2012. The DBMA did not explicitly state whether a diagnosis was required for a chronic disease to be reported; thus, information reported was based on the judgement of the participant. The longitudinal validity and interpretability sample self-completed the same measures in the clic environment at baseline and at three-month follow-up either in the clinic environment or at home for participants who were going to be unable to attend the clinic at the three-month follow-up.
Outcome measures
The DBMA lists 21 physical conditions (Table 2, first column) and also has an open-ended section where other conditions can be added. Each condition is rated as present or absent. If present, the condition is rated on a 5-point Likert scale (1 = limits daily activities not at all to 5 = limits daily activities a lot), with higher scores indicating greater disease burden. The revised 21-item English version16 of the original 23-item questionnaire 1 was used for this study. The total score for impact on daily activities is the sum of recorded scores for each condition.1,16 The DBMA has no theoretical ceiling as participants can list an unlimited number of ‘other’ conditions in the open-ended section. 1 Non-responses (i.e. no score for the absence of the listed condition and impact on daily activities) were coded as missing values. A single assessor, who was a research assistant for the study with a clinical background in clinical measurements for acute and chronic diseases, administered the questionnaires for the reproducibility component and administered most of the questionnaires for the longitudinal validity and interpretability component, although another two assessors were involved.
Reproducibility of the Disease Burden Morbidity Assessment (DBMA) using a one week retest.
Note: Dashes indicate where there were too few cases for a specific disease for values to be calculated.
No.: number; obs.: observations; SD: standard deviation; SEM: standard error of measurement; SDC: smallest detectable change; N/A: not applicable; BP: blood pressure; RA: rheumatoid arthritis; Circ: circulation; CH Failure: congestive heart failure; IQR: interquartile range; ICC: Intraclass correlation coefficient; ANOVA: analysis of variance.
an = 62.
bn= 56.
The eight dimensions of the SF-36 were used as an external criterion of health-related quality of life for testing longitudinal validity as per previous studies.1,17 The SF-36 dimensions were scored on a 0 to 100% scale with 0 indicating the worst function and 100 the best function. The item on self-reported health perception was used as a criterion for determining health stability between baseline and retest as this measure has been found to be a simple, integrative patient-centred assessment for the evaluation of illness in the context of multimorbidity over and above psychosocial measures. 23 The SF-36 has been shown to have acceptable reliability and construct validity in Australian chronic disease populations.24,25
Statistical analysis
Descriptive statistics were used to describe the study sample. Data were checked regarding whether the assumption of missing completely at random was met using Little’s Completely Missing at Random (CMAR) test. Reproducibility includes both reliability and agreement. 26 Reliability of the total score for impact on daily activities and number of comorbidities was assessed using ICCs, using a two-way, random effects model, and associated CIs. An ICC of greater than 0.7 was considered acceptable for research purposes. 21 Visual inspection of a Bland–Altman plot and linear regression was conducted to determine whether proportional bias was contributing to the ICCs. Agreement was assessed using: percentages of agreement within one and two points between test and re-test on the DBMA; and the standard error of measurement (SEM) and smallest detectable change (SDC) for items with normally distributed residuals. The SEM was calculated as √σ2 (where σ2 is the mean square error term from the ICC ANOVA).21,26 The SDC was calculated as 1.96 × √2 × SEM. Interpretability was determined by the percentage of participants who had the lowest and highest total score and the distribution of scores22 at baseline and three-month follow-up, as participants who already have the highest or lowest scores at any point cannot change further in the respective direction impacting on responsiveness, and if large numbers of participants have the same scores they cannot be differentiated from each other influencing reliability. 27 Stability as a necessary prerequisite for examining test–retest reliability 27 was examined using median or mean individual item and total impact on daily activities scores of the DBMA and the self-reported health perception item of the SF-36. The distribution of scores, mean, and standard deviation of the DBMA sample total scores, median (IQR) of the DBMA individual items, and missing values were also calculated as indicators of interpretability. 26
For consistency with previous recommendations, missing SF-36 data were imputed using the mean values of the remaining dimension items28 when less than 50% of the items were missing. When 50% or more of items were missing, data were removed from analyses. Convergent and divergent construct longitudinal validity was determined by examining hypothesised correlations between changes in the DBMA and changes in the SF-36 dimensions using Pearson's product-moment correlation coefficients (or Spearman's correlation coefficients for gender subgroup analyses), determined a priori (Table 3). Internal consistency of the total number of comorbidities and the total impact on daily activities was examined using Chronbach’s alpha. Statistical analyses were performed using IBM SPSS Statistics for Windows, Version 23.0. (Armonk, New York: IBM Corporation).
Hypotheses regarding longitudinal validity and interpretability of the DBMA using correlations with the SF-36 dimensions. a
aHypotheses for longitudinal validity were based on expected correlations between changes in SF-36 dimensions with changes in the DMBA total scores and individual items. Hypothesis testing was based on studies of the DBMA wherever possible, as different measures of multimorbidity have beeen found to result in varying levels of physical health-related quality of life. 30 As few studies of temporal patterns exist for people with multimorbidity expected changes were based on cross-sectional as well as longitudinal studies and expert opinion.
Results
A summary of the participant flow for the study is presented in Figure 1. The assumption of data missing completely at random for the number of comorbidities and severity of comorbidities was supported for the DBMA (Little’s MCAR test = 0.07 to 0.95). Internal consistency of the number of comorbidities score was 0.60 and of the total impact on daily activities score was 0.69.
Reproducibility component
Of the participants who consented to participation (n = 62 of 78 approached), 90% participated. No data were missing for participants who completed the DBMA at baseline or one-week retest. Participants who completed the initial test and a one-week retest (n = 56) had a mean (SD) age of 61 (12) years, and the majority (n = 34, 61%) were female. The mean (SD) number of comorbidities was 7.5 (2.9), and all participants had two or more comorbidities.
For the total score of the impact of all conditions on daily activities, test–retest reliability was acceptable (ICC = 0.77) (Table 2). The SDC varied from 1.69 to 3.80 on individual items. The reliability of the total number of chronic conditions (Cohen’s kappa = 0.27) indicated this index was not as stable as the total impact on daily activities over the test–retest period (Table 2). A Bland–Altman plot of the total number of chronic conditions and accompanying regression analysis (Figure 2) indicated proportional bias did not appear to be present. The majority of individual items and the total impact on daily activities median and mean scores were relatively stable (i.e. no change for 15 out of 19 relevant individual items, mean (SD) total impact on daily activities score 18.02 (10.49) at baseline and 18.11 (12.32) at retest). Self-reported health perception was relatively stable between baseline and retest (mean (SD) = 3.61 (0.70), 3.55 (0.79), respectively) based on the patients who completed the health perception item of the SF-36 at both timepoints.

Bland–Altman Plot of test–retest reliability for total number of comorbidities.
Longitudinal validity and interpretability component
Participants (n = 307/351, 87%) who completed the DBMA at baseline and at three-month follow-up (Figure 1) had a mean (SD) age of 59 (14) years and the majority (n = 178, 58%) were female. DBMA missing data at baseline for 10 participants are reported in Table 4. No DBMA data were missing for participants who completed the DBMA at three-month follow-up. The mean (SD) number of comorbidities at baseline was 7.0 (3.0), and n = 346 (99%) had two or more comorbidities. The most common condition at baseline was being overweight (76%). Ratings of the impact on daily activities (Table 4) were skewed with the impact of individual conditions in the lower range for the majority of participants, with median ratings across the 21 listed conditions of 3 or less (on the 5-point scale) for all but the rheumatology item (median 4 at baseline).
Disease burden morbidity assessment (DBMA) medians and interquartile ranges at baseline and three-month follow-up (longitudinal validity study component).
IQR: interquartile range.
aConditions listed by participants in the “other” section of the DBMA but not included in the table were depression, post-polio syndrome, epilepsy, anxiety, schizophrenia, neck pain, vertigo, obstructive sleep apnoea, hemochromatosis, non-alcoholic fatty liver disease, ross river fever, spinal fusion, diverticulitis, Reynaud’s syndrome, benign intercranial hypertension, glaucoma, gout, tinnitus, carpal tunnel, and kidney disease. A further rare condition was documented by a participant but has not been specified to prevent any risk of the patient being identified.
bDBMA missing data at baseline included back data for four participants, thyroid and back problems for one participant, rheumatoid arthritis for one participant, vision problem and chronic bronchitis or emphysema for one participant, osteoporosis for one participant, vision and hearing problems for one participant, diabetes for two participant, asthma for one participant.
The strength and direction of correlations between changes in SF-36 dimensions with DBMA total scores and individual items are displayed in Table 5. The longitudinal validity hypotheses were supported or partly supported (Table 3), although some unexpected higher strength correlations were observed for conditions in three of the six hypotheses. With respect to interpretability, a low percentage of participants had the lowest possible total number of comorbidities at baseline and three-month follow-up (1.4 and 3.9%, respectively) and for the total score of impact on daily activities at baseline and three-month follow-up (0.6 and 1.3%, respectively) indicating floor effects were unlikely to have impacted on estimates. Analyses pertaining to the hypothesis regarding the subgroups of males versus females were likely adequately powered as data for greater than 50 participants were available for each subgroup 19 (n = 178 females, n = 129 males).
Correlations between changes from baseline to three months on the disease burden morbidity assessment (DBMA) and changes in the 36-item Short Form Health Survey SF-36 physical and mental functioning dimensions (longitudinal validity study component).
Note: 36-item Short Form Health Survey: SF-36.aCorrelations have only been calculated for items with greater than n = 20.
bNo SF-36 data missing at baseline for participants who were included at baseline and three-month follow-up, one participant had all SF-36 data missing at three-month follow-up, three participants had data for two dimensions missing at three-month follow-up, and one participants had data for one dimension missing at three-month follow-up.
*p < 0.05, **p < 0.01.
Discussion
This study indicated that the DBMA had acceptable psychometric properties for evaluating disease burden in people aged 18 years or older with chronic disease and multimorbidity in several areas (e.g. longitudinal validity of individual items and test–retest reliability of the total score of the impact on daily living). The sample used to test the DBMA was unique in that individuals with a large number of conditions were included, which has not been the focus of previous work. Interpretability was supported based on the small percentage of participants with the lowest scores.
Variability in the impact of some diseases on daily activities was likely responsible for the low reliability coefficients for these items (i.e. diabetes, stomach problem, vision problem, osteoarthritis) rather than measurement error. Under-reporting of stigmatizing chronic diseases and variability in patient-reports of conditions with highly subjective symptoms such as pain in arthritis and headache 6 may have impacted the findings in this study. It was also plausible that variability of the impact of individual conditions such as overweight and hard of hearing was due to variability in the impact of other comorbid conditions that was difficult to isolate or that ratings across multiple items were influenced by dominant conditions at that time. For example, separating out the impact of pain underlying osteoarthritis from the impact of pain underlying cancer might be difficult for patients with both conditions. This reasoning is potentially supported by the unexpected relatively high correlations between changes in the SF-36 bodily pain dimension and the impact of conditions that have not typically been conceptualised as pain-centric conditions (colon problem, blood circulation problem, overweight, hard of hearing). Short-term fluctuations in disease burden attributable to individual conditions pose an important challenge to those seeking to quantify disease burden among people with chronic conditions; however, findings from this study indicated that the DBMA total score of the impact on daily living was acceptable for picking up changes in the study sample.
Interestingly, the reliability of using a total count of conditions using the DBMA was not supported in people with a large mean number of chronic conditions in this study. Counts of the number of chronic conditions have been important when quantifying multimorbidity in prior studies. 4 Acceptable agreement of more than 90% for all but one item has been reported for a patient-report comorbidity measure based on the Charlson index in an inpatient sample, although only 3 out of 18 individual items had test–retest reliability coefficients above 0.70.33 Variability in the total number of conditions reported by chronic disease patients over the one-week period in this study may be related to inaccurate memory of conditions (recall bias),34 and variability in the setting or in reporting due to short-term fluctuations in symptoms (e.g. arthritis may have been reported on days where stiffness was felt by patients even if there was no formal diagnosis of arthritis35). The reliability results using the total score of impact on daily activities were consistent with results of testing using the DBMA-Fr, albeit that our test–retest reliability results were marginally lower than in that study (ICC = 0.77 vs. 0.8618).
Limitations and future directions
Support exists regarding the validity of summing scores of the impact of single conditions on daily activities in a population sample of adults 50 years or older with four or fewer comorbidities.36 Further, new work that applied multi-trait multi-method analysis has supported the validity of using a summed disease burden score for a disease burden measure with multimorbid conditions using a list of 35 conditions.37 However, further work is recommended to confirm the use of DBMA total scores. Although the impact on daily activities may be apportioned across conditions in cases where people find it difficult to attribute symptoms to individual diseases,38 the potential difficulty that patients face in judging how much impact they apportion to each condition may be considered a limitation of the DBMA. Similarly, patients may inadvertently attribute the impact of multiple chronic conditions to a single condition which may complicate disease burden estimates.35 However, it is noteworthy that in both of these scenarios, the validity of the total score of the impact of individual conditions would be supported.
The focus of the DBMA was on physical conditions. Mental health conditions could be reported in the ‘other’ open-ended section of the measure but should be measured separately1,18 for those seeking to capture the impact of all comorbidities as the prevalence of these conditions was likely under-reported. Additionally, the relatively small sample size for testing reproducibility for some individual items may have meant those item findings were influenced by unusual reporting behaviour from a small number of participants, thus further work is required to confirm (or refute) some of this study’s findings. Further content validation is required that involves target groups and examines temporal trends. A further consideration is that the DBMA may perform differently in people with few comorbidities as mean number of comorbidities in the study sample was seven, thus further testing should be conducted with this group.
Footnotes
Acknowledgements
The authors gratefully acknowledge the contribution of Lynette Mackey de Paiva and Christine Woods who assisted in the collection and collation of data and the clinic staff who assisted with coordination of the study.
Declaration of conflicting interests
The author(s) declared the following potential conflicts of interest with respect to the research, authorship, and/or publication of this article: At the commencement of the study, Kerrie-anne Frakes was the manager of the clinic at which the study was conducted and Zephanie Tyack was a research fellow in the health service where the clinic was located.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This study was funded by the Queensland Government, Health Practitioner Research Scheme, and the Central Queensland Hospital and Health Service. The funding bodies did not have any input into the study design, conduct or reporting. SMM is supported by a National Health and Medical Research Council (of Australia) fellowship (#1090440).
