Abstract
Although intervention dose—defined as the quality and quantity of an intervention and participation—might be key to understanding why some multisite quality improvement (QI) initiatives work and others do not, evaluations rarely consider dose, and there is no widely accepted method for measuring it. In this exploratory study, the authors examined the literature on QI dose, identified four methods for measuring QI dose, applied them to 14 communities participating in a QI initiative, examined whether the dose scores aligned with perceptions of QI dose among individuals knowledgeable of the initiative, and report on lessons learned. They conclude it is feasible to measure QI dose and found a high level of concordance between scores on a comprehensive dose measure and knowledgeable informants’ perceptions. However, measuring QI dose presents many challenges, including subjective decisions about the elements of dose to include in a measure and the need for extensive data collection.
Introduction
In response to high-profile reports showing deficiencies in the delivery of health care in the United States (Agency for Healthcare Research and Quality, 2011; Mangione-Smith et al., 2007; McGlynn et al., 2003), there have been a number of large, multisite initiatives in the public and private sectors to improve care quality. Examples include Medicare’s coordinated care demonstration and the Institute for Healthcare Improvement’s 100,000 Lives Campaign. Many of these quality improvement (QI) efforts include evaluations to determine whether the interventions were successful at improving health care or health outcomes.
However, evaluations of QI interventions often ignore the influence of intervention dose in their assessments (Legrand et al., 2012; Reed et al., 2007). We define intervention dose as the quality and quantity of an intervention and participation. Although the term “dose” more commonly refers only to the quantity of an intervention, our definition acknowledges that other characteristics of the intervention may also influence its effectiveness (O’Neill, 2011).
The absence of a consideration of dose in evaluations of QI interventions is an important oversight since the implementation of a QI intervention is likely to vary across intervention sites with respect to the quality and quantity of both the intervention and participation. Understanding the dose of a QI intervention may be central to explaining why an intervention succeeded at improving outcomes in some sites and not others. A better understanding of dose may facilitate comparisons of different types of QI interventions (e.g., practice coaching vs. learning collaboratives) and explorations about the dose needed to see a significant change in outcomes (Reed et al., 2007). This information, in turn, may provide more instructive information to organizations seeking to implement QI interventions (Jackson & Williams, 2015).
A major challenge to undertaking this work is that there is currently no widely accepted method of measuring dose (Huber, Hall, & Vaughn, 2001), and prior research on the subject offers little guidance. Evaluations that consider dose are often focused on a single intervention at the practice or hospital level. For these studies, dose is often measured by simply counting the number of exposures to the intervention (Charles, Pitt, Halliday, & Amor, 2013). For multisite program evaluations, measuring dose becomes more complicated as improvement work across sites may involve numerous interventions, sometimes taking place across multiple departments or organizations. Frequently, the intervention dose is crudely measured as a single binary variable indicating if the site participated in the program or not, assuming dose is constant across all intervention sites (Brock et al., 2013; Schonlau et al., 2005).
The purpose of this effort was to advance the field’s understanding of QI dose for multisite program evaluations. First, we explored the concept of dose in the literature. Second, we identified four approaches for measuring QI dose and applied them to interventions adopted under the Aligning Forces for Quality (AF4Q) program, an initiative of the Robert Wood Johnson Foundation that offered financial support and technical assistance to multistakeholder alliances to improve health care quality within their communities (Painter & Lavizzo-Mourey, 2008). Third, we examined the extent to which scores from our measurement approaches aligned with opinions of an informed group of respondents involved with managing the operations of the AF4Q program. Fourth, we report on the feasibility and challenges of measuring dose in an effort to provide lessons learned for evaluators, designers, and funders of multisite QI initiatives.
New Contribution
Our study contributes to the literature on measuring the dose of multisite QI interventions and is unique in two respects. First, to our knowledge, it is among the first articles to investigate the dose of a broad, multisite QI initiative in which intervention sites were given considerable latitude in terms of how to meet program expectations. Second, our study considers multiple ways to measure QI dose and demonstrates how one constructs and applies measures of dose that can have important implications for research findings. Based on our work, we offer suggestions for advancing the field in measuring QI dose. Ultimately, we seek to bring the field closer to the development of a robust measure of dose, which would provide practical, actionable information for evaluators of QI interventions and organizations seeking to adopt QI interventions.
The Aligning Forces for Quality Program
The purpose of the AF4Q program was to “reweave the fabric of their health care systems to be stronger, more resilient, and of higher quality across the full continuum of care” (Painter & Lavizzo-Mourey, 2008, p. 1461). Under the program, funding was directed to multistakeholder alliances, consisting of payers, providers, consumers, and purchasers, that facilitated improvement by securing and coordinating resources, promoting collaboration across providers, disseminating information, and prioritizing common goals and initiatives (Harvey, Beich, Alexander, & Scanlon, 2012). The foundation provided funding and technical assistance to the alliances, and the alliances were expected to meet specified goals and objectives in several programmatic areas including QI. The initiative began in 2006 with four alliances that were invited to participate based on their history of community collaboration. By 2007, the foundation added 10 additional alliances through a competitive application process. (Two more communities were added in 2010, but are not included in this analysis due to their limited time in the program.) The alliances represented a diverse set of communities across the nation, including whole states, and metropolitan and rural areas (e.g., Minnesota, Cleveland, Humboldt County, California). A more thorough description of the AF4Q program can be found elsewhere (McHugh et al., 2012; Scanlon et al., 2012b).
Grantees were broadly charged with helping providers in the community improve the quality of ambulatory chronic illness care. The alliances were given considerable latitude in pursuing their work. For example, alliances could establish their own activities related to chronic care improvement, partner with other organizations, or use a combination of the two approaches. Improving care management processes, encouraging the adoption of patient-centered medical homes, and reducing readmissions were the most commonly identified foci of alliances’ (and their partners’) interventions. The most widely used interventions were learning collaboratives, where local providers gather (in person or virtually) to discuss QI in a noncompetitive environment. Alliances also had the flexibility to focus on any number of health conditions, but diabetes and congestive heart failure were the most commonly targeted areas. More detailed information on alliances’ approaches to QI has been reported elsewhere (McHugh et al., 2012; Scanlon et al., 2012b).
Previous Efforts to Assess Intervention Dose
In February 2015, we conducted a PubMed search for articles published since 1990 that empirically measured the strength or quality of interventions in multisite QI initiatives beyond simply counting the number of participants or initiatives. We used the search terms “dose,” “strength,” “scope,” “intensity,” “implementation,” and “reach” with the term “quality improvement” and the terms “program evaluation” or “process evaluation.” Results, provided in Appendix A, showed that few have attempted to measure the strength or quality of QI interventions. None attempted to validate their dose measure. Most of the approaches used a combination of counts of QI interventions or participants in combination with authors’ qualitative assessment of the intensity of the intervention. Two used a structured scoring worksheet to facilitate interrater reliability (Pearson et al., 2005; Peikes, Chen, Schore, & Brown, 2009).
Notably, our literature search also found several commentaries that did not empirically measure dose but described important elements of dose that should be taken account in measurement (Cheadle et al., 2012; Hasson, 2010; Huber et al., 2001; Reed et al., 2007). Drawing from both the empirical articles and commentaries, we developed a list of possible dimensions to be included in a measure of dose (Table 1). While all of these dimensions of dose may be important, without further evidence, it is difficult to know which are most important to measure. There are trade-offs between using a simple method, such as counting the number of initiatives or participants, and a more comprehensive method that includes multiple dimensions and tries to capture implementation intensity. It is generally easy for researchers to obtain reliable data on the number of QI interventions employed. However, assessments of the intervention scope and the intensity of participant contact may require substantially more data collection and more subjective decisions by researchers, which may be prone to bias or error. Multiple raters would be needed to assure validity and reliability, which would add substantial cost to the evaluation. Therefore, understanding which dimensions most accurately reflect QI dose becomes especially important.
Dimensions of Intervention Dose.
Based on the elements of dose identified through the literature review, we identified four different approaches for measuring QI dose ranging from simple to comprehensive. Approach 1, Quantity, is simply a count of the number of QI interventions that were undertaken as part of AF4Q. This approach could be called the “more is better” method, as it gives credit to sites that engaged in the most QI interventions. However, the Quantity Approach does not take into account the size of the interventions or their potential to drive improvements. Approach 2, Intensity and Reach, relies on investigators’ judgment to classify the intensity and reach of the communities’ QI interventions. Intensity refers to the likelihood that a given intervention will change an outcome, with the assumption that more intense interventions will have a greater impact on the outcome (Duhon, Mesmer, Atkins, Greguson, & Ollinger, 2009). Reach refers to the percent of local providers (hospitals and physicians or physician practices) touched by the QI intervention.
Approach 3, Duration and Scope, takes into account the length of time in which the QI intervention was active. Scope refers to the intervention’s comprehensiveness, for example, the number of different outcome measures the intervention is designed to address. Finally, Approach 4 is the Comprehensive Approach and incorporates all of the elements of the previous three approaches.
Method
Data
In order to apply these four approaches to the AF4Q program, we drew on a number of data sources collected as part of the AF4Q evaluation. They included transcripts from two rounds of site visits conducted with alliance leaders in 2006 and 2010, as well as telephone interviews with alliance leaders conducted annually between 2010 and 2013. The interviews were recorded, transcribed, and uploaded into Atlas.ti, a qualitative analysis software program (Version 7.1, Scientific Software Development, 2014). Transcripts were reviewed and coded with predefined global codes that aligned with the AF4Q’s programmatic areas (e.g., QI, public reporting, consumer engagement, equity). All passages coded with the QI global code were reviewed. More information on the collection and analysis of interview data may be found elsewhere (Scanlon et al., 2012a).
Additionally, we reviewed funding proposals, work plans, and progress reports submitted by the alliances to the foundation between 2006 and 2013. Based on these data and the interviews, we developed a summary of the major QI interventions that the alliances pursued under AF4Q. In 2013, the individuals identified by alliance directors as being responsible for overseeing the alliances’ QI interventions were asked to verify and/or make corrections to our summaries.
Application of the Four Approaches
Approach 1: Quantity
We reviewed the summaries of the major QI interventions and ranked communities from high to low based on the number of QI interventions undertaken. We then divided them into approximately three equal groups. Communities in the “high-quantity” category had six or more QI interventions, communities in the “medium-quantity” category had four or five interventions, and communities in the “low-quantity” category had fewer than four QI interventions. We then assigned scores of 3, 2, or 1 to the high-, medium-, and low-quantity communities, respectively.
Approach 2: Intensity and Reach
Two AF4Q program evaluators (MM, JH) independently reviewed the summaries of QI interventions for each community and assigned communities to one of the three categories, based on whether the communities generally used high-, medium-, or low-intensity interventions. To facilitate interrater reliability and determine cutoff points, the evaluators were guided by examples of low-intensity interventions (e.g., educational sessions), medium-intensity interventions (e.g., learning collaboratives with intervention training and monitoring), or high-intensity interventions (e.g., individualized practice coaching) during their scoring, similar to an approach used by Wickizer et al. (1998). We then assigned scores of 3, 2, or 1 to the high-, medium-, and low-intensity communities, respectively.
Next, we ranked the communities based on the percentage of providers reached by the QI interventions and assigned them into three approximately equal groups (high, medium, and low, corresponding to 3, 2, and 1 scores, respectively). Communities scoring in the “high-reach” category had several large interventions involving more than 70% of local physicians or hospitals. Communities in the “low-reach” category were typically conducting pilot interventions with less than 25% of local providers participating. To combine the intensity and reach scores (giving equal weight to all), we simply aggregated them. Total possible scores for the Intensity and Reach Approach ranged from 2 to 6.
Approach 3: Duration and Scope
Duration refers to the total number of calendar years in which AF4Q QI interventions were active. For example, a community with two interventions that began in 2010 and ended in 2013 would sum to 8 years (i.e., two interventions each active for 4 years). Communities were assigned into three approximately equal groups—high, medium, and low duration, corresponding to scores of 3, 2, and 1, respectively. Communities in the “high-duration” group had durations of 15 or more years; communities in the “low-duration” group had durations of less than 10 years.
Scope refers to the extent to which the alliances adopted interventions targeted to address six different QI outcomes included in the AF4Q logic model: care coordination, patient satisfaction and experience, receipt of recommended care for chronic illness, emergency department visits and hospitalizations, inpatient processes of care, and adoption of health information technology. We reviewed the summaries of the QI interventions for each community and counted the number of outcomes targeted across all interventions. All communities targeted either three or four outcomes. Communities that targeted three outcomes received a score of 1, and communities that targeted four outcomes received a score of 2. To combine the duration and scope scores, we simply aggregated them. Total possible scores for the Duration and Scope Approach ranged from 2 to 5.
Approach 4: Comprehensive
The Comprehensive Approach takes into account all of the dose elements listed above. To combine the quantity, intensity, reach, duration, and scope dimension scores, we simply aggregated the dimension scores. Total possible scores for the Comprehensive Approach ranged from 5 to 15.
All scoring was conducted independently by two investigators (MM, JH). Initial agreement between the two researchers was 94%. All disagreements were resolved through conversation and review of the data until consensus was reached. The scoring sheet that guided the investigators’ decisions is attached as Appendix B.
Knowledgeable Informants’ Perceptions of Dose
To assess the convergent validity of our dose scores, we asked a group of individuals close to the AF4Q program for their opinions about which communities had the highest and lowest QI dose. These individuals were involved with managing the operation of the AF4Q program, but were not directly responsible for the alliances’ QI efforts, and played no role in the development or evaluation of the approaches to measuring dose examined in this article. We used purposive sampling in an effort to target respondents who were most likely to be familiar with alliances’ QI activities and who represent key stakeholder groups. Individuals were contacted by e-mail and asked to participate in a brief survey. Once they agreed, they were asked the following question: Relative to other AF4Q alliances, how would you rate the quantity and quality of each alliance’s AF4Q-related quality improvement (QI) interventions as a whole? For each community, possible responses included above average, average, below average, and do not know. The survey was approved by the The Pennsylvania State University Institutional Review Board.
Out of 23 individuals who were invited to take the survey, 13 (57%) completed it. Five respondents were from the National Program Office (38%), three from the foundation (23%), three were technical assistance providers (23%), and two were from the foundation’s communications firm (15%).
Analysis
We report total scores for each community under each of the four dose measurement approaches. Additionally, we categorized communities as low, medium, or high dose under each of the dose measurement approaches based on the relative score of each community.
Similarly, using data from the survey of knowledgeable informants, we categorized communities into roughly three equal groupings (above average, average, and below average) based on their relative scores. We calculated the level of concordance between scores from the dose measures and survey of knowledgeable respondents using Goodman–Kruskal’s (1954) gamma and the asymptotic standard error (ASE).
Results
Consistency of Dose Scores
Communities’ final scores across the four approaches are shown in Table 2. The scores for each dose element varied across communities. High-, medium-, and low-dose communities under the four approaches are also identified in Table 2. Categorization was consistent across the four approaches for three communities: Community I was a low-dose community under all four approaches, and Communities A and J were medium dose. There was strong agreement (i.e., consistent categorization across three approaches) for four communities: Community G was a low-dose community under three approaches, and Communities B, H, and M were high-dose communities. There was partial agreement for three communities: Communities C and E were consistently categorized as low- or medium-dose communities, and Community D was consistently categorized as a medium- or high-dose community across all four approaches. There were also clear inconsistencies across the approaches. For example, four communities (F, K, L, and N) were high-, medium-, or low-dose communities, depending on the approach used.
Final Scores, by Approach to Measuring Dose.
Concordance Between Dose Scores and Results From the Survey of Knowledgeable Informants
Drawing on results from the survey of knowledgeable informants, we categorized the communities as above average, average, or below average in terms of QI dose (see Appendix C). Table 3 shows the cross frequencies for the communities based on findings from the dose measures and the survey of knowledgeable respondents. For example, one community was a high-dose community under the Quantity Approach and was also considered an above-average community by knowledgeable informants. Table 3 also presents the level of concordance. There was poor concordance (gamma = −0.0455, ASE = 0.3191) between the Quantity Approach to dose measurement and knowledgeable respondents’ perceptions, much higher concordance with the Intensity and Reach Approach (gamma = 0.7917, ASE = 0.1424) and Duration and Scope Approach (gamma = 0.8065, ASE = 0.1956), and the highest concordance with the Comprehensive Approach (gamma = 0.9231, ASE = 0.0888).
Cross-Frequency of Dose Categories From the Four Approaches and the Survey of Knowledgeable Informants.
Limitations
There are several limitations worth noting. First, there is a myriad of different ways that we could have constructed the dose measures. Since our work is exploratory, we simply chose four approaches that varied in complexity. Second, our measures of dose—even the comprehensive measure—did not include all dimensions of dose that have been discussed in the literature. Despite our extensive data collection, our team did not have enough information to sufficiently assess participant engagement or intervention quality. Additionally, there may be other important elements of dose, for example, how well alliances’ QI interventions complement each other, that are absent from the literature, and therefore, our analysis. Third, our scores were based on the subjective assessments of reviewers. Furthermore, variation in scores may be a function of the scale of the rating system. Variation may have been greater if we used a 10-point scale for each dose element. Fourth, we checked the performance of our dose measures against knowledgeable respondents’ perceptions. Knowledgeable respondents’ perceptions is not necessarily a “gold standard,” but should theoretically be correlated with a valid measure of dose. Finally, our finding that the Comprehensive Approach was best aligned with perceptions of knowledgeable respondents may not be generalizable to other multisite QI initiatives. The QI intervention under investigation here was unique in its use of multistakeholder alliances as the instigators of change, focus on community-level interventions, and generous investment from the foundation.
Discussion
In the absence of a widely accepted approach for measuring QI dose, we reviewed the literature on QI dose, identified four methods for measuring dose, and applied them to 14 communities in the AF4Q program. Our effort resulted in a number of lessons learned.
First, relying on extensive data collected as part of the AF4Q evaluation, we found that it was feasible to assign values for the elements of dose deemed important in the literature (e.g., amount, duration, scope of activities). However, the data we used to assign values were collected longitudinally and drawn from multiple sources (e.g., interviews with site leadership, proposals, work plans). The extensive data needed to assign values to the various dose elements may be one reason that so few evaluation efforts include dose measurement in their analysis. Smaller evaluations with fewer opportunities for data collection may not have access to similar data. To measure dose, evaluators and their funders must make a substantial investment in data collection that includes longitudinal data collection and multiple investigators to make interrater comparisons.
Second, although there were some consistencies among the communities that were deemed to be “high dose” or “low dose” under the three approaches, ultimately the different approaches yielded different sets of “high-dose” and “low-dose” communities. This demonstrates that the way in which one defines dose can influence conclusions about which interventions sites are “high dose” or “low dose.”
Third, there was considerable variation in scores across communities when multiple elements of dose were taken into account. The differences in scores across sites suggests that dose may be an important contributor to the effectiveness of QI interventions.
Fourth, our work highlights challenges that investigators may encounter as they measure QI dose. For example, researchers need to make decisions about which elements of dose to include in their analyses. These decisions may vary depending on the effort under evaluation. For example, a highly prescriptive QI program with a defined set of activities and consistent start and end dates across intervention sites may not need to include counts or duration of QI activities in the dose measurement. Researchers investigating programs with fewer guidelines or restrictions on QI interventions may find a larger number of elements to be relevant. This suggests that there may not be one method of dose measurement that can be universally adopted by the field. Additionally, there is no clear cutoff points between high- and low-dose scores. As a result, in addition to the subjectivity of scoring for many of the dose elements, final identification of high- and low-dose communities is based on subjective decisions of researchers.
Fifth, the high level of concordance between our Comprehensive Approach to measuring dose and knowledgeable respondents’ perceptions suggests that the Comprehensive Approach may be the best approach for measuring QI dose for the AF4Q program. The low level of agreement between our Quantity Approach and knowledgeable respondents’ perceptions suggests that evaluators of complex, multisite QI initiatives should not simply count activities as a proxy for QI dose. At least for the AF4Q program, the more dimensions of dose that were included, the stronger the agreement with knowledgeable respondents’ perceptions.
Finally, our work suggests a number of areas for future research. The tremendous growth in the QI field since the publication of To Err is Human (Kohn, Corrigan, & Donaldson, 2000) has resulted in a need to understand the factors associated with success. While there appears to be agreement in the literature that QI dose is important, the limited number of attempts to measure dose across intervention sites suggests a need for greater dialogue among researchers about how to approach dose measurement. Such a conversation might be facilitated by foundations and federal agencies that support QI research. Additionally, measurement of dose, including our own work, has so far been conducted retrospectively. Future evaluations of multisite QI interventions should include dose measurement in the design and planning stage and collect data prospectively.
Finally, while there is an assumption in the field that QI dose can influence outcomes, empirical evidence is limited. Further study is needed to investigate the link between QI dose and outcomes, and how other factors known to influence outcomes (e.g., resources, leadership) might moderate that relationship. Exploring the linkage between dose measures and outcomes was beyond the scope of this article, but will be examined as final AF4Q program outcomes become available. If ultimately we believe that greater dose is associated with greater impact, then understanding the factors driving QI dose would be important to program planners and policy makers.
Footnotes
Appendix B
Appendix A
Previous Efforts to Measure the Dose of a QI Intervention.
| Author (year) | Description of measure reported in the article | Definition/method(s) | Intervention sites | QI focus |
|---|---|---|---|---|
| Shortell et al. (1995) | Degree of QI implementation | Employee assessments across five areas used for the Malcolm Baldridge National Quality Award: leadership, information and analysis, human resources management, quality management, and strategic quality planning | Hospitals | Improving clinical efficiency through organizational culture and quality improvement implementation |
| Wickizer et al. (1998) | Intensity and exposure | 1. Authors’ qualitative ratings of intervention intensity (high, medium, low) 2. Number of exposures to the intervention per 1,000 population |
Communities | Reducing risk factors for adolescent pregnancy, cancer, cardiovascular disease, injuries, and substance abuse |
| Montori et al. (2002) | Intensity | Item count of the number of entries in the medical record that were related to the intervention | Primary care practices | Improving diabetes outcomes through practice guidelines, support for self-management, and clinical information systems |
| Pearson et al. (2005) | Intensity | 1. Count of the organization’s change activities 2. Authors’ qualitative rating of the depth of change |
Health care organizations | Improve care for congestive heart failure, diabetes, depression, and asthma through implementation of the chronic care model |
| Peikes et al. (2009) | Program effort and number of program enrollees | Program effort measured through authors’ qualitative rating of strength of 10 different aspects of the QI program (e.g., program staffing, patient education) on a scale of 1 to 5. | Hospitals | Reducing hospitalizations and Medicare expenditures and improving quality of care for chronically ill Medicare beneficiaries through care coordination. |
| Bosworth et al. (2010) | Dose and intensity | 1. Frequency of use of the intervention and number of written materials provided at the intervention sites 2. Attendance on QI conference calls |
VA facilities | Improving blood pressure control through an evidence-based nurse-led patient self-management intervention |
| Nutting et al. (2010) | Percent of intervention components implemented | The proportion of the intervention components that the practices implemented | Physician practices | Improving primary care through the implementation of patient centered medical homes |
| Van der Veer et al. (2011) | Exposure | Amount of time spent on QI activities | Hospital intensive care units (ICUs) | Improving the quality of ICU care through feedback data, establishment of a QI team, and educational outreach visits |
| Huis et al. (2013) | Dosage and coverage | 1. Dosage: Dichotomous variable indicating whether QI activities were delivered as often and as long as planned 2. Coverage: Extent to which the intended target group received the improvement activities |
Hospitals | Improving nurse hand hygiene through provider education and feedback |
| Halladay et al. (2014) | Degree of QI implementation | Practice coaches score the implementation of each activity (disease registry; use of a care protocol, use of planned care templates, and development of self-management support tools for patients) on a scale from 0 to 5, where 0 indicates the practice has no activity in the area and 5 indicates the activity has been implemented practice-wide | Physician practices | Improving chronic disease care through learning collaboratives and practice coaches |
Note. QI = quality improvement.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by the Robert Wood Johnson Foundation, Grant Number 70877.
