Abstract
Background. This article identifies the most influential methods reports for group-randomized trials and related designs published through 2020. Many interventions are delivered to participants in real or virtual groups or in groups defined by a shared interventionist so that there is an expectation for positive correlation among observations taken on participants in the same group. These interventions are typically evaluated using a group- or cluster-randomized trial, an individually randomized group treatment trial, or a stepped wedge group- or cluster-randomized trial. These trials face methodological issues beyond those encountered in the more familiar individually randomized controlled trial. Methods. PubMed was searched to identify candidate methods reports; that search was supplemented by reports known to the author. Candidate reports were reviewed by the author to include only those focused on the designs of interest. Citation counts and the relative citation ratio, a new bibliometric tool developed at the National Institutes of Health, were used to identify influential reports. The relative citation ratio measures influence at the article level by comparing the citation rate of the reference article to the citation rates of the articles cited by other articles that also cite the reference article. Results. In total, 1043 reports were identified that were published through 2020. However, 55 were deemed to be the most influential based on their relative citation ratio or their citation count using criteria specific to each of the three designs, with 32 group-randomized trial reports, 7 individually randomized group treatment trial reports, and 16 stepped wedge group-randomized trial reports. Many of the influential reports were early publications that drew attention to the issues that distinguish these designs from the more familiar individually randomized controlled trial. Others were textbooks that covered a wide range of issues for these designs. Others were “first reports” on analytic methods appropriate for a specific type of data (e.g. binary data, ordinal data), for features commonly encountered in these studies (e.g. unequal cluster size, attrition), or for important variations in study design (e.g. repeated measures, cohort versus cross-section). Many presented methods for sample size calculations. Others described how these designs could be applied to a new area (e.g. dissemination and implementation research). Among the reports with the highest relative citation ratios were the CONSORT statements for each design. Conclusions. Collectively, the influential reports address topics of great interest to investigators who might consider using one of these designs and need guidance on selecting the most appropriate design for their research question and on the best methods for design, analysis, and sample size.
Keywords
Introduction
Murray et al. 1 recently reviewed the development of methods for the design and analysis of group- or cluster-randomized trials, individually randomized group treatment trials, and stepped wedge group- or cluster-randomized trials based on a review of reports published through 2018. This article provides an update to that article, reviewing reports published through 2020. After describing the key features for these three types of randomized trials, the most influential reports for the development of these methods are identified and described.
Key features
Parallel group- or cluster-randomized trials
A group-randomized trial randomizes groups rather than individuals to study conditions and outcomes are measured for participants from each group.1–7 In the parallel group-randomized trial considered here, there is no crossover of groups to a different study condition during the trial. The parallel group-randomized trial is the best comparative design available when there is a good reason for randomization of groups rather than individuals. The usual reasons are (1) concern for contamination across conditions if delivered within the same group or (2) the use of a group-based intervention.
The key feature of the group-randomized trial is the randomization of groups to study conditions. Outcomes on participants from the same group are expected to be positively correlated as a result of common exposures, shared experience, or participant interaction. 8 This correlation violates the assumption of independence of errors that underlies the familiar analytic methods for individually randomized controlled trials.1–7 This correlation is often measured by the intraclass correlation.
Murray et al. 9 recently characterized the design features for group-randomized trials involving cancer or cancer-related outcomes. Most group-randomized trials compared two study conditions using a pretest–posttest design. Some used a posttest-only design, while others included multiple pretest and/or posttest measures. Most employed a cohort design observing the same participants from each group at each measurement occasion. Others employed a cross-sectional design observing different participants from each group at each measurement occasion. Still others included both a cohort design and a cross-sectional design in the same study. Most employed some form of restricted randomization, such as stratification, matching, or constrained randomization.
Despite the availability of many books and hundreds of papers on the design and analytic methods for group-randomized trials, reviews have regularly shown that a large proportion of published group-randomized trials fail to account for the intraclass correlation in either sample size calculations or in data analysis.9–20 This is problematic because undersized studies will have insufficient power for a valid analysis and because invalid analyses will have an inflated type 1 error rate.9,12–19,21–37
Stepped wedge group-randomized trials
Stepped wedge group-randomized trials have increased in popularity over the last 15 years and are now commonly used to evaluate interventions that test new approaches to delivery of care. 38 Stepped wedge group-randomized trials are more complex than either parallel group-randomized trials or individually randomized group treatment trials and they also face greater risks for bias.39,40
The key feature of the stepped wedge group-randomized trial is the crossover of each group from the control to the intervention condition in a random order and on a staggered schedule. 41 Observations from participants from the same group will be positively correlated as in parallel group-randomized trials; however, the impact of the intraclass correlation is reduced in the stepped wedge group-randomized trial because the groups are crossed with study conditions rather than being nested within study conditions. Unlike parallel group-randomized trials, the intervention effect is confounded with calendar time.41,42 Moreover, the effect of the intervention may vary depending on how much time has passed since the intervention was introduced;42,43 that is important in a stepped wedge group-randomized trial because they are often longer than a parallel group-randomized trial. Finally, the pattern of correlation over time can be complex because stepped wedge group-randomized trials involve repeated measurements on the same groups and sometimes on the same participants, often for a prolonged period.44–46
The most common design for a stepped wedge group-randomized trial is a complete design; here, data are collected when all groups are in the control condition, again in each group when one or more groups cross over to the intervention condition, and usually again after all groups are in the intervention condition. The incomplete design is less common; here, data are not collected from all groups at all steps. 47 Stepped wedge group-randomized trials vary in the number of groups that crosses over in each step, the number of steps, and in the time between steps. They may employ restricted randomization, as described above, to balance groups in the sets to be randomized to the steps. In the continuous recruitment short exposure design, participants are recruited continuously and exposed for only a short period; participants may be measured only once or repeatedly. In the closed cohort design, participants are identified at the beginning of the study, participate throughout the study, and are measured repeatedly. In the open cohort design, many participants are identified at the beginning of the study, but some may leave while other participants are recruited over time; participants may be measured only once or at multiple occasions.
The literature on the design and analytic methods for stepped wedge group-randomized trials has developed rapidly over the last 15 years. Reviews have noted deficiencies in reporting of study design, sample size, analytic methods, and ethical conduct.48–57
Individually randomized group treatment trials
An individually randomized group treatment trial differs from an individually randomized controlled trial in that the method by which the intervention is delivered creates a level of intraclass correlation that otherwise would not exist. 34 Correlated observations can result if participants receive at least some of their intervention in a group format (e.g. attend the same weight-loss class), if participants share the same interventionist (e.g. have the same instructor, therapist, or surgeon), or if participants interact with one another in some other way that is related to the method in which the intervention is delivered (e.g. through a virtual chat room created for participants in the same study condition). The individually randomized group treatment trial is the best comparative design available if randomization of individuals is possible but it is necessary or more efficient to deliver at least some of the intervention in a group format or through a shared interventionist.
The key feature of the individually randomized group treatment trial is that the method by which the intervention is delivered generates some level of correlation among outcomes taken on groups of participants within the same study condition, creating the same type of intraclass correlation seen in parallel group-randomized trials. Investigators must account for the intraclass correlation in the sample size to avoid low power and in the analysis to avoid type I errors.34,58–62 If the method of intervention delivery creates multiple overlapping groups or if the group structure changes over time, the situation is even more complicated, further increasing the risk of an inflated type I error rate if the investigators do not account for the complex pattern of correlation.63–66
The methods literature for individually randomized group treatment trials is much more limited than for the other designs considered here and the issues are not widely recognized. Reviews suggest that most investigators who employ the individually randomized group treatment trial design are not aware of it and do not use appropriate methods for sample size or analysis.34,66–69
Influential methods reports for these designs
Methods
The primary objective of this article was to identify the most influential reports in the development of methods for group-randomized trials and related designs through 2020. In the development of an earlier paper, 1 PubMed was searched to identify all methods papers related to these designs through 2018. Search terms used in the initial PubMed search were (“Randomized Controlled Trials as Topic” (MeSH Terms) AND “Cluster analysis” (MeSH Terms) AND (“sample size” (MeSH Terms) OR “computer simulation” (MeSH Terms) OR “research design” (MeSH Terms) OR “Analysis of Variance” (MeSH Terms) OR “Epidemiologic Research Design” (MeSH Terms) OR “Random Allocation” (MeSH Terms)) AND (“1975/01/01” (PDAT): “2018/12/31” (PDAT)) OR ((cluster *(Title) OR group*(Title) OR community(Title)) AND (random*(Title) OR RCT(Title)) AND (analysis* (Title) OR design(Title) OR method(Title) OR sampl* (Title)) AND (“1975/01/01” (PDAT): “2018/12/31” (PDAT)). Additional searches were conducted for all papers published by each of the first authors identified in the initial search.
The results were augmented by articles, books, chapters, and reports known to the authors of the earlier report. PubMed was then searched for other papers by any of the authors of the identified reports, yielding a total of 4514 candidates. The author reviewed each report and any that were not focused on design or analytic methods for the designs of interest were excluded, leaving 924 reports; 797 focused on parallel group-randomized trials, 49 on individually randomized group treatment trials, 74 on stepped wedge group-randomized trials, and 4 addressed all three designs.
Using the same methods, another 119 reports were identified through 2020, bringing the total to 1043 reports through 2020. That total included 873 reports focused on group-randomized trials, 55 focused on individually randomized group treatment trials, 105 focused on stepped wedge group-randomized trials, and 10 that addressed two or more of these designs.
Citation counts and the relative citation ratio (RCR) were used to assess the influence of each report, accessing data on 27 April 2021. The RCR uses relative citation rates to measure influence at the article level, standardized across areas of science defined by the relative citation network (the articles cited by other articles that also cite the reference article). 70 In brief, the citation rate of the reference article is compared to the citation rates of the relative citation network. RCRs change over time as the relative citation network grows, particularly in the first few years after a report is published; as a result, the RCR for some of the reports included in the earlier paper had changed by the time the data were accessed for this article 2 years later.
The RCR was available for 898 reports; for the remaining 145 reports, the Web of Science, Scopus, and Google Scholar were used to obtain citation counts. Group-randomized trial reports (N = 32) and reports addressing two or more of the study designs (N = 0) with an RCR ≥ 7.98 (99th percentile for all PubMed entries) or without an RCR but with ≥ 200 citations were retained. Individually randomized group treatment trial reports (N = 7) with an RCR ≥ 3.45 (95th percentile) or without an RCR but with ≥ 100 citations were retained. Stepped wedge group-randomized trial reports (N = 16) with an RCR ≥ 4.91 (97.5th percentile) or without an RCR but with > 150 citations were retained. The citation count thresholds generally discriminated between the reports that fell above and below the RCR thresholds; the sliding scale reflected the number of reports for each design with many more for group-randomized trials and many fewer for individually randomized group treatment trials. This provided 55 methods reports related to these designs that were deemed most influential (Tables 1–3). These reports do not necessarily represent the current state of the science; recent summaries of the state of the science are available elsewhere.1,71–73
Influential reports for group-randomized trial methods sorted by RCR and citation count with explanation for selecting items; RCR and count data were retrieved on 27 April 2021.
Reports in bold type were not included as influential reports in the earlier paper (1).
Influential reports for individually randomized group treatment trial methods sorted by RCR and citation count with explanation for selecting items; RCR and count data were retrieved on 27 April 2021.
Reports in bold type were not included as influential reports in the earlier paper (1).
Influential reports for stepped wedge group-randomized trial methods sorted by RCR and citation count with explanation for selecting items; RCR and count data were retrieved on 27 April 2021.
Reports in bold type were not included as influential reports in the earlier paper (1).
The next three sections identify the influential reports for the three designs. The earliest reports are mentioned first in each section and the remaining reports are grouped by theme.
Parallel group-randomized trials
In 1978, Cornfield published the first methods paper for trials involving group randomization in the biomedical literature. He identified two penalties associated with group randomization: extra variation attributable to the group and limited degrees of freedom for the test of the intervention effect. 99 Both penalties must be addressed in studies that randomize groups rather than individuals.
Several influential papers addressed general issues for parallel group-randomized trials. A 1997 commentary drew attention to the methodological issues inherent in these trials. 77 A 1997 paper presented features for the optimal design of a parallel group-randomized trial. 98 A 2003 paper reported on the use of parallel group-randomized trials for evaluating the effectiveness of change and improvement strategies. 87 A 2013 paper reported on methods for process evaluation. 81 A 2018 paper reported on the use of parallel group-randomized trials in dissemination and implementation research. 93
Others presented methods for sample size calculations based either on the intraclass correlation76,78,88,89,97 or on methods based on the coefficient of variation. 79 Eldridge et al. 90 and Rutterford et al. 95 addressed the question of unequal sample size at the group or cluster level. Rutterford et al. 95 also addressed the effect of attrition, non-compliance, adjustment for baseline covariates, and repeated measures on sample size estimation.
Others focused on specific analytic methods. Those included methods for random coefficient or growth curve models, 92 binary data, 80 ordinal data,85,91 Poisson regression, 83 meta-analysis,82,100 mediation analysis,84,94 and survival analysis. 96
The first textbook on the design and analysis of parallel group-randomized trials was published in 19982 followed in 2000 by the second. 3 Three subsequent textbooks also met the criteria to be judged an influential report.4,7,113
The first CONSORT statement on parallel group-randomized trials was published in 200474 and provided a checklist to identify the methodological information to include in trial reports. An update was published in 2012. 75
Murray et al. 86 published a review of methodological issues in parallel group-randomized trials in 2004, summarizing work on both design and analytic methods.
Stepped wedge group-randomized trials
Though the concept of the stepped wedge design was introduced in 1987, 114 Hussey and Hughes 41 published the first analytic methods for that design in 2007. Copas et al. 47 later delineated three types of stepped wedge group-randomized trials and discussed the number and length of the steps, incomplete and complete designs, and randomization methods, including restricted randomization methods. Hemming et al. 42 provided guidance on the rationale, design, analysis, and reporting for this design. They noted that the stepped wedge group-randomized trial is particularly well suited for evaluations of health service delivery interventions. Barker et al. 53 reviewed the statistical methods used in stepped wedge group-randomized trials. Hemming et al. 111 summarized methods appropriate for stepped wedge group-randomized trials with repeated cross-sectional samples.
Several papers addressed sample size methods. Woertman et al. noted that the sample size will depend on group size, the intraclass correlation, the number of steps, the number of baseline measurements, and the number of measurements between steps; 106 a subsequent letter 115 corrected an important error in that paper which was accepted by the authors. 116 Baio et al. 110 presented simulation methods for sample size estimation. Hemming et al. 107 considered power for several different stepped wedge group-randomized trial designs, including incomplete cross-sectional designs, designs with multiple levels of nesting, and complete designs. Girling and Hemming 112 proposed an algorithm to optimize the design of stepped wedge group-randomized trials and showed that for large studies, the best design may be a hybrid of parallel group-randomized trial and stepped wedge group-randomized trial design components. Hemming and Taljaard 108 compared power for parallel group-randomized trials and stepped wedge group-randomized trials when the number of groups is fixed and reported that the parallel group-randomized trial tends to be more efficient when the intraclass correlation is small and that the stepped wedge group-randomized trial tends to be more efficient when the intraclass correlation is large, dependent on group size. Hooper et al. 109 presented formulaic methods for sample size estimation based on the intraclass correlation and the cluster and individual autocorrelations. Kasza et al. 44 reported on the effect of decaying correlation over time on sample size and power.
Several state-of-the-practice reviews of stepped wedge group-randomized trials have been published,48–50 reporting wide variation in data analytic methods and reporting standards. The recent CONSORT statement for this design 105 was written in part to try to reduce that variation.
Individually randomized group treatment trials
Several early reports addressed the risk of type 1 error in studies in which individually randomized participants receive their intervention in a group format or from a shared interventionist. These papers appeared in the biomedical, 103 psychological,68,102 and educational literature. 104
Roberts and Roberts 58 presented the first report to address the analytic challenges specific to individually randomized group treatment trials in 2005. Consistent with guidance for parallel group-randomized trials, they noted that mixed models specifying the groups as levels of a random effect provided an appropriate analysis.
In 2008, Boutron et al. 60 extended the CONSORT statement to non-pharmacologic interventions which include many individually randomized group treatment trials. They pointed to the need to provide details on how the intervention was delivered (e.g. individually, in groups, through a common interventionist) and to address the implications for the analysis of having correlated observations within one or more study conditions. Boutron et al. 101 published an update to that CONSORT statement in 2017.
Summary
This article identifies the 55 most influential reports contributing to the methods for the evaluation of group- or cluster-randomized trials, individually randomized group treatment trials, and stepped wedge group- or cluster-randomized trials through 2020, adding 2 years of data to an earlier paper. 1 There was substantial overlap with the 50 reports listed in the earlier paper, but there was also turnover in the list of influential reports, based on changes to the RCRs and citation counts as often happens with the passage of time. There were six new parallel group-randomized trial reports, one new individually randomized group treatment trial report, and five new stepped wedge group-randomized trial reports included here. Of the original 50 reports, 3 individually randomized group treatment trial reports, 2 stepped wedge group-randomized trial reports, and 2 reports that addressed more than two designs no longer met the inclusion criteria and so were not included in this report. Reports that had been included previously but failed to meet the inclusion criteria for this update generally had low but qualifying RCRs for the original report and those values declined with the additional follow-up time. RCRs are generally stable after a few years and most of the reports for which the ratio changed were relatively new in the original report.
The RCR and citation counts were used to gauge the influence of a report on the work published later. Influence does not equate to quality, so there is no claim that the most influential reports are also the highest quality reports.
Many of the influential reports were early publications that drew attention to the issues that distinguish these designs from the more familiar individually randomized controlled trial. Others were textbook treatments that covered a wide range of issues for these designs. Others were “first reports” on analytic methods appropriate for a specific type of data (e.g. binary data, ordinal data), for features commonly encountered in these studies (e.g. unequal cluster size, attrition), or for important variations in study design (e.g. repeated measures, cohort versus cross-sectional). Many presented methods for sample size calculations. Others described how these designs could be applied to a new area (e.g. dissemination and implementation research). Among the most influential reports were CONSORT statements which provide guidance on how to present the methods and results from a study based on its design. Collectively, they address topics of great interest to investigators who might consider conducting a group- or cluster-randomized trial, an individually randomized group treatment trial, or a stepped wedge group or cluster-randomized trial and need information to guide their planning for design, analysis, and sample size. As a set, they make an excellent reading list for anyone interested in learning about the methods used in these designs and their appropriate applications in public health and medicine.
Footnotes
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
