Abstract
Background
Cluster randomized trials are designed to evaluate interventions at the cluster or group level. When clusters are randomized but some clusters report no or non-analyzable data, intent-to-treat analysis, the gold standard for the analysis of randomized controlled trials, can be compromised. This article presents a very flexible statistical methodology for cluster randomized trials whose outcome is a cluster-level proportion (e.g. proportion from a cluster reporting an event) in the setting where clusters report non-analyzable data (which in general could be due to nonadherence, dropout, missingness, etc.). The approach is motivated by a previously published stratified randomized controlled trial called, “The Randomized Recruitment Intervention Trial (RECRUIT),” designed to examine the effectiveness of a trust-based continuous quality improvement intervention on increasing minority recruitment into clinical trials (ClinicalTrials.gov Identifier: NCT01911208).
Methods
The novel approach exploits the use of generalized estimating equations for cluster-level reports, such that all clusters randomized at baseline are able to be analyzed, and intervention effects are presented as risk ratios. Simulation studies under different outcome missingness scenarios and a variety of intra-cluster correlations are conducted. A comparative analysis of the method with imputation and per protocol approaches for RECRUIT is presented.
Results
Simulation results show the novel approach produces unbiased and efficient estimates of the intervention effect that maintain the nominal type I error rate. Application to RECRUIT shows similar effect sizes when compared to the imputation and per protocol approach.
Conclusion
The article demonstrates that an innovative bivariate generalized estimating equations framework allows one to implement an intent-to-treat analysis to obtain risk ratios or odds ratios, for a variety of cluster randomized designs.
Keywords
Introduction
Randomized controlled trials are the gold standard of public health research. Typically, the unit of randomization and analysis is the individual. Cluster randomized trials are trials designed to evaluate interventions that operate at a cluster level. 1 Such studies may manipulate the physical or social environment such that intervention cannot feasibly be delivered to individuals.2–5 Examples include interventions delivered to schools, workplaces, or hospitals.6–10 The unit of analysis for cluster randomized trials may be the individual or the cluster.
In order to uphold randomization and ensure unbiased (causal) effects are estimated in the primary outcome analysis, all randomized trials should be analyzed on the principle of intent to treat. Under intent to treat, participants or clusters are analyzed as members of the treatment group to which they were randomized regardless of their adherence to, or whether they received, the intended treatment.11–13 Thus, it ignores nonadherence, protocol deviations such as randomization errors, withdrawal, dropout, and anything that happens after randomization, most of which are inevitable in human trials. 14 In the setting of cluster randomized trials, a design effect is typically included to model the correlation between individuals within a cluster. 15 However, applying the intent-to-treat principle to a cluster randomized can be more difficult because the issues of non-compliance, dropout, and missing data can be more complex. In fact, research has shown the intent-to-treat principle is more difficult to adhere to in cluster randomized trials, as loss of an entire cluster (versus say, one individual) both compromises inference and potentially decreases statistical power.8,16–18 Researchers have presented remedies for this issue including randomizing clusters only when the first participant is included to prevent empty clusters (index case concept), 17 or for trials already in progress, using ad hoc missing data or propensity score methods to accommodate missing outcome data at the individual or cluster level. Others propose the use of an Expectation-Maximization algorithm to obtain unbiased intent-to-treat principle estimates. 15
The current research is motivated by a stratified cluster randomized trial where 8 of 50 randomized clusters reported 0/0 proportion data. The Randomized Recruitment Intervention Trial (RECRUIT, ClinicalTrials.gov Identifier: NCT01911208) has been previously described in great detail. 10 Briefly, RECRUIT was the first multi-site randomized controlled trial to examine the effectiveness of a trust-based continuous quality improvement intervention aimed at health care workers, to increase minority recruitment into clinical trials. Four multi-site randomized controlled trials (parent trials) supported by three National Institutes of Health participated in RECRUIT. Fifty sites (i.e. clusters in this setting) within the 4 parent trials were randomized, 24 to intervention, and 26 to no intervention. Sites, the unit of analysis, were matched within parent trial on site characteristics. Overall, 26 intervention and 24 control sites were enrolled. The primary outcome was the site-level (cluster-level) proportion minority enrollment (to produce 50 outcomes for analysis). A simple schematic of the design can be viewed in Figure 1 of the design paper. 10 Briefly, the study was powered to detect a 0.10 absolute difference in intervention versus control proportions of minorities recruited assuming a 2-sample test and intra-cluster (site) correlation of 0.10. 10 A generalized estimating equation (GEE) approach to model the proportion minority enrolled was pre-specified to account for clustering of people within a site; GEE provides consistent standard errors even when the correlation structure is incorrectly specified 19 and is appropriate for 50 or more clusters in a cluster randomized trial. 20
An unforeseen issue that arose mid-trial was that 8 of 50 sites were unable to enroll any patients. This led to an analytic challenge since the proportion minority enrolled for those eight sites returned non-analyzable data. As omitting eight randomized sites could compromise the final analysis, this article proposes a remedy that enables the inclusion of all sites. This article proposes and assesses via simulation, a remedy for analyzing the effect of the intervention on a cluster-level outcome measured as either a proportion or a count, when some clusters report non-analyzable data, or do not report outcomes. Specifically, an exact statistical approach using GEEs for estimation is developed, tested via simulation, and applied to the RECRUIT study. The RECRUIT study is used as an example throughout the following development to demonstrate the flexibility of the approach to both straightforward and more complex cluster randomized designs.
Methods
The relevant Institutional Review Board provided approval for the RECRUIT Study. The below formulation will be described in terms of the RECRUIT data example for ease of presentation, but is generalizable to any cluster randomized trial where the outcome is a cluster-level proportion. Since the cluster in RECRUIT is a site, the term “site” is used henceforth in place of “cluster.” Let
The sites with no participants enrolled provided an undefined value for the proportion of minorities enrolled (0/0) at that clinic. The following presents several possible approaches to the analysis in the presence of clusters that report no or undefined data; the last approach derived is the true intent-to-treat analysis, as it analyzes all clusters as they were randomized.
Method 1: a GEE model for individual-level outcomes
The primary analysis plan assumed all clusters would report whether or not a person enrolled was a minority. The pre-specified analysis was therefore a binomial GEE with logit link function to model the individual-level outcome,
Sites enrolling no individuals (i.e.
To facilitate the pre-planned primary analysis while still including the eight sites with
Method 2: a generalized linear model for cluster-level outcomes
An alternative procedure is to consider modeling
where
The parameter of interest is still
Method 3: intent-to-treat analysis for cluster-level outcomes
Since none of the aforementioned approaches are ideal methods of analysis due to either imputation or omission of randomized units, a novel intent-to-treat GEE approach was formulated based on a joint analysis of two correlated components: the numerator and the denominator of the outcome proportion. This simple approach treats the cluster-level outcome as a bivariate vector, representing a numerator and a denominator produced from each site (cluster). The formulation is used for convenience because as shown below, a simple reparameterization enables a cluster randomized intent-to-treat analysis, whereas any other obvious approach would require omission of sites, imputation, expectation maximization (EM), or otherwise. 15
The trick to enabling intent-to-treat GEE analysis when some clusters do not provide analyzable proportion outcome data is by jointly modeling the numerator and denominator “components” such that sites that report no outcome, still enter the statistical model. The outcome is therefore now defined as a bivariate vector,
Specifically, consider the joint model of the mean of
In this model, the interaction between
From these formulae, the treatment effect on the proportion of interest for the cluster randomized trial is simply
Therefore, by modeling the cluster-level count outcomes, this model formulation allows for the estimation of the intervention effect on the proportion of interest.
For the RECRUIT stratified cluster randomized trial, assume two parent trials, trial A and trial B. Simply incorporate a parent trial indicator
The parameter of interest is still
Thus, we have
where
It is clear from above that
To conduct the analysis under the model given by equations (3) or (5), one adopts existing GEE software.
19
Specifically, an overdispersed Poisson distribution or a negative-binomial model with working independence correlation structure and log link may be specified. As the GEE model focuses on estimating the marginal means, it makes little assumptions about the joint distribution of
For all three GEE-based methods (i.e. available cases GEE, imputation GEE, intent-to-treat GEE), empirical standard error estimates are adopted. It is known that the empirical standard errors may be slightly biased downward, and the bias becomes more noticeable when the number of clusters is smaller than 50. 27 Researchers have discussed several bias correction methods for the empirical standard error estimates, which are implemented in the simulation study and application using the R package geesmv. 27
Simulation study
Simulation setup
A simulation study is conducted to evaluate the performance of the intent-to-treat GEE approach versus the alternative approaches. The R code for data generation and model fitting is presented in the Supplementary Material (See Supplemental Material). The goal of the study is not necessarily to demonstrate exceptional performance, but to show near-equivalence between the methods such that the intent-to-treat GEE analysis may be used in place of the others in settings of either missing or non-analyzable outcome data. To parallel the RECRUIT design, two parent trials (trial A and B) with a total of 50 sites within those trials are assumed. The intervention indicator variable,
for
First, consider a simpler scenario without a parent trial effect, setting
Next, consider the scenario with a parent trial effect, with
All simulated data are generated using the statistical package R. The simulated datasets are then analyzed using the function glm in R for the two GLM models and using the gee and geesmv package in R for the three GEE models.
Simulation results
Table 1 displays the result from all five approaches presented in section “Methods” under the typical cluster randomized trial scenario without a parent trial effect in terms of bias, relative bias, standard deviation of the estimates, average of estimated standard errors, relative standard errors, and 95% coverage probability. Note that for the two approaches based on the GEE model with binomial distribution, namely the Imputation-GEE and the available cases-GEE,
Simulation results for the scenario with a trial effect.
Table 1 indicates that the intent-to-treat GEE approach has comparable bias under a variety of null and clinically meaningful non-null intervention effects, and correlation parameters (columns). The bias tends to be downward in scenarios without treatment effects and upward in scenarios with positive treatment effects. However, the degree of bias is small; the relative bias is smaller than 3.2% in all settings. The relative bias increases slightly with the intracluster correlation coefficient for all methods, likely because that larger correlations correspond to a reduction in the effective sample size.
29
The standard deviation and average of estimated standard error reflect one another reasonably well, as evidenced by relative standard errors that are close to 1. The 95% coverage probabilities are near the nominal level for all but the two GLM approaches. Similarly, the type I error rates, which correspond to 1 minus the coverage probability when
Table 2 shows the results under the scenario with a parent trial effect. In the presence of strata (trial), the data generation scheme satisfies models (2) and (5) but not the binomial GEE model, as it is difficult if not impossible to generate data that satisfy the assumptions of all models simultaneously. Therefore, the two binomial GEE approaches are not implemented here to avoid unfair comparisons. The intent-to-treat GEE continues to perform very satisfactorily in terms of bias and average of estimated standard errors, with coverage rates that are close to the nominal level of 95%. By comparison, the GLM approaches give coverage rates that are lower than the nominal level, likely because its standard error estimates tend to be smaller than the empirical counterparts.
Simulation results for the scenario with a trial effect.
The conclusion of the simulation study is that the performance of the intent-to-treat GEE approach is competitive when compared to alternative methods and therefore could be used in place of imputation and available case analysis in order to include all clusters that were randomized.
Data application
The approaches above are applied to the RECRUIT study using a sensitivity analytic approach, to complement the analysis reported in the primary paper. From exploratory model fit criteria, it was determined a NB distribution was appropriate. As presented in the “Methods” section, each model fit includes the three trial indicators to adjust for these design effects. The intent-to-treat GEE additionally includes the interactions between treatments and trial as necessary to produce the intent-to-treat intervention effect.
Table 3 presents the intervention effect, the standard error of the intervention effect, the 95% confidence interval, and associated p-values resulting from applying the five methods to RECRUIT. The estimated coefficients correspond to
RECRUIT data analysis results for the intervention effect.
SE: standard error (without correction); CI: confidence interval; Interaction: component interaction from equation (3); GEE: generalized estimating equations; GLM: generalized linear model.
Coefficient corresponds to
Discussion
This article presents a novel intent-to-treat approach to analyzing cluster randomized trials in the difficult setting where a binary cluster-level outcome is non-analyzable. The simple and easy-to-implement approach decomposes the proportion outcome into numerator and denominator counts to facilitate an exact method such that all randomized units can be included in the analysis and rate ratios for the intervention effect can be produced. This is achieved by modeling the bivariate “count” vector via a GEE approach for clustered count data, or any other statistical method that can accommodate the now-bivariate count outcome in the context of clustered data (e.g. a generalized linear mixed model). Unlike previous work, the solution does not require ad hoc missing data methods such as imputation or expectation maximization and is applicable to the very common setting when cluster-level proportions or counts are the outcome of interest.15,17 On the other hand, simulation results suggest that the available case binomial GEE is also acceptable if the pre-specified analysis plan does not require the inclusion of all randomized units.
In cluster randomized trials where the group rather than individual is randomized, the intent-to-treat principle is often challenging to implement because of the lack of statistical methods to handle empty clusters. Oftentimes, clusters are discarded from the analysis. 17 While methods of imputation have been proposed, there is currently no clear solution to the current problem in the literature.15,17 The solution presented in this article is viable, easy to implement, and as shown via the application in this article, can be adapted to even more complex designed such as stratified cluster randomized trials. The approach could potentially be adaptable to other complex designs such as stepped wedge trials, where clusters are randomized to different sequences over time; more research into this would be needed as such trials have the additional complexity of potential confounding by time.30,31 Furthermore, if one wanted to obtain the intervention versus control comparison for the OR rather than the RR, then the bivariate method would simply be applied to the minority and non-minority count (rather than the minority and total count). In this case, one obtains ORs instead of RRs under the NB distributional assumption.
There are limitations to the current approach. Primarily, it is only applicable when individual-level covariates are not important to the overall study hypotheses, as the method only accommodates cluster-level (and not individual-level) covariates. Furthermore, as with any method, it should be applied in the context of a sensitivity analysis, for example, in comparison to per protocol analysis of the data. When inference is similar for per protocol, imputation, and intent-to-treat GEE approaches, the authors recommend intent-to-treat GEE be reported in conjunction with those analyses.
Footnotes
Acknowledgements
The authors acknowledge the entire Randomized Recruitment Intervention Trial Study team.
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by the National Institutes of Health award NIH/NIMHD (Grant number U24MD006941).
Trial registry
ClinicalTrials.gov Identifier: NCT01911208
Supplemental material
Supplemental material for this article is available online.
