Abstract
Stepped wedge cluster randomized trials are often analysed using linear mixed effects models that may include random effects for cluster, time and/or treatment. We investigate the impact of misspecification of the random effects structure of the model. Specifically, we considered two cases of misspecification of the random effects in a cross-sectional stepped wedge cluster randomized trials model – fit a linear mixed effects model with random time effects but the true model includes random treatment effects (case 1) or fit a linear mixed effects model with random treatment effect but the true model includes random time effects (case 2) – and derived the variance of the estimated treatment effect under misspecification. We defined two measures of the effect of misspecification: validity and efficiency. Validity is the ratio of the model-based variance of the treatment effect from the mis-specified model divided by the true variance of the treatment effect from the mis-specified model (based on a sandwich estimate of the variance). Efficiency is the ratio of the model-based variance of the treatment effect from the correctly specified model divided by the true variance of the treatment effect from the mis-specified model. We found that validity is less than 1.0 (anti-conservative) in almost all situations investigated with the exception of case 1 with two sequences, when validity could be greater than 1.0. Efficiency is less than 1 in all cases and depends on the intracluster correlation coefficient, the relative magnitude of the variance of the misclassified variance component, and the number of sequences. In general, there is no universal recommendation as to the most robust approach except for the case of a classic stepped wedge cluster randomized trial with only 2 sequences, where fitting a random time model is less likely to lead to anti-conservative inference compared with fitting a random intervention model.
Introduction
Cluster randomized stepped wedge designs implement an intervention across participating clusters in phases (Figure 1). Following a period of baseline data collection, clusters are randomly assigned to ‘sequences’ (with 1 or more clusters to a sequence) and the intervention is implemented at a different time in each sequence (sequence 1 clusters implement the intervention at time 1, sequence 2 clusters at time 2, etc.). Typically, once the intervention is implemented, it remains ‘on’. The wedge-shaped appearance of the staggered intervention start times across the clusters gives the design its name.1,2

Stepped wedge design with 3 sequences, 4 clusters per sequence (dashed lines within cluster) and 4 time periods. Shaded cells correspond to the intervention period.
A consequence of this design is that the intervention effect in a stepped wedge trial (SWT) is partly confounded with time – early on in the trial, most clusters remain in the standard-of-care condition, and later on, most clusters are in the intervention condition. Thus, any analysis of the intervention effect must adjust for time in some fashion. In addition, the clustered nature of the design (including multiple observations over time within a given cluster and, possibly, repeated observations on individuals) requires careful consideration of the correlation structure of the observations. Often, adjustments for time and correlation are modelled using a linear mixed effects modelling framework.3–6 In this presentation, we investigate the impact of misspecification of the random effects structure of the model.
Mixed effects models for the analysis of SWTs
We consider cross-sectional SWT designs with M unique treatment sequences, each observed over J time points. We assume that each sequence contains N clusters and each cluster contains K (different) individuals at each point in time. We assume that each individual is observed only once.
Let
where
In addition to a random cluster effect, denoted by
Model misspecification in SWT
We considered two cases of misspecification of the random effects in an SWT model. In the first case, the researcher fits the random time effect model, but the true model is the random treatment effect model (case 1). In the second case, the researcher chooses the random treatment effect model, but the true model is the random time effect model (case 2).
In general, the goal of an SWT is to estimate the treatment effect. Specifically, one would like an (asymptotically, at least) unbiased estimate of the treatment effect and a correct estimate of the sampling variance of the estimated treatment effect. In addition, one would prefer an estimate of the treatment effect with the smallest possible variance to increase precision and power. When the model is correctly specified, the maximum likelihood estimator satisfies these criteria. However, under a mis-specified model the maximum likelihood estimator of the parameters converges to the value
where
To evaluate the variance of the treatment effect estimate under misspecification according to the two criteria noted above (correct variance and minimum variance), we defined two measures of the effect of misspecification: validity and efficiency. Validity is defined as the ratio of the model-based variance of the treatment effect from the mis-specified model divided by the true variance of the treatment effect from the mis-specified model (i.e. the variance of the treatment effect estimate over repeated sampling when the data are generated from the true model and the treatment effect is estimated using the mis-specified model). The latter is based on a sandwich estimate of the variance, which provides a consistent estimate of the true variance of the estimated treatment effect under misspecification of the variance components. 10 That is
where
Drawing on the concept of asymptotic relative efficiency, efficiency is defined as the ratio of the model-based variance of the treatment effect from the correctly specified model divided by the true variance of the treatment effect from the mis-specified model. That is
where
We summarize our findings in Table 1. Further details are available in Voldal et al. 11 In case 1 (true model is random intervention but random time is fit), validity can be greater than (2 sequences) or less than (>2 sequences) 1.0 depending on the number of sequences. Validity deviates further from 1.0 as the ACC increases, as K increases and usually as the proportion of the cluster variance attributable to the mis-specified variance component increases. In case 2 (true model is random time but random intervention is fit), validity is less than 1.0 in all cases examined. Again, validity deviates further from 1.0 as the ACC increases, as K increases and as the proportion of the variance attributable to the mis-specified variance component increases. Thus, for classic SWT designs (except for the special case of 2 sequences), there is no modelling approach that will guarantee nominal or conservative inference in the face of random effect misspecification. In addition, some of these trends are sensitive to details of the SWT design, so do not hold universally.
Summary of findings for the classic stepped wedge design.
ACC: average cluster correlation.
As noted above, efficiency is always less than 1.0 (less efficient) under model misspecification. Generally, efficiency decreases (resulting in lower power/precision) with increasing ACC and as the proportion of the cluster variance attributable to the mis-specified variance component increases. In case 1, designs with more sequences generally (but not always) have less efficiency loss than designs with fewer sequences. In case 2, number of sequences has less impact on efficiency.
Conclusion
The analytic framework developed here allows one to quickly calculate the validity and efficiency of treatment effect estimates under different scenarios of misspecification without simulations. This can help a researcher understand the risks of choosing one model over another, particularly when convergence issues prevent fitting a model that includes all possible random effects.
In general, there is no universal recommendation as to the most robust approach, although one model may be more robust than the other for particular SWT designs (e.g. in a classic 2-sequence design, fitting a random time model is less likely to lead to anti-conservative inference compared with fitting a random intervention model; note that this may not be true for non-classic designs, however). 11 Our results also suggest that using robust variances for inference is recommended when possible.
Footnotes
Acknowledgements
This is a summary of a presentation made at the University of Pennsylvania’s 13th Conference on Statistical Issues in Clinical Trials – Cluster Randomized Clinical Trials: Opportunities and Challenges. The material is authored by the current authors and accepted in Statistics in Medicine. 11
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship and/or publication of this article: Research reported in this paper was supported in part by the National Institutes of Health under award number AI29168. The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health.
