Abstract
Background/aims
Cluster randomized trials are popular in health-related research due to the need or desire to randomize clusters of subjects to different trial arms as opposed to randomizing each subject individually. As outcomes from subjects within the same cluster tend to be more alike than outcomes from subjects within other clusters, an exchangeable correlation arises that is measured via the intra-cluster correlation coefficient. Intra-cluster correlation coefficient estimation is especially important due to the increasing awareness of the need to publish such values from studies in order to help guide the design of future cluster randomized trials. Therefore, numerous methods have been proposed to accurately estimate the intra-cluster correlation coefficient, with much attention given to binary outcomes. As marginal models are often of interest, we focus on intra-cluster correlation coefficient estimation in the context of fitting such a model with binary outcomes using generalized estimating equations. Traditionally, intra-cluster correlation coefficient estimation with generalized estimating equations has been based on the method of moments, although such estimators can be negatively biased. Furthermore, alternative estimators that work well, such as the analysis of variance estimator, are not as readily applicable in the context of practical data analyses with generalized estimating equations. Therefore, in this article we assess, in terms of bias, the readily available residual pseudo-likelihood approach to intra-cluster correlation coefficient estimation with the GLIMMIX procedure of SAS (SAS Institute, Cary, NC). Furthermore, we study a possible corresponding approach to confidence interval construction for the intra-cluster correlation coefficient.
Methods
We utilize a simulation study and application example to assess bias in intra-cluster correlation coefficient estimates obtained from GLIMMIX using residual pseudo-likelihood. This estimator is contrasted with method of moments and analysis of variance estimators which are standards of comparison. The approach to confidence interval construction is assessed by examining coverage probabilities.
Results
Overall, the residual pseudo-likelihood estimator performs very well. It has considerably less bias than moment estimators, which are its competitor for general generalized estimating equation–based analyses, and therefore, it is a major improvement in practice. Furthermore, it works almost as well as analysis of variance estimators when they are applicable. Confidence intervals have near-nominal coverage when the intra-cluster correlation coefficient estimate has negligible bias.
Conclusion
Our results show that the residual pseudo-likelihood estimator is a good option for intra-cluster correlation coefficient estimation when conducting a generalized estimating equation–based analysis of binary outcome data arising from cluster randomized trials. The estimator is practical in that it is simply a result from fitting a marginal model with GLIMMIX, and a confidence interval can be easily obtained. An additional advantage is that, unlike most other options for performing generalized estimating equation–based analyses, GLIMMIX provides analysts the option to utilize small-sample adjustments that ensure valid inference.
Keywords
Introduction
Increasingly in health-related research, cluster randomized trials, or group randomized trials, are used to randomize entire clusters, or groups, of subjects to trial arms as opposed to individually randomizing subjects.1–6 Such trials are conducted, for instance, because of the need or desire to apply an intervention to the entire cluster due to natural grouping or to avoid contamination. Popular examples of possible clusters include, but are not limited to, communities, schools, and medical care practices.
The number of clusters and their sizes depend on the setting. For instance, Van Breukelen and Candel 7 summarize results from reviews on primary care cluster randomized trials reported by Adams et al. 8 and Eldridge et al., 9 with the median number of clusters ranging from 31 to 41, and a median of the mean number of subjects per cluster equaling 23 8 or 32. 9 Alternatively, in many cluster randomized trials, the number of clusters is small, whereas the size of each cluster can be large. For example, Roetzheim et al.10,11 and Lee et al. 12 report results from a repeated cross-sectional cluster randomized trial with focus on three cancer screening outcomes, which were subject-level binary outcomes of whether or not a given subject had received a mammography, a papanicolaou (pap) smear, or a fecal occult blood test. The intervention and control conditions were only randomized to four clinics each, and data were obtained from 51 to 127 subjects per clinic, depending on the time period and outcome type. 12
In this article, we focus on situations in which subjects contribute binary outcomes; for example, an indicator for having been screened for a particular type of cancer. Because outcomes from subjects within the same cluster tend to be more alike than outcomes from subjects within other clusters, for example, due to differences in clinical practices, an exchangeable correlation arises that is measured via the intra-cluster, or intra-class, correlation coefficient (ICC). For more detail, a review on different definitions for the ICC can be found in Eldridge et al. 13 In this article, we focus on ICCs corresponding to the proportions scale.
The ICC is important because it must be appropriately accounted for in the design of cluster randomized trials and during the statistical analysis of data resulting from such a trial. According to the extension of the Consolidated Standards of Reporting Trials (CONSORT) statement to cluster randomized trials, 14 ICC estimates should be reported, which are very helpful for planning future cluster randomized trials. 6 For example, Hade et al. 15 reported 138 unadjusted and adjusted ICC estimates, from 12 different studies, corresponding to different types of cancer screening, and Taljaard et al. 16 reported 86 estimated ICC values for variables corresponding to maternal and perinatal health. In addition, there has been an increase in reporting of ICCs in primary care settings. 17 Furthermore, databases of ICC values have even been created for certain types of studies; for example, a surgical-based ICC database.6,18
Due to the importance of reporting reliably estimated ICC values, numerous ICC estimation methods have been developed and studied against each other. For summaries, see, for instance, the references cited here.6,19,20 Although multiple estimators have been shown to work well, such as the well-known ANOVA estimator19–23 which was used by Hade et al., 15 to our knowledge many of these estimators cannot be used in the context of practical data analyses in which marginal models, fit using generalized estimating equations (GEE), 24 do or do not adjust for covariates.
With GEE, the exchangeable correlation, or the ICC in our context, has traditionally been estimated using the method of moments. This is the case, for instance, in the popular R package geepack25–28 as well as SAS GENMOD. 29 However, this approach is well known to produce negatively biased ICC estimates, as demonstrated in, for instance, Wu et al. 20 and later in our simulation study. The resulting problem is that the use of such negatively biased estimates to plan future cluster randomized trials could result in the calculation of a sample size (number of clusters or the sizes of the clusters) that is too small or alternatively an over-calculation of power for a fixed number of clusters and their sizes.
Due to the limitations of well-known and routinely used estimators, in this article we demonstrate the ability of SAS GLIMMIX 29 to estimate the ICC, on the proportions scale, in a GEE-based analysis. In short, residual pseudo-likelihood is used via SAS GLIMMIX to estimate the ICC as opposed to the method of moments. Although our main focus in this article is demonstrating the ability of this procedure to estimate the ICC, our desire to demonstrate how well GLIMMIX works in this regard is also due to this procedure’s ability to implement small-sample bias corrections to empirical standard error estimates and the allowance of user-defined degrees of freedom in order to ensure valid inference. We discuss this issue in detail later in the “Discussion” section and in an online supplementary appendix.
We note that the construction of confidence intervals for the ICC has received a considerable amount of attention in the cluster randomized trial literature. For example, see Chakrabory and Hossain 30 and Braschel et al. 31 Methods for confidence interval construction have been established for the ANOVA and methods of moments estimators,30,32 but not for the residual pseudo-likelihood estimator. Therefore, we study a general confidence interval construction approach of Demetrashvili et al. 33 that is naturally applicable for use when estimating the ICC with residual pseudo-likelihood.
Next, we briefly give details on GEE-based analyses and discuss ICC estimation with the method of moments and residual pseudo-likelihood. Furthermore, as a standard of comparison, we also describe ANOVA estimators due to their popularity, their proven ability to estimate the ICC well, as well as the capability to use this type of estimator even when adjusting for cluster-level covariates in the marginal model. Via a simulation study, we then demonstrate the utility of the residual pseudo-likelihood method of GLIMMIX in terms of estimating the ICC relative to ANOVA-based and method of moment estimators. Furthermore, we study the confidence interval construction approach of Demetrashvili et al. 33 to determine how well it works when the ICC is estimated with residual pseudo-likelihood. The ICC estimation approaches are then contrasted via application to analyses of data from the cancer screening study of Roetzheim et al.10,11 Finally, we provide concluding remarks, including a discussion on the additional practical advantages of using GLIMMIX.
Methods
Notation
For simplicity, we assume two trial arms, each consisting of
The working covariance structure of
Marginal modeling via GEE
To fit a marginal model utilizing GEE, the estimate for
in which
in which
ICC estimation
With our focus on GEE, the ICC corresponds to a common exchangeable correlation parameter which, in general, is given by
The traditional approach to obtaining
where
Due to the popularity and utility of the ANOVA estimator,19–23 as well as the possibility to use it with GEE when no subject-level covariates are incorporated within the marginal model, we later utilize it as a standard of comparison in our simulation study. When a trial arm indicator is the only covariate in the marginal model, the estimate for
Let
When other cluster-level covariates are included in the marginal model, an alternative formula can be derived based on the ANOVA method 38 that still has the form given in equation (2), although formulas differ as follows
and
In this article, we propose the use of residual pseudo-likelihood via the GLIMMIX procedure in SAS.
29
In short, Taylor series expansions are implemented such that pseudo, or transformed, response data are created to be used in a weighted linear mixed model, and therefore SAS documentation refers to the corresponding marginal model as a “GEE-type” model.
29
Due to the use of this weighting, the estimated compound symmetry and residual covariance parameters,
As most cluster randomized trials do not involve a large number of clusters, we suggest the use of the default restricted maximum likelihood 42 for covariance parameter estimation, that is, estimation of the ICC. Details on residual pseudo-likelihood are provided in the online supplementary appendix, and in-depth information is provided by Fitzmaurice et al. 42 and SAS’s User’s Guide. 29
We note that methods for constructing confidence intervals corresponding to ICC estimates from the residual pseudo-likelihood approach are not established. However, Demetrashvili et al. 33 proposed an approach to obtain confidence intervals with approximately correct coverage in the context of variance component models. Therefore, due to the use of restricted maximum likelihood, and based on the form for the ICC estimate given in equation (3), the approach of Demetrashvili et al. 33 naturally can be applied when utilizing residual pseudo-likelihood to estimate the ICC. For details, we refer the reader to Demetrashvili et al. 33 and the SAS code that is available in Supplemental Material.
Simulation study
We conduct a simulation study to demonstrate how well the residual pseudo-likelihood approach works relative to the method of moments and ANOVA methods in terms of estimating the ICC. Specifically, we compare empirical biases. We note that for each of these methods, if a negative ICC estimate is produced, we use an ICC value of 0 as is often done in practice. 6 Furthermore, we assess the approach of Demetrashvili et al. 33 to constructing a confidence interval for the ICC based on the residual pseudo-likelihood estimation method. Specifically, we compare empirical coverage probabilities of these 95% confidence intervals with the nominal 0.950 value.
For each setting, 10,000 datasets are generated from either an unadjusted or adjusted marginal logistic regression model. The unadjusted model is given by logit
Empirical biases of estimated intra-cluster correlation coefficients (ICCs) from 10,000 datasets generated for each setting (row) from the unadjusted marginal logistic regression model with given coefficient of variation (CV) of cluster sizes (mean size is 100), true ICC values, number of clusters per trial arm
Empirical coverage probabilities (CPs) of 95% confidence intervals (CIs) corresponding to the residual pseudo-likelihood estimator and construction approach of Demetrashvili et al. 33 are also presented in the last column.
ANOVA 1: The traditional ANOVA estimator that cannot account for covariate adjustment.
ANOVA 2: The ANOVA estimator that can account for covariate adjustment and utilizes leverage values.
MM Common: The common method of moments estimator given in equation (1).
MM Corrected: The corrected method of moments estimator of Lu et al. 36
RPL: residual pseudo-likelihood from GLIMMIX.
Empirical biases of estimated intra-cluster correlation coefficients (ICCs) from 10,000 datasets generated for each setting (row) from the unadjusted marginal logistic regression model with given coefficient of variation (CV) of cluster sizes (mean size is 30), true ICC values, number of clusters per trial arm
Empirical coverage probabilities (CPs) of 95% confidence intervals (CIs) corresponding to the residual pseudo-likelihood estimator and construction approach of Demetrashvili et al. 33 are also presented in the last column.
ANOVA 1: The traditional ANOVA estimator that cannot account for covariate adjustment.
ANOVA 2: The ANOVA estimator that can account for covariate adjustment and utilizes leverage values.
MM Common: The common method of moments estimator given in equation (1).
MM Corrected: The corrected method of moments estimator of Lu et al. 36
RPL: Residual pseudo-likelihood from GLIMMIX.
Empirical biases of estimated intra-cluster correlation coefficients (ICCs) from 10,000 datasets generated for each setting (row) from the adjusted marginal logistic regression model with given mean cluster size
Empirical coverage probabilities (CPs) of 95% confidence intervals (CIs) corresponding to the residual pseudo-likelihood estimator and construction approach of Demetrashvili et al. 33 are also presented in the last column.
ANOVA 2: The ANOVA estimator that can account for covariate adjustment and utilizes leverage values.
MM Common: The common method of moments estimator given in equation (1).
MM Corrected: The corrected method of moments estimator of Lu et al. 36
RPL: Residual pseudo-likelihood from GLIMMIX.
We only present results from the unadjusted model settings in which
Similar to our application example and the numbers of clusters and their sizes reported in Adams et al. 8 and Eldridge et al., 9 we focus on general cluster randomized trial settings with either a small number of large clusters or a medium number of smaller clusters. Specifically, the former (latter) settings comprise either 5 or 10 (15 or 20) clusters in each trial arm, with a mean cluster size of 100 (30). Due to computational time, we do not study larger clusters or a larger number of clusters. However, we do note that as the number of clusters increases, the degree of bias in the ICC estimation approaches decreases. We utilize true ICC values of 0.01 and 0.10, as most estimated ICCs reported from such cluster randomized trials tend to not be large.6,8,9,17 Furthermore, the coefficients of variation of cluster sizes were 0 and 1, with sizes generated from a normal distribution and then rounded to the nearest integer.44,45 More details are provided in the online supplementary appendix.
Data are generated from the given marginal models via the use of a beta-binomial distribution. R version 3.4.3 25 and SAS 29 are used to conduct the study. Details on data generation and software use are provided in the online supplementary appendix.
Application
To illustrate the different ICC estimation methods assessed in our simulation study, we analyze data on females from the Roetzheim et al.10,11 and Lee et al. 12 cluster randomized trial. As a reminder, there are eight total clinics, or clusters, such that four are in each of the control and intervention arms, and three different outcomes are subject-level binary indicators of whether or not a given subject had received a mammography, a papanicolaou smear, or a fecal occult blood test. We utilize data from baseline and 2 years post baseline. Although this study utilizes a repeated cross-sectional design, some subjects are observed at both points in time simply by chance. The number of subjects in each clinic contributing data with respect to a papanicolaou smear, mammography, and fecal occult blood test at 2 years post baseline ranges from 66 to 85, 111 to 122, and 112 to 126, respectively. The proportion of women receiving a papanicolaou smear, a mammography, and a fecal occult blood test ranges from 0.38 to 0.55, 0.50 to 0.71, and 0.04 to 0.22, respectively, in the control arm and 0.15 to 0.71, 0.45 to 0.77, and 0.14 to 0.37, respectively, in the intervention arm.
For each outcome, we fit two marginal logistic regression models:
Unadjusted: logit
Adjusted: logit
Within these models,
Results
Simulation study
As expected, the ANOVA estimators perform very well overall, as they often provide unbiased ICC estimates or estimates with little bias. Furthermore, both types of moment estimators typically provide negatively biased ICC estimates. Overall, the residual pseudo-likelihood estimator performs very well. It notably outperforms the moment estimators that are traditionally used with GEE analyses in practice. Furthermore, it tends to work almost as well as the ANOVA estimators. When the true ICC value is 0.01, the residual pseudo-likelihood estimator yields unbiased results with the exception in some settings of the slight simulated bias of only 0.001. Although there are a few settings in which small negative bias exists with the residual pseudo-likelihood estimator when the true ICC value is 0.10, this bias is not nearly as great as the bias in the moment estimators, and the bias in the ANOVA estimators is only slightly smaller. Furthermore, there are settings in which the residual pseudo-likelihood estimator performs best.
For all estimators, bias tends to decrease as the number of clusters increases. Also, the closer marginal probabilities are to the edge of the parameter space, that is, the further probabilities are from 0.5, the greater the potential bias in any of these estimators. We note that this bias is more notable with the larger true ICC value of 0.1. We also note that allowing marginal probabilities to vary across clusters, as opposed to having equal marginal probabilities across clusters, does not impact our findings.
The confidence interval construction approach of Demetrashvili et al.
33
tends to work well, but not perfectly, when applied with the residual pseudo-likelihood estimator. When bias is negligible, for example,
Application
ICC estimates are provided in Table 4, corresponding to each outcome type, model, and estimation method. In general, results are similar to findings in the simulation study. The method of moments estimators yield the smallest ICC values. Alternatively, the residual pseudo-likelihood estimates are similar to the estimates produced by the ANOVA methods. Based on the results from our simulation study, this is the result we are both expecting and hoping for.
Estimates of the intra-cluster correlation coefficient from analyses of the cancer screening data.
Unadjusted marginal logistic regression model: logit
Adjusted marginal logistic regression model: logit
Logit
ANOVA 1: The traditional ANOVA estimator that cannot account for covariate adjustment.
ANOVA 2: The ANOVA estimator that can account for covariate adjustment and utilizes leverage values.
MM Common: The common method of moments estimator given in equation (1).
MM Corrected: The corrected method of moments estimator of Lu et al. 36
RPL: Residual pseudo-likelihood from GLIMMIX.
FOBT: Fecal occult blood test.
Discussion
In this article, we demonstrate that ICC estimation via residual pseudo-likelihood using the GLIMMIX procedure in SAS works very well. Specifically, it results in either unbiased estimates or estimates with small bias. Furthermore, any bias is notably smaller in magnitude than the negative bias resulting from the method of moments approaches. Furthermore, residual pseudo-likelihood estimates are often comparable to ANOVA estimates.
The construction of confidence intervals for the ICC has also received a considerable amount of attention in the cluster randomized trial literature. Therefore, we studied the applicability of the construction approach of Demetrashvili et al. 33 in our context of estimating the ICC with residual pseudo-likelihood. In our simulation study, this approach is shown to often, but not always, produce near-nominal coverage probabilities as long as the bias in the residual pseudo-likelihood estimator is negligible. Therefore, this approach is recommended for use in practice, although it should be kept in mind that coverage may not be perfect. This is especially the case when notable bias exists with the estimator, although in practice it will not be known if any bias is present. We also note that this is a limitation with any ICC estimator and confidence interval construction approach.
In addition to providing better ICC estimates than other approaches that are commonly utilized with GEE in practice, GLIMMIX has additional practical advantages. First, user-defined degrees of freedom and small-sample corrections to empirical standard errors can be employed. Second, the ICC structure can be flexibly modeled. 42 A discussion on the importance of these advantages, with details, is provided in an online supplementary appendix and is based on the references cited here.20,32,34,36,38,40–45,47–58
In Supplemental Material, we provide SAS code to demonstrate how to conduct a GEE-type analysis with GLIMMIX. We focus on ICC estimation via residual pseudo-likelihood, and we provide code to calculate a confidence interval for the ICC based on the method of Demetrashvili et al. 33 Furthermore, we show how to utilize corrections to the empirical sandwich estimator and user-defined degrees of freedom. Furthermore, we demonstrate how to obtain a separate ICC estimate with corresponding confidence interval for each trial arm.
A limitation of estimating a separate ICC for each trial arm is that fewer clusters will be utilized to estimate each ICC, and therefore estimation will be less precise. Therefore, estimates for a common ICC as well as separate ICCs for each trial arm could be reported. Furthermore, both crude and adjusted estimates can be reported, as in Hade et al.6,15 We note that ICC estimates will often, but not always, be smaller when adjusting for covariates in the regression model.8,15,17,59 For example, this is seen with both pap smears and mammographies in our application example. Studies should also report the method for obtaining ICC estimates in order for researchers to judge if the ICC was estimated using a reliable approach; for example, an estimate from the method of moments will likely be negatively biased, whereas an estimate from the use of residual pseudo-likelihood will often be approximately unbiased. Finally, due to the potentially high variability in ICC estimates, approaches to utilize information from reported ICC estimates in order to design future cluster randomized trials have been developed. For instance, see Turner et al. 60
Although the use of generalized linear mixed models 42 is also popular, in this article, we focus on the use of GEE to fit a marginal model, as population-average interpretations from statistical models are more natural in cluster randomized trial settings in which clusters are only observed within one trial arm during the study.4,32,61 Furthermore, with GEE a correct specification of the distribution for the outcomes is not required. We point out that the ICCs for binary data can be on the proportions scale or the log odds scale if using a logit link function, and ICCs for different scales are not equivalent.13,17 Specifically, the ICC calculated from a GEE-type marginal model is on the proportions scale and is not equivalent to the ICC calculated from a mixed-effects logistic model. More detail is provided in the online supplementary appendix.
Supplemental Material
803635_supp_mat – Supplemental material for A readily available improvement over method of moments for intra-cluster correlation estimation in the context of cluster randomized trials and fitting a GEE–type marginal model for binary outcomes
Supplemental material, 803635_supp_mat for A readily available improvement over method of moments for intra-cluster correlation estimation in the context of cluster randomized trials and fitting a GEE–type marginal model for binary outcomes by Philip M Westgate in Clinical Trials
Footnotes
Acknowledgements
We thank an anonymous associate editor and two reviewers for their constructive comments that helped improve this paper.
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
