Abstract
In cluster randomized trials, the intraclass correlation coefficient (ICC) is classically used to measure clustering. When the outcome is binary, the ICC is known to be associated with the prevalence of the outcome. This association challenges its interpretation and can be problematic for sample size calculation. To overcome these situations, Crespi et al. extended a coefficient named R, initially proposed by Rosner for ophthalmologic data, to cluster randomized trials. Crespi et al. asserted that R may be less influenced by the outcome prevalence than is the ICC, although the authors provided only empirical data to support their assertion. They also asserted that “the traditional ICC approach to sample size determination tends to overpower studies under many scenarios, calling for more clusters than truly required”, although they did not consider empirical power. The aim of this study was to investigate whether R could indeed be considered independent of the outcome prevalence. We also considered whether sample size calculation should be better based on the R coefficient or the ICC. Considering the particular case of 2 individuals per cluster, we theoretically demonstrated that R is not symmetrical around the 0.5 prevalence value. This in itself demonstrates the dependence of R on prevalence. We also conducted a simulation study to explore the case of both fixed and variable cluster sizes greater than 2. This simulation study demonstrated that R decreases when prevalence increases from 0 to 1. Both the analytical and simulation results demonstrate that R depends on the outcome prevalence. In terms of sample size calculation, we showed that an approach based on the ICC is preferable to an approach based on the R coefficient because with the former, the empirical power is closer to the nominal one. Hence, the R coefficient does not outperform the ICC for binary outcomes because it does not offer any advantage over the ICC.
1 Introduction
Cluster randomized trials are increasingly being used in health research. In such a setting, clusters of individuals are randomly allocated to different arms. 1 Clusters may be families, schools, worksites, medical practices, towns or other social units. In such trials, outcomes for individuals from the same cluster are more similar than are outcomes for individuals from different clusters.
The intraclass correlation coefficient (ICC) is classically used to measure this resemblance. It can be defined as the proportion of total variance due to between-cluster variation or the correlation between any two members of the same cluster.1,2 An ICC equal to 0 indicates independence among individuals of a cluster, whereas an ICC equal to 1 indicates that individuals from a given cluster have identical outcomes.
The Consolidated Standards for Reporting of Trials (CONSORT) extension for cluster randomized trials recommends reporting a measure of intracluster correlation, such as the ICC, for each primary outcome. 3 This has actually two aims. The first is that it may help interpret the results of the trial. Indeed, the assessed intervention may affect the level of clustering, and this result is important for a complete interpretation of trial result. When the outcome is binary, the ICC is known to be associated with the prevalence of the outcome. 4 As the prevalence increases from 0 to 0.5, the ICC increases. Because of this association, ICC values are expected to differ when prevalences differ, even when the clustering level remains identical. This association with prevalence challenges the interpretation of the ICC because ICC values do not just depend on clustering level.
The second reason for providing clustering estimates is that such values are of help for sample size calculation of future studies. Yet, the association between the ICC and the outcome prevalence can be problematic in sample size calculation if the study to be planned is expected to have prevalences different from those from which we derived ICC estimates.
To overcome these situations, Crespi et al. 5 extended a coefficient named R, initially proposed by Rosner 6 for ophthalmologic data, to cluster randomized trials. R is defined as a ratio for which the numerator is the conditional probability that a member of a cluster has the outcome given that another member of the cluster also has the outcome, and the denominator is the outcome prevalence. Crespi et al. asserted that R may be less influenced by the outcome prevalence than the ICC. To support this assertion, the authors provided an illustration with an example, stating that mathematical proof was not possible. Moreover, they proposed sample size formulas using R coefficients or ICCs and used these formulas to calculate required sample size in diverse situations. They concluded that “the traditional ICC approach to sample size determination tends to overpower studies under many scenarios, calling for more clusters than truly required”. However, this latter conclusion was based on sample size calculations without any consideration of empirical power.
The aim of this study was to investigate whether R is indeed independent of the outcome prevalence. We also investigated which sample size calculation, based on the R coefficient or the ICC, provides the empirical power closest to the nominal power.
We define R in section 2 and provide its estimator in section 3. In section 4.1, we explore theoretically the relation between R and the outcome prevalence in the special case of clusters of size 2. In section 4.2, we report a simulation study to explore the situation of cluster sizes greater than 2, both fixed and variable. An illustration using real data is provided in section 5. In section 6.1, we compare sample size calculation using R or the ICC in another simulation study and in section 6.2, we illustrate the asymmetry of R in sample size calculation. We conclude with a short discussion in section 7.
2 Definitions
In this section up to and including section 4, we will consider one arm composed of k clusters of size
2.1 R as defined by Rosner
Rosner
6
worked on methods for analysing ophthalmologic data. In this special case, the cluster unit is the individual, with two observations (eyes) per individual. Rosner defined R as
2.2 Crespi’s extension of R
Crespi et al. extended the formula from Rosner to the case of clusters of fixed size (m) potentially greater than 2
5
R can be seen as a quantification of how much or less likely a member of a cluster is to be successful given that another member of the cluster is successful.
3 Estimating R
3.1 The Rosner R estimator:
In the special case of clusters of fixed size m = 2, Rosner
6
showed that the maximum likelihood estimator of R is
3.2 The Crespi R estimator:
Under the common correlation model, the ICC has been defined as
7
Therefore, from equation (2) we derive that
Crespi et al. suggested that R be estimated by using
Many methods have been proposed to estimate the ICC for binary outcomes. 8 Simulation results reported by Ridout et al. showed that the ANOVA estimator, the Fleiss-Cuzick estimator and some of the moment estimators performed well in terms of bias, standard deviation and mean square error.
When using the Fleiss-Cuzick estimator and considering clusters of size 2, the R estimator using the approach proposed by Crespi et al. is equivalent to its maximum likelihood estimator (3) as proposed by Rosner. Therefore, we used the Fleiss-Cuzick estimator defined as
4 Relation between R and outcome prevalence
4.1 Exploration of the symmetry of the R estimator around a prevalence of 0.5
In this section, we consider a dataset e containing ke clusters, each with two observations of a binary outcome. The estimated prevalence of success is
Let us now consider that we are interested in measuring the clustering for failure rather than success; thus, all 0 values are replaced by 1 and vice versa. The associated estimate prevalence of failure is
It can be shown that
For example, if we consider the situation of 100 clusters with
Moreover, this asymmetry is counterintuitive because we would expect that the degree of intracluster resemblance to be the same whether we consider the resemblance in success or failure for a given dataset.
4.2 Simulation study
In the previous section, we showed that R is associated with outcome prevalence, for clusters of size 2. We then investigated the shape of the relation between R and prevalence. We considered the most general case of cluster sizes greater than 2, both fixed and variable. To this end, we conducted a simulation study according to the following principle. We generated correlated binary data with pre-specified outcome prevalence p and intraclass correlation
We first specified
4.2.1 Simulation plan
Steps of the data generation for each pair For the following cluster sizes:
Variable cluster sizes: simulate ni cluster sizes, Fixed cluster sizes: set For each cluster, simulate For each individual, simulate For each individual, simulate Calculate Xij according to equation (8).
We varied p between 0.01 and 0.99. Statistical analyses were conducted with the three following steps:
Estimate p as Estimate Calculate
All negative values of
We generated 50,000 datasets for each scenario and for each value of p, we summarized results by computing
Simulations were run considering three initial values of
4.2.2 Simulation results
Figure 1 displays three plots

Theoretical intraclass correlation coefficient (ICC)
5 Example
To empirically illustrate the relation between R and the success prevalence, we used data from the Health Services Research Unit in Aberdeen (https://www.abdn.ac.uk/hsru/what-we-do/tools/index.php#panel177). These data provide estimates of ICCs from changing professional practice studies. Clusters were hospitals, hospital units, hospital directorates, general practices, physicians or pharmacies.
We used 145 ICCs from binary outcomes and associated prevalence values. ICCs ranged from 0 to 0.659 (median 0.057, interquartile range [IQR] 0.012–0.105). Prevalence values ranged from 0.032 to 0.995 (median 0.452, IQR 0.209–0.819). We estimated R by using equation (5). R values ranged from 1 to 4.217 (median 1.084, IQR 1.006–1.243).
We plotted R on the logarithm scale as a function of prevalence (Figure 2). The strength of the association between R and prevalence was estimated by the Spearman correlation coefficient. The estimated correlation coefficient was −0.721 (95% CI −0.833 to −0.580,

Association between R and prevalence by using data from the Health Technology Assessment review. The Spearman correlation coefficient was estimated at −0.721 (95% CI −0.833 to −0.580).
6 Sample size considerations using R
6.1 Empirical power in sample size calculation when using R
Crespi et al. 5 proposed sample size formulas with the R coefficient or ICC and used them to calculate required sample sizes in diverse situations. The authors considered three approaches: (1) one based on two R coefficients, that is, one for each arm (R-based approach); (2) one based on two ICCs (ICC A approach) and (3) one based on an ICC assumed to be common to the two arms (ICC B approach). They also considered two situations: (1) the prevalence levels of the study to be planned differ from those of the study previously conducted and from which R coefficients and ICCs have been estimated and (2) the prevalence levels are identical. The required number of clusters differed according to the approach used for sample size calculation. Focusing on the first situation (i.e. change in prevalence level between the previously conducted study and the planned one), the authors concluded that: (1) “when moving from high to moderate or from moderate to low prevalence, the ICC approaches can grossly overpower the study, calling for many more clusters than required to achieve desired power” and that (2) “when moving to a higher prevalence setting, ICC A underpowers the study”. However, the approach used by Crespi et al. is debatable. Indeed, they derived three required sample sizes using the three previously cited approaches. Then, using the sample size formula based on two R coefficients, they derived power associated with the three sample sizes previously calculated. As a consequence, they observed a power close to 80% for the approach based on two R coefficients (slight deviations from 80% are due to rounding of the number of clusters). For the two other approaches, power was greater than 80% because the required number of clusters was higher when using an approach based on ICCs than the approach based on the R coefficient. However, doing so does not demonstrate anything, because Crespi et al. did not check whether the empirical power actually equals the nominal power. Therefore, we investigated whether Crespi et al.’s assertions were correct, estimating the empirical power associated with each situation they considered. Indeed, claiming that approach A is overpowered only because it requires a larger sample size than approach B is not correct. This is true only if the empirical power associated with approach B equals the nominal one, which was not verified by Crespi et al. Therefore, we performed a simulation study to assess which approach is preferable.
6.1.1 Simulation plan
Let us consider that a two-arm cluster randomized trial has already been previously conducted and prevalences were p1 and p2 in arm 1 and 2, respectively, with associated R1 and R2 coefficients. Given that p1, p2, R1 and R2, ρ1 and ρ2 can be derived by using equation (5). In practice, p1, p2, R1, R2, ρ1 and ρ2 can be replaced by their associated estimates, and this holds true for the upcoming issues. We plan to conduct a new study with expected prevalences R-based approach
ICC A approach
ICC B approach
with m the fixed cluster sizes, α and β the type I and type II error, respectively;
Of note, for the latter formula we computed 3. We estimated the empirical power associated with each sample size. For this, we used the same simulation plan as that described in section 4.2. Considering p1 and ρ1, we derived, using formula (7), an estimate of
Simulations were run considering the same scenarios as Crespi et al., when clusters are of size m = 20. We considered the situation in which prevalences in the future study are expected to be different from those of the previously conducted study (from
All programming was implemented by using R software, v3.6.1. Code is available on https://github.com/Mbekwe/Simulation-study-with-R-software.git.
6.1.2 Simulation results
When prevalences in the previously conducted study were high and those in the future one were expected to be moderate (first row of Figure 3), the sample size computed by using the R-based approach did not reach the theoretical power of 80%. The empirical power was always largely lower than 80%. Conversely, when using ICC-based approaches, the empirical power was close to 80%, even closer when we considered two ICCs rather than a common one. When we moved from moderate to low prevalence (second row of Figure 3), the R-based approach still led to fewer clusters than necessary. When using ICC-based approaches, this led to more clusters than necessary but with an empirical power closer to 80% than with the R-based approach. When we moved from low to moderate prevalence (third row of Figure 3) or from moderate to high prevalence (fourth row of Figure 3), the sample size computed using the R-based approach was greater than necessary: the empirical power was always greater than 80%. Conversely, using ICC-based approaches still led to empirical power close to 80%. Use of the R-based approach was under-powered (when moving from high to moderate prevalence or from moderate to low prevalence) or over-powered (when moving from low to moderate prevalence or from moderate to high prevalence). In all cases, using the R-based approach was worse than using ICC-based approaches. In the no prevalence change setting (Figure 4), the R-based approach and ICC A approach were identical and their empirical power was similar to that of the ICC B approach. Thus, sample size calculation using an ICC approach performs better than using the R approach.

Comparison of three approaches (R-based, ICC A and ICC B) in empirical power estimation. These three approaches were used to compute sample size. A total of 5000 datasets were simulated for each combination of p1, p2,

Comparison of three approaches (R-based, ICC A and ICC B) in empirical power estimation. These three approaches were used to compute sample size. A total of 5000 datasets were simulated for each combination of p1, p2,
6.2 Asymmetry when using R in sample size calculation
Another drawback of the R approach is its asymmetry in sample size calculation. To illustrate this point, let us consider the situation presented in section 4.1. A previously conducted study in which successes were considered had a prevalence of 0.15, an ICC of 0.29, and a R coefficient of 2.64. Suppose we want to detect an increase of 10 percentage points in the proportion of success in a future study. To detect an increase in success rate from 0.15 to 0.25, with a power of 80% at the 5% level, 179 clusters for 358 individuals would be needed for each arm. Let us now focus on failures rather than successes. The previously conducted study had a failure rate of 0.85, an ICC of 0.29 and a R coefficient of 1.05. If we now want to detect a decrease from 0.85 to 0.75 in the failure rate (equivalent to the increase in success), with a power of 80% at the 5% level, 149 clusters for 298 individuals would be needed for each arm. Therefore, using the R coefficient would lead to two different required sample sizes, although intuitively, they should be equal.
7 Discussion
In this paper, we have described the R coefficient and explored its association with the outcome prevalence by using mathematical developments for fixed cluster sizes of 2 and using simulations for the most general cases of fixed and variable cluster sizes greater than 2.
We show that the R coefficient decreases with increasing prevalence, so R depends on the outcome prevalence.
Furthermore, R is not symmetrical around a prevalence of 0.5. Thus, R performs even worse than the ICC to measure clustering in the sense that for a given dataset, the R value is not the same when we are interested in success or failure for the same variable. Consequently, R is not an appropriate coefficient if one wants an index independent of the outcome prevalence. Sample size calculation using an approach based on the ICC appears to be preferable because the empirical power is closer to the nominal one versus an approach based on the R coefficient. Moreover, when using R, the sample size differs depending on whether we focus on success or failure. Even if both R and ICCs have limits, our results encourage the use of the ICC over the R coefficient for binary outcomes. Further work is needed to explore or develop other measures, notably the tetrachoric correlation coefficient, which can be used to quantify clustering without being influenced by the outcome prevalence, thus allowing a direct comparison of clustering for outcomes with different prevalences.
Supplemental Material
SMM900200 Supplemental material - Supplemental material for Is the R coefficient of interest in cluster randomized trials with a binary outcome?
Supplemental material, SMM900200 Supplemental material for Is the R coefficient of interest in cluster randomized trials with a binary outcome? by Ariane M Mbekwe Yepnang, Agnès Caille, Sandra M Eldridge and Bruno Giraudeau in Statistical Methods in Medical Research
Footnotes
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
Supplemental material
Supplemental material for this article is available online.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
