Abstract
Traditional methods based on the assumption that the treatment distribution is a pure shift of the control distribution may not always hold. The possibility that an individual from the treatment group may not respond to the treatment motivates the use of a mixture distribution for the treatment group. This paper considers two test procedures based on the Wilcoxon rank-sum statistic for a group sequential design to detect the one-sided mixture alternative. Error spending functions are used for the allocation of error rates at each stage. The two tests are evaluated individually in determination of critical values and arm sizes and asymptotic multivariate normality is shown to hold for both. Upon comparison, the tests are presented to be asymptotically equivalent. Both test statistics maintain the Type I error rate even if
Introduction
The setting in this article is that of testing the efficacy of a treatment versus a control. Clinical trials are a popular and effective design used to perform these tests while ensuring safety of the participants and general population. Group sequential clinical trials take these ideals a step further by incorporating the possibility of terminating the trial early by rejecting the null hypothesis due to strong evidence of efficacy or accepting the null hypothesis for futility. Instead of a one-time test to evaluate a treatment, group sequential methods conduct hypothesis tests at a predetermined number of stages. That is, a stage represents each time analysis is performed. The trial may terminate at an earlier stage than the final one if there is sufficient evidence to make a decision in the hypothesis test. One way of achieving the boundaries for such tests is through error spending methods. Lan and Demets 1 developed rejection boundaries based on an error spending function that controls the rate at which the Type I error is spent. Similar methods are developed by Chang et al. 2 to include an error spending function for the Type II error rate.
Traditional methods of evaluating a treatment versus control assume that the treatment effect should be represented by an additive shift of the control distribution. In this pure shift situation, all treated individuals have responses that come from the shifted distribution. However, this may not always be the case. Good 3 and Boos and Brownie 4 both show instances of “nonresponders” to a given treatment. For additional evidence of this phenomenon, see Jeske and Yao 5 and Peng et al. 6 We define nonresponders as individuals that have been given the treatment and were unaffected by it. Therefore, this treatment group will have responses that come from the same distribution as the control group responses. The proportion of nonresponders is considered unknown, and with their existence, the treatment distribution will be represented by a mixture of the control distribution and a location shift applied to the control distribution.
For the formulation of hypothesis tests, we define the distribution function of the control group as
Peng et al.
6
work in this context and develop a group sequential test based on the difference of means. They work under the assumption that the control and treatment group observations follow a normal distribution and a mixture of normal distributions, respectively, but their work can be used for other distribution functions if the central limit theorem is invoked. R code is available from their paper to set up the design for user-inputted choices of various parameters. While working in a fixed sample setting, Jeske and Yao
5
remove the normal theory assumption in favor of a less restrictive assumption of the control distribution being some continuous cumulative distribution function,
In the group sequential setting, there is no literature aimed at detecting the mixture alternative with a general
The rest of the article is organized as follows. In Section 2, we introduce our proposed test procedure including discussion of the test statistic, its distribution, the normal approximation of its distribution, and computation of arm size. Section 3 examines the power and robustness of the test procedure. Section 4 investigates a comparison of the test statistic used by Shuster et al. 8 and the test statistic, we introduce in Section 2. In Section 5, we explore estimation of the treatment effect via extension of the estimator proposed by Jeske and Yao. 5 We conclude with a brief summary in Section 6.
Group sequential average rank procedure
As we look to extend fixed sample methods into the group sequential space, an issue to be determined is how to combine information across the stages in the statistic. In this section, we develop a group sequential average rank (SAR) procedure.
Test statistic and critical values
The SAR test statistic is calculated by finding the Wilcoxon rank-sum statistic for each stage using only the observations within that stage and averaging these statistics across stages. We will focus on the scenario where the number of observations in the treatment group at each stage,
In order to determine the mean and variance of the SAR statistic, we must first obtain the mean and variance of
The general forms of the mean and variance, respectively, of
When implementing the SAR procedure, we wish to set up a group sequential design to detect a treatment effect. This may be done in a variety of ways including pre-fixed boundaries or error spending boundaries. We will focus on the latter in this article. The design should be able to reject the null hypothesis for efficacy and accept the null hypothesis for futility. This setup will involve both upper and lower critical values unless the preset maximum number of stages is reached, where a single critical value will be used to reach a final accept or reject decision. Figure 1 displays a visual of this process.

An example of the critical values for the test statistic of a four-stage group sequential trial.
The critical values
The resulting Type I and Type II error rates for each stage are calculated as follows
Once the desired error spending functions are set, the critical values are found using the joint distribution of
Under the alternative hypothesis, the joint distribution of the SAR test statistic depends on the choice of
We suggest using the simple algorithm visualized in the flowchart in Figure 2 to find the critical values. The algorithm starts at an arm size

The flowchart of the algorithm used to find critical values when simulating the exact joint distribution of the test statistic.
The exact joint and marginal distributions of the SAR are discrete with the number of possible values increasing as
The Wilcoxon rank-sum statistic at stage
Let
The existence of the multivariate normal approximation for
Using the normal approximation to obtain the critical values and arm sizes for a given alternative will follow a similar algorithm as simulating the exact joint distribution, as we need to reach the
Comparison of simulated distribution and normal approximation
Due to the large number of simulations needed for a good simulation approximation of (7) to (10), the multivariate normal approximation offers a great increase in computation speed. Figure 3 highlights this computation time difference. Both approaches recommend the same arm sizes in each example and the critical values become virtually indistinguishable for large

Comparison of algorithms for the sequential average rank (SAR) test statistic for various scenarios using 100,000 simulations for the exact simulation with
Determining arm size
The focus will be on the standardized location-scale family of distributions defined by Jeske and Yao.
5
For the normal distribution,
We begin with a demonstration of the arm sizes needed to detect certain alternatives. Various possibilities of the mixture alternative are shown in Table 1 along with the required arm size to detect each alternative. Note the decrease in arm size if the proportion of responders increases. The cases with
Arm sizes per group per stage necessary to detect the mixture alternative with sequential average rank (SAR) using equal arm sizes
in the control and treatment groups for
and
power with
determined by simulation of the joint distribution. Arm sizes determined by the normal approximation are in parentheses. The values for the proportion of responders,
, and size of the shift,
, are displayed.
shows the results for a fixed sample design.
Arm sizes per group per stage necessary to detect the mixture alternative with sequential average rank (SAR) using equal arm sizes
It can be seen in the arm size tables that for a given alternative, there is an ordering of the distributions in terms of arm size needed to detect the same treatment effect. This ordering of the distributions is explained by Jeske and Yao
5
with the ordering of the Kullback–Leibler distance between
The arm size results determined via simulation may differ slightly from other runs due to the number of simulations. This variation may also be an explanation for the differences in results between simulation and the normal approximation, especially in large arm size situations. Otherwise, the arm sizes determined by both algorithms are nearly the same, illustrating the accuracy of the multivariate normal approximation.
The average sample number shows what the total sample size will be on average for a given design. It accounts for the random number of stages associated with the trial. Average sample numbers for a variety of designs are displayed in Table 2. The values may be calculated by simulating each setting until termination, finding the average stage at which it stopped, then multiplying the average by the sum of the number of observations per stage of the two groups,
Average sample number to detect the mixture alternative with sequential average rank (SAR) using equal arm sizes in the control and treatment groups for
Equal arm sizes in the control and treatment groups may not always be feasible due to practical limitations such as finding sufficient participants for the treatment group. In these cases, the experimenter may get the arm sizes necessary to detect a given design alternative for a specified ratio between the groups. Table 3 displays the treatment arm sizes and average sample numbers for various design alternatives where the control group arm size is two times the treatment group arm size.
Treatment arm size with average sample number (in parentheses) to detect the mixture alternative with sequential average rank (SAR) for
Next, we wish to investigate the power of the SAR. Here we set the design alternative with a particular
The alternative hypothesis used in this article is a specific case of testing if
The power curves for each test statistic and the four different design scenarios are shown in Figure 4. Clearly, each test statistic provide essentially the same power curves. Since the procedure from Peng et al.
6
does not use

Power curves for the sequential average rank (SAR), log-rank, MaxCombo, and Peng et al. test statistics from a design alternative with number of stages
The SAR, log-rank, and MaxCombo are able to take a hypothesized distribution of the data into account, thus adapting its design to the data. This results in different critical values (not shown) as well as different arm sizes needed to detect the alternative. The noticeable difference between the methods is the arm size. The method from Peng et al.
6
will recommend the same arm size for different
It can be seen that the order of distributions from greatest to least arm size is normal, logistic, Laplace,
In this section, we discuss the robustness of the SAR procedure. We address the concern of the procedure’s performance if a hypothesized value of the design alternative, namely
First, consider
The same reasoning can be applied to the situation where
Average treatment effect
Instead of thinking of the treatment effect as the pair
Table 4 presents the arm sizes necessary to detect several different design alternatives based on
Arm sizes per group per stage necessary to detect the mixture alternative with sequential average rank (SAR) using equal arm sizes
in the control and treatment groups for
and
power with
determined using the normal approximation. The values for the proportion of responders,
, and
, where
is the size of the shift, are displayed.
shows results for a fixed sample design.
Arm sizes per group per stage necessary to detect the mixture alternative with sequential average rank (SAR) using equal arm sizes
The second method we wish to consider to combine the observations at each stage we will deem the group sequential rerank (SR) procedure. This method is explored by Spurrier and Hewett 7 for two stages and general alternative. It is used by Shuster et al. 8 for ordinal categorical data with an arbitrary number of stages and general alternative. We investigate the test statistic in an arbitrary number of stages with continuous data under a mixture alternative.
Test statistic and critical values
For this method, we will again be getting the Wilcoxon rank-sum statistic at each stage. However, the SR statistic at stage
Let
The covariance of the SR statistic between stages
Like the SAR, we consider a multivariate normal approximation to circumvent the need for running the large number of simulations that would be necessary to enumerate the exact distribution. It is clear that the marginal distribution of the SR statistic at each stage is asymptotically normal both under the null and alternative hypothesis, given that each is a standardized Wilcoxon rank-sum statistic. The limiting multivariate normal distribution of the SR statistic has been investigated and shown to hold by several authors.18,8,7
Figure 5 shows a comparison of simulation and normal approximation critical values, arm sizes, and computation time. In the various scenarios, both algorithms give the exact same arm size. There are some small differences in the critical values in the smaller arm size situations. The critical values become indistinguishable between the two algorithms as the arm size increases. As with the SAR, we see a great improvement in computation time by using the normal approximation. The multivariate normal approximation may be reasonable for

Comparison of the multivariate normal approximation and simulation of the exact distribution for the sequential rerank (SR) using 100,000 simulations with
With both the SAR statistic and SR statistic based on the Wilcoxon rank-sum statistic, it is of interest to see how they compare and whether one may be preferred over the other. We begin by looking at the limits of the mean, variance and covariance of each statistic. At each stage,
Next, we examine the covariance of the test statistics for stages
For a further comparison of the test statistics, Table 5 shows the arm sizes needed to detect various alternatives for the SR statistic. The arm sizes in parentheses are those needed by the SAR statistic. Almost all comparisons are equal. The SR statistic requires one fewer observation per group in only a handful of scenarios.
Arm sizes necessary per group per stage to detect the mixture alternative using the SR statistic with equal arm sizes
for
and
power with
determined by the normal approximation. The arm sizes for the SAR statistic determined by the normal approximation are in parentheses. The values for the proportion of responders,
, and size of the shift,
, are displayed.
shows results for a fixed sample design.
Arm sizes necessary per group per stage to detect the mixture alternative using the SR statistic with equal arm sizes
SR: sequential rerank; SAR: sequential average rank.
With the evidence of asymptotic equivalence and equality of arm sizes in Table 5, the SAR statistic is the recommended test statistic over the SR statistic. An added benefit of the SAR statistic is the computational convenience of calculating the Wilcoxon rank-sum statistic with only the observations at each stage, whereas all observations need to be used throughout the entirety of the experiment with the SR statistic.
An R function provided in the Supplemental Material with user documentation can be used to calculate the arm size and critical values for various design alternatives using the four choices of
Estimation
The treatment effect in our setting is
Method of moments
First, we consider a method of moments (MoM) estimator that builds upon the estimator proposed by Jeske and Yao.
5
These estimates use the sample mean and sample variance of the control and treatment groups,
Due to the randomness of the sample statistics, it is possible for the estimators to be outside the parameter space. In order for the estimators to appropriately represent the null scenario, if either estimator is equal to zero then the estimate of the other will also be set to zero. For the MoM estimators, this situation would occur only if
Our MoM estimators for
Constrained
-means
Next, we explore estimating the treatment effect using the constrained
With this constrained algorithm, we can utilize both the control and treatment observations together in the
Occasionally for two clusters, the cluster that includes the control group observations will also include the largest observations of the treatment group. This contradicts the assumption that the largest observations in the treatment group would come from a shifted version of
Under the mixture alternative, there should be two clusters. Let the cluster with the control group observations be Cluster 1 with mean
Results
This section will discuss the results of simulating estimates of
RMSE values for estimating
RMSE values for estimating
based on 1000 simulations where the data are generated from the
, proportion of responders
, and shift size
values indicated in the table with
. The design alternative uses
2,
0.7,
1.5,
Normal,
0.01,
0.1,
, and
resulting in an arm size of 17 for both the control and treatment groups.
RMSE values for estimating
RMSE: root mean squared error; MoM: method of moments; CKM: constrained
Represent the lower RMSE between the MoM and CKM estimators for each combination of
We also consider the RMSE when estimating
RMSE values for estimating
RMSE: root mean squared error; MoM: method of moments; CKM: constrained
Represent the lower RMSE between the MoM and CKM estimators for each combination of
In this article, we argue that the traditional assumption of the pure shift model to represent the treatment distribution may not always be correct. A mixture model may be a more appropriate representation of the treatment group to account for the possibility that there are individuals who are unaffected by the treatment. This phenomenon was introduced by Good 3 and Boos and Brownie. 4 The Food and Drug Administration 22 and Lubich et al. 19 provide further evidence and examples of such nonresponders. If nonresponders exist but the experimenter erroneously works under the pure shift assumption, a design will be underpowered.
We present a novel method of using the Wilcoxon rank-sum statistic in group sequential clinical trials with the SAR test procedure. It’s usefulness over existing methods if the experimenter has some idea about
We have shown how error spending methods may be used to set up a group sequential design with the SAR test procedure. However, the SAR test may be used with other group sequential boundaries since the distribution of the test statistic is unchanged by the choice of boundary.
Two estimators are evaluated for the treatment effect
Supplemental Material
sj-pdf-1-smm-10.1177_09622802231184634 - Supplemental material for Wilcoxon rank-sum tests to detect one-sided mixture alternatives in group sequential clinical trials
Supplemental material, sj-pdf-1-smm-10.1177_09622802231184634 for Wilcoxon rank-sum tests to detect one-sided mixture alternatives in group sequential clinical trials by Dylan C Friel and Daniel R Jeske in Statistical Methods in Medical Research
Supplemental Material
sj-r-2-smm-10.1177_09622802231184634 - Supplemental material for Wilcoxon rank-sum tests to detect one-sided mixture alternatives in group sequential clinical trials
Supplemental material, sj-r-2-smm-10.1177_09622802231184634 for Wilcoxon rank-sum tests to detect one-sided mixture alternatives in group sequential clinical trials by Dylan C Friel and Daniel R Jeske in Statistical Methods in Medical Research
Footnotes
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship and/or publication of this article.
Supplemental material
Supplemental material for this article is available online.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
