A parallel randomized trial is frequently used to investigate the treatment effectiveness as compared to the gold standard. In early phase trials, a group sequential design has the potential to reduce the expected sample size as compared to the traditional one-stage design, and protect participants when a new treatment is not as effective as expected. When the outcome is binary, a group sequential design based on exact binomial distribution is preferable as compared to the asymptotic limiting distribution. To improve the design efficiency, we propose to develop new parallel two-stage adaptive design and promising zone design allowing sample size adjustment in the second stage based on the outcome from the first stage. The conditional probability is guaranteed in the proposed designs when a trial proceeds to the second stage. All these designs control the type I error rate, but only the proposed two designs guarantee the conditional probability constraint. We used a real example from a completed cancer trial to illustrate the application of the proposed designs. The adaptive design substantially increases unconditional power but requires a large sample size as compared to the group sequential design. The promising zone design achieves a good balance between statistical power and the expected sample size.
Randomized clinical trials are commonly used in assessing treatment effectiveness as comparing a new treatment to the gold standard. In early phase trials, as compared to traditional randomized designs, group sequential (GS) designs have the potential to reduce sample size and protect patients when a new treatment is not effective. For a study with a binary outcome, the binomial distribution is preferable to calculate the exact type I error (TIE) rate and statistical power instead of asymptotic limiting distributions. Kepner1 was among the first to develop a two-stage GS design by using the exact binomial distribution for a study allowing early stopping after the first stage. In practice, it would be reasonable to allow early stopping due to futility after the first stage to protect patients in the case that a new treatment is not as effective as expected. The GS design based on exact statistical distribution controls the TIE rate, ensuring rigorous decision-making without relying on asymptotic approximations.1 Despite its advantages, the GS design has some limitations. It is not adaptive with the sample size and decision thresholds which cannot be modified based on interim results. This limits the ability to better utilize emerging trends from the observed data.
To improve the efficiency of a clinical trial, adaptive designs may be utilized by enabling sample size re-estimation and adaptive decision-making after interim analyses.2–4 Adaptive designs have the potential to make clinical trials more flexible by utilizing accumulating data to modify an on-going trial based on pre-specified rules.5 The GS designs often control for the unconditional power, but their conditional probability (CP) could be low in some scenarios. For each possible early stage outcome, adaptive designs allow sample size re-estimation to make sure that the conditional probability is above the pre-specified threshold with the adjusted sample size. With the sample size re-estimation, its associated threshold value is also adjusted. All these adaptations are pre-specified to have valid study designs.
In a two-stage adaptive design with sample size re-estimation, the second stage sample size could be very large when the first stage treatment effect size is small. Due to the restraints of budgets and study timeline, it may not be feasible to request a very large sample size after the first stage. For that reason, the promising zone design was developed to further extend adaptation by allowing adjustments to the second-stage sample size when interim results suggest potential efficacy.6–8 Under this framework, an interim analysis is conducted to assess whether the treatment effect falls into the pre-specified promising zone based on the conditional power.7 The promising zone is defined as a specific range of interim conditional probability values at which increasing the sample size is warranted to maintain a high probability of trial success.9 If the interim result is too weak, a trial is stopped early for futility after the first stage. When the computed conditional probability falls within the promising zone, the second stage sample size is adjusted to meet the conditional probability requirement.7 This approach ensures that trials have adequate conditional probability to detect a meaningful treatment effect by dynamically modifying the second-stage sample size.10 By incorporating the promising zone strategy based on the GS design, we enhance its efficiency, improve decision-making, and reduce the likelihood of inconclusive results, thus offering a robust framework for clinical trials evaluating treatments with binary outcomes.
In a recent systematic review of clinical trials designed by the promising zone design, almost 50% of them had the binary outcome as the primary outcome.11 Many of these promising zone designs for binary outcomes were developed by using asymptotic limiting distributions. The currently available promising zone design based on the exact binomial distribution was proposed for a single-arm study.12 Given the importance of randomized trials in drug development, it is essential to develop a new two-stage promising zone design with binary outcomes using the exact binomial distribution.
The rest of the article is organized as follows. In Section 2, we first introduced the two-stage parallel randomized GS design with binary outcomes. Based on the first stage sample size and its associated threshold value, we then proposed to develop new adaptive design and promising zone design based on the exact binomial distribution for a parallel randomized study with a fixed sample size ratio between the treatment group and the control group. In Section 3, we compared the performance of the three designs (the GS design, the adaptive design, and the promising zone design) with regard to TIE, statistical power, and the expected sample size (ESS). In a two-stage design allowing early stopping due to futility after the first stage, the final sample size depends on whether a trial proceeds to the second stage or not. The ESS is the expected number of patients in such trials: the first stage sample size plus the ESS in the second stage.13 In this article, we also include the traditional one-stage design in the comparison. Exact approach is used in the one-stage design to calculate sample size.14,15 We used a completed real trial to illustrate the application of the proposed designs. In the last section, we provided some comments and suggestions.
Methods
To evaluate the effectiveness of a new treatment as compared to the gold standard, a parallel randomized two-stage design may be utilized to improve the study design efficiency as compared to the traditional randomized design. Suppose is the response rate of the treatment group (g = T) and that of the control group (g = C). We are interested in testing a one-sided hypothesis: , with the statistical power being computed at , where is the rate difference between the two groups. The TIE is calculated under the null hypothesis with .
The one-stage design has no interim analysis or sample size adjustment. We use the exact test comparing two proportions based on the Boschloo’s approach in the R function “power.exact.test” to compute the TIE and statistical power,16 and identify the smallest sample size that achieves the actual statistical power above the nominal level while the TIE is controlled.
In a two-stage design, suppose is the number of responses among participants who are assigned to group in the -th stage (). The total sample sizes in the first stage and the second stage are and , respectively. It follows that the maximum possible sample size is .
We first introduced the existing two-stage GS design based on exact binomial distributions developed by Kepner.1 He developed the GS design with equal allocation between the two groups, which can be readily extended to a study with a fixed sample size ratio. We then proposed to develop new adaptive and promising zone designs allowing sample size re-estimation after the first stage.
Group sequential design by using exact binomial distributions
In a two-stage GS design with equal allocation (e.g. T:C = 1:1), suppose and are the first stage and second stage response differences between the treatment group and the control group. For a study with a fixed sample size ratio (e.g. T:C=2:1), one may use the rate difference for the treatment-control difference in each stage. For simplicity, we used the equal allocation here to introduce these designs.
In this article, we are interested in a study design allowing early stopping due to futility after the first stage. In a two-stage design, the tail probability function can be calculated as:
where and are the threshold values for the responses in the first stage and the total responses from both stages combined, where and . It should be noted that the tail probability is always calculated by using the exact binomial distribution instead of the asymptotic distribution in this article. The tail probability can be expressed as the weighted sum of probabilities across all possible outcomes of in stage 1, which is given by:
In this equation, is the probability of observing in stage 1, and represents the conditional probability of meeting the threshold in stage 2, given the outcome from the first stage. These two quantities are calculated by using exact binomial distributions, such as where is the probability mass function of a binomial distribution and is an index function.
From the tail probability function, we have the TIE and statistical power constraints as:
where and are the nominal type I and II error rates, and is the expected rate improvement of a new treatment as compared to the control. To search for the optimal GS design, it is computationally intensive as one has to search for all possible combinations of , , and given and the aforementioned constraints. In the GS design by Kepner,1 a constraint of the relationship between and was added in the design search: . When the smallest and were determined, there could be a few sets of and that meet both TIE and statistical power requirements. Among them, the one with the largest and the smallest were chosen as the final threshold values to minimize the indecision probability which is the probability of a trial that proceeds to the second stage but ultimately fails to demonstrate treatment efficacy: .1
Adaptive design
Conditional probability represents the likelihood of achieving the overall success criterion in the second stage, given the observed outcome from stage 1. A large conditional probability ensures a great likelihood of meeting the study’s objective(s) if a trial progresses to later stages. In the GS design, the optimal design ensures that the unconditional statistical power is greater than or equal to while the overall TIE is less than or equal to . However, it is possible that the calculated conditional probability may be very low for some scenarios. To address scenarios with low conditional probability values, the sample size and threshold for the second stage may be re-estimated to guarantee the nominal conditional probability level.
When , a trial proceeds to the second stage. Among the first stage outcome , suppose there are cases whose conditional probability is below the nominal level of . There are a total of first stage outcome cases whose conditional probability is lower than . The conditional probability is a non-decreasing function of given the second stage sample size and its threshold value . For the remaining cases with , their conditional probability values are above . As their conditional probability is large enough, the planned second stage sample size is sufficient.
When the first stage outcome is between and , the re-estimated and given the first stage outcome are determined to meet the following two constraints:
where is the second stage response difference between the treatment group and the control group with the new sample size , and accounts for the permissible TIE adjustment, with being the actual TIE rate of the GS design. To avoid being overly conservative while still controlling the TIE, we allocate the remaining error equally across the cases by setting . The first TIE constraint is used to control for the TIE of the adaptive design. When several combinations of and satisfy the two constraints, we choose the combination that minimizes both and given the first stage outcome where .
Once we have determined the second stage adaptive sample size and threshold value for , we calculate the ESS under the null hypothesis for the adaptive design. For each possible first stage outcome , the sample size is if a trial is stopped for futility, or if a trial proceeds to the second stage. That sample size is weighted by its probability in computing the ESS, which is presented as:
When the conditional probability is very low for the cases having slightly above , the new may have to increase significantly to meet the nominal level, potentially affecting the overall efficiency and feasibility of the study. For example, as illustrated in Figure 1 for a study with the design parameters , the first stage threshold value of the GS design is . Suppose a study has the first stage outcome . Then, the re-estimated of the adaptive design increases to 546 which is more than four times the planned of the GS design. This significant increase in the second stage sample size may lead to several challenges in conducting a trial, such as increased costs, a longer study duration, and higher resource demands. To address that, we further consider the promising zone design allowing sample size adjustment for the cases with the conditional probability above the pre-specified acceptable level.
The second stage sample size per group as a function of the first stage outcome for a balanced randomized study with , , and . The GS design is from the article by Kepner (2010). The second stage sample size here is the sample size per group. The sample size for the one-stage design is 134 per group.
Promising zone design
In the promising zone design, let the minimal acceptable condition probability be (e.g. 50%, 40%).6,17 If is greater than or equal to but its conditional probability falls below , a trial is stopped after the first stage due to futility. Suppose there are cases whose conditional probability is below among the cases with , a trial is stopped after the first stage when . Similar to the adaptive design, the second stage sample size remains unchanged when a study has a large conditional probability with . For the cases in the promising zone with , the re-estimated sample size and threshold value are determined by meeting the following two constraints:
Where , and is the TIE rate of a GS design with the updated first stage threshold value parameter that replaces in the original GS design.
We selected the smallest values of and for configurations among in the promising zone design when multiple values meet the TIE and statistical power constraints. The ESS for the promising zone design is calculated as follows:
The monotonic relationship between the second stage sample size and the first stage outcome should be respected in adaptive and promising zone designs.
The second stage sample size is a non-increasing function of the first stage outcome when a trial proceeds to the second stage.
Please find the detailed proof in Appendix. This theorem shows the monotonic relationship between the second stage sample size and the first stage outcome when a trial proceeds to the second stage for the adaptive design and the promising zone design.
Results
We compared the performance of the two new designs (the adaptive design and the promising zone design), the existing GS design by using the exact binomial distribution, and the exact one-stage design with from 10% to 70% and at 15% or 20%. The GS designs with equal allocation were obtained directly from the tables in the article by Kepner1 for a study to attain 80% statistical power at the significance level of . The first stage sample size and its associated threshold value were used in searching for the adaptive design and the promising zone design. For the promising zone design, we first use the in comparison with other designs. In the last paragraph of this section, we investigate the effect of (30% to 50%) on the statistical power and ESS in Subsection 3.2.
In Figure 2, we presented the minimum, maximum, and average conditional probability values for the GS design to detect a 20% improvement in the treatment group as compared to the control group when the first stage response difference is between and that is the range for the adaptive design having sample size adjustment. It is expected that the maximum conditional probability is below the nominal level as the conditional probability at these values is always less than . We observed that the minimum conditional probability values are often below 40% for these cases with the range from 19.2% to 34.7%. The average conditional probability of these cases is between 47% and 56% for the configurations in Figure 2. As compared to the GS design, the proposed new adaptive and promising zone designs always have the minimum conditional probability above the nominal level for the cases that proceed to the second stage.
The average conditional probability and its range for the GS design for the scenarios that a trial proceeds to the second stage with in a balanced study to detect 20% difference between the two groups given and .
We also studied the minimum, maximum, and average conditional probability for the GS design when . The minimum conditional probability could be lower than 20% in some cases, and the overall average conditional probability for the GS design when is 4% lower than that when in Figure 2. We also presented these minimum conditional probability values of the GS design in Table 1. These findings provide insights into scenarios where the GS design often fails to maintain adequate conditional probability when .
Comparison between the three designs (the GS design, the adaptive design, and the promising zone design) with regard to TIE, unconditional power, and the ESS under the null hypothesis for studies to detect a 20% or 15% difference between the treatment group and the control group to achieve 80% power given .
GS design
Adaptive design
Promising zone design
TIE
Power
ESS
TIE
Power
ESS
TIE
Power
ESS
N
0.1
0.347
0.0395
0.8031
60.0
0.0383
0.8660
88.7
0.0364
0.8279
64.5
110
0.2
0.214
0.0404
0.8015
98.0
0.0378
0.8764
157.9
0.0344
0.8064
88.3
140
0.3
0.209
0.0416
0.8014
118.8
0.0401
0.8753
184.4
0.0352
0.7891
100.2
154
0.4
0.268
0.0434
0.8019
122.5
0.0419
0.8686
174.1
0.0382
0.8057
114.2
160
0.5
0.192
0.0459
0.8017
127.4
0.0442
0.8774
196.9
0.0404
0.8041
111.2
156
0.6
0.233
0.0498
0.8004
109.3
0.0477
0.8726
157.9
0.0433
0.8033
100.0
138
0.7
0.307
0.0448
0.8023
88.7
0.0427
0.8678
114.4
0.0381
0.8020
82.7
102
0.1
0.242
0.0463
0.8029
104.7
0.0411
0.8760
176.8
0.0356
0.7716
84.3
180
0.2
0.118
0.0447
0.8005
177.6
0.0428
0.8818
322.8
0.0375
0.7807
136.7
230
0.3
0.130
0.0427
0.8001
220.8
0.0405
0.8806
380.5
0.0368
0.7796
169.9
268
0.4
0.179
0.0471
0.8005
227.0
0.0457
0.8762
360.3
0.0403
0.7893
189.8
288
0.5
0.262
0.0499
0.8002
210.4
0.0484
0.8665
295.1
0.0438
0.7990
194.8
272
0.6
0.249
0.0447
0.8007
204.6
0.0435
0.8691
285.5
0.0395
0.8000
185.6
248
0.7
0.073
0.0459
0.8001
196.1
0.0449
0.8853
347.4
0.0382
0.7839
148.0
204
The sample size in the last column is the total sample size of the one-stage design.
The calculated ESS and unconditional power of the three designs as a function of for a fixed sample size ratio study (T:C1:2 and 2:1) to detect 20% difference between the two groups given and .
In addition to the minimum conditional probability of the GS design, we also presented the TIE rate, statistical power, and ESS of the three designs with equal allocation in Table 1 when and 15%. All these three designs control for the TIE with the actual TIE rate being less than or equal to the nominal level . In most scenarios, the sample size of the one-stage design exceeds the ESS of the GS design by approximately 20% to 30%. Notably, when and , the one-stage design sample size exceeds the ESS of the GS design by 83.3%, highlighting that even a two-stage GS design can offer substantial sample size savings as compared to the one-stage design. As compared to the GS design, the adaptive design’s statistical power is 7.3% higher, at the cost of larger sample sizes using the adaptive design (231 on average for the adaptive design VS 148 for the GS design). The GS design and the adaptive design have the same threshold value of the first stage outcome. With the increased second stage sample size in the adaptive design to meet the conditional probability requirement, it is expected that both ESS and the statistical power go up. Meanwhile, the promising zone design may have a shorter range of where a trial proceeds to the second stage as compared to the other two designs. For that reason, the unconditional power of the promising zone could be slightly below the nominal level as observed in one configuration when and several configurations when . We found that the promising zone design can reduce the ESS by more than 40% as compared to the adaptive design, and has fewer ESS by 14% on average as compared to the GS design.
For a study with a fixed sample size ratio, we presented the ESS and statistical power comparison between the four designs with in Figure 3 when the sample size ratio of (left) and 2:1 (right). All three designs control the TIE rate with the actual TIE rate being bounded by the nominal level. The results from these two cases are very similar to each other. The findings are similar to the findings in Table 1 for a study with equal allocation between the two groups. The promising zone design has a good balance with regard to ESS and statistical power, although the unconditional power of this design could be slightly below the nominal level by 2% in some cases. In contrast, the one-stage design consistently achieves the statistical power of 80% or above, but at the cost of significantly larger ESS similar to the results of equal allocation, demonstrating its relative conservativeness and inefficiency. The conditional probability values of the promising zone design and the adaptive design are all above the nominal level while the GS design could have the average conditional probability below 60% for a study that proceeds to the second stage. Similar results were found for the cases with a fixed sample size ratio of 2:1 or 1:2 when .
Effect of the actual rate difference
When the actual rate difference observed in a study is lower than , we present the statistical power of the four designs in Figure 4 by using the sample size that is calculated with the expected rate difference between the two groups. When the actual rate difference equals the true value, all the designs have the statistical power being close to or above the nominal level. Adaptive designs have the statistical power above the nominal level due to the large ESS of this design as compared to others (see Figure 3). In the case of with , the statistical power of these designs is reduced by 21% on average with the range from 19.9% to 22.3%. As the actual rate difference is further reduced to 50% of , the statistical power could be lower by 47% as compared to the power when is the expected rate difference. Similar results are observed for the cases with .
Statistical power is plotted as a function of the control group rate for the four designs when the true effect is 100%, 75%, or 50% of the assumed design effect (top row , bottom row ).
In Figure 4, the computed statistical power is similar between the GS design and the promising zone design. As suggested by the reviewers, we also compare their ESS when their unconditional power is 80% or above for the cases of with . When the unconditional power of the proposed design is below the nominal level, we first search for the GS design with the nominal power level of 80.1%. Then, we obtain the proposed design, and calculate the statistical power of the design. We increase the nominal power for the GS design by 0.1% each time until the statistical power of the proposed design is 80% or above. Similarly, we decrease the nominal level of the GS design by 0.1% when the proposed design has the statistical power above 80%. The results show that the adaptive design consistently requires larger sample size than GS design, with an average increase of 45.7 patients (25.3%), ranging from 6.1 patients (3.1%) to 72.6 patients (34.5%). In contrast, the promising zone design reduces the sample size relative to the GS design, with an average decrease of 23 patients (12.1%) and reductions ranging from 3 patients (1.4%) to 41.1 patients (21%). Therefore, when all three designs achieve similar TIE control and power, the adaptive design increases ESS substantially, whereas the promising zone design provides substantial sample size savings.
Effect of on the promising zone design
When the CP threshold value is set as 50%, a trial needs to have the CP above 50% after the first stage in order to move forward to the second stage. In practice, that requirement may be adjusted given the fact that the first stage sample size is often not large enough. In Table 4, we present the actual TIE, statistical power, and the ESS of the promising zone with , 40%, and 30% when the expected rate difference is and 15%. The ESS goes up as decreases from 50% to 30%. For all the cases in Table 4, the average ESS is 126.4 when , and it goes up to 177.9 when . With the increase of ESS when , it statistical power is increased by 6% on average as compared to the cases having .
Real data example
We used a randomized phase II study for patients with relapsed/refractory classical Hodgkin’s lymphoma (HL)18 to illustrate the application of the proposed designs. That was a balanced randomized study comparing the overall response rate between the gemcitabine, vinorelbine, and liposomal doxorubicin (GVD) plus SGN-30 group and the GVD group. That study was designed to detect a 15% improvement in the treatment group with the estimated response rate in the control group as 70%, given and .
We applied the four designs to this real example, and presented the detailed study designs in Table 2. When a trial has the first stage outcome , it proceeds to the second stage with an additional 31 patients per group if the GS design is utilized. In this scenario, the conditional probability is below 30% with the planned second stage sample size. In order to achieve the conditional probability of 80%, a total of 254 patients would be needed in the adaptive design given . This could cause a longer study if this occurs when the adaptive design is used. The promising zone design would determine to stop the trial due to futility when the conditional probability is below 50%. In Table 3, we presented the actual TIE, unconditional power, and the ESS for the three designs for this example. All designs control the TIE. Their unconditional power values are above the nominal level for this example. The GS design requires the smallest ESS (105), followed by the promising zone design (109), and the adaptive design (160). The exact one-stage design requires the sample size of 110. This finding is consistent with the results in the numerical studies: the one-stage design, the GS design and the promising zone have similar sample size when is large (e.g. 70%).
Detailed study designs to detect a 15% increase in overall response rate with the control rate of 70% to achieve 80% statistical power given .
r
0.19
0
.
0
.
0
.
1
0.28
31
6
127
13
0
.
0
0.40
31
6
104
11
0
.
1
0.52
31
6
79
9
79
9
2
0.64
31
6
63
8
54
7
3
0.75
31
6
46
7
46
7
4
0.84
31
6
31
6
31
6
The GS design is (. This is a balanced study. We present the sample size per group here.
Comparison between the three designs to detect a 15% increase in overall response rate with the control rate of 70% to achieve 80% statistical power given with regards to TIE, statistical power, and ESS.
Design
ESS
Power
TIE
Group sequential
104.9
0.8008
0.1395
Adaptive
159.2
0.8688
0.1328
Promising zone
108.3
0.8143
0.1262
The GS design is (. The total sample size of the one-stage design is 110.
Comparison of different for the promising zone design with regard to TIE, unconditional power, and ESS under the null hypothesis for studies to detect a 20% or 15% difference between treatment group and the control group to achieve 80% power given .
Promising zone design
TIE
Power
ESS
TIE
Power
ESS
TIE
Power
ESS
0.1
0.0364
0.8279
64.5
0.0364
0.8279
64.5
0.0383
0.8660
88.7
0.2
0.0344
0.8064
88.3
0.0365
0.8411
107.2
0.0374
0.8634
131.8
0.3
0.0352
0.7891
100.2
0.0375
0.8236
115.3
0.0388
0.8483
135.5
0.4
0.0382
0.8057
114.2
0.0402
0.8343
130.1
0.041
0.8547
150.3
0.5
0.0404
0.8041
111.2
0.0419
0.8336
127.7
0.0431
0.8544
147.5
0.6
0.0433
0.8033
100.0
0.0456
0.8363
115.0
0.0469
0.8585
134.5
0.7
0.0381
0.8020
82.7
0.0408
0.8432
96.5
0.042
0.8678
114.4
0.1
0.0356
0.7716
84.3
0.0389
0.8251
108.3
0.0406
0.8581
140.2
0.2
0.0375
0.7807
136.7
0.0400
0.8151
156.9
0.0422
0.8581
216.3
0.3
0.0368
0.7796
169.9
0.0391
0.8312
212.8
0.0401
0.8485
240.5
0.4
0.0403
0.7893
189.8
0.0435
0.8347
232.3
0.0445
0.8502
259.7
0.5
0.0438
0.7990
194.8
0.0457
0.8227
213.7
0.0479
0.8558
264.3
0.6
0.0395
0.8000
185.6
0.0406
0.8251
205.2
0.0429
0.8588
254.9
0.7
0.0382
0.7839
148.0
0.0408
0.8179
164.2
0.0435
0.8602
212.5
The GS design and the one-stage design do not control the conditional probability, while the proposed two designs do. The adaptive design achieves the highest power but demands a substantially larger ESS, whereas the promising zone design strikes a balance, offering a small increase in ESS while maintaining robust statistical power.
Discussion
The promising zone design effectively ensures that conditional probability remains above the nominal level when a trial proceeds to the second stage, making it a valuable approach for adaptive clinical trials. However, this design may lead to the reduction in the unconditional statistical power, sometimes falling below the target . The primary reason for this power loss is that sample size adjustments are restricted to cases where interim results fall within the pre-defined promising zone, excluding cases with low conditional probability values.19 To mitigate this power loss, one may consider the approach of lowering the threshold for expanding the second-stage sample size which could help improve overall power by allowing more scenarios to proceed to the second stage.
The proposed adaptive and promising-zone designs allow early stopping for futility when a new treatment is not effective as expected. The proposed methods can be modified to allow early stopping for futility or efficacy. An additional parameter for the first stage outcome is needed to define the tail probability as:
When the first stage outcome is +1 or higher, a trial is stopped for efficacy. It can be observed from the results by Kepner,1 the futility stopping threshold value is often not the same between the two designs: the design allowing futility stopping only, and the design allowing futility or efficacy stopping. From a few designs allowing futility or efficacy stopping given , we find that both the GS design and the adaptive design could be more affected by the cases near the futility boundary. However, the promising zone is less likely to be affected as the number of scenarios that have sample size adaptation depends on the value. For a study with a small sample size, it is often preferable to continue a study with the planned second stage sample size when the outcome from the first stage is promising to provide more information for future studies.
The GS design is a minimax type design with the smallest and meeting several constraints. We may search for the optimal GS with increased and , and we then use the new GS design to search for the promising zone design. With the increased sample size, we expect that the unconditional power may be increased. We explored one particular case with and given and . The GS design for a study with equal allocation is (, , , r) = (49, 49, -3, 10). The original promising zone design has the unconditional power of 78.07% as seen in Table 1. When we increased the sample size by 1 per group at each stage with , the new GS design has the updated threshold values: and . Then, the new promising zone design based on and has the unconditional power of 82.2% with the TIE rate of 0.041. The ESS of the new promising zone design is 159.1 which is larger than the ESS of the original promising zone design, but it is still lower than that of the original design with the ESS of 177.6.
In the proposed adaptive and promising zone designs, the conditional probability is calculated by using the expected response rate of the treatment group and the observed first stage outcome. Such conditional probability is frequently used in GS and adaptive two-stage designs when the exact binomial distribution is used.1,4,20,21 Alternatively, one may consider using the conditional power based on the stage 1 treatment effect size.22 In such designs, the asymptotic limiting distribution of the first stage outcome and that of the both stages are often used to compute the TIE rate and the statistical power asymptotically. We consider developing a new randomized study with exact binomial distributions instead of asymptotic limiting distribution in the future.
The multi-arm multi-stage (MAMS) trial design has gained significant attention in clinical research due to its ability to efficiently evaluate multiple treatment arms while allowing for interim adaptations. Applying the promising zone design to MAMS trials presents a compelling future direction, as it introduces adaptive sample size modifications based on CP, enabling more refined interim decisions. By integrating conditional probability based decision rules into the multi-arm framework, promising treatments could continue with optimized sample size allocation, improving efficiency while maintaining statistical rigor.
Footnotes
Acknowledgments
The authors are very grateful to the Editor, Associate Editor, and two reviewers for their insightful comments that help improve the manuscript, and their encouragement to develop more adaptive promising zone designs for clinical trials. Shan’s research is partially supported by the National Institutes of Health under the Award Number R01AG070849, R03AG083207, and R03CA248006.
ORCID iD
Guogen Shan
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
Declaration of conflicting interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Appendix
In the promising design, is the threshold value for the first stage outcome. When , a trial is stopped for futility with no participants in the second stage. Otherwise, a trial proceeds to the second stage with additional participants. When is large enough (e.g. ), its conditional probability is always above the nominal level and is equal to the planned second stage sample size. Therefore, we have when .
When , let and be the new threshold value and new second stage sample size satisfying the TIE and power constraints:
For the case with the first stage outcome , we set
The TIE rate with these and values is calculated as
It is easy to show that
It follows that
Thus, the TIE constraint is satisfied.
It is straightforward to show
Therefore, should be less than or equal to . This property can be easily extended to the adaptive design by replacing with 0 and with , respectively. It is also similar to prove this property for a study with other fixed sample size ratios.
References
1.
KepnerJL. On group sequential designs comparing two binomial proportions. J Biopharm Stat2010; 20: 145–159.
2.
JinHWeiZ. A new adaptive design based on Simon’s two-stage optimal design for phase II clinical trials. Contemp Clin Trials2012; 33: 1255–1260.
3.
BhattDLMehtaC. Adaptive designs for clinical trials. N Engl J Med2016; 375: 65–74.
4.
ShanGWildingGEHutsonAD, et al.Optimal adaptive two-stage designs for early phase II clinical trials. Stat Med2016; 35: 1257–1266.
5.
PallmannPBeddingAWChoodari-OskooeiB, et al.Adaptive designs in clinical trials: why use them, and how to run and report them. BMC Med2018; 16: 29.
6.
ChenYHDeMetsDLLanKK. Increasing the sample size when the unblinded interim result is promising. Stat Med2004; 23: 1023–1038.
7.
MehtaCRPocockSJ. Adaptive increase in sample size when interim results are promising: a practical guide with examples. Stat Med2011; 30: 3267–3284.
8.
JennisonCTurnbullBW. Adaptive sample size modification in clinical trials: start small then ask for more?Stat Med2015; 34: 3793–3810.
9.
MehtaCBhingareALiuL, et al.Optimal adaptive promising zone designs. Stat Med2022; 41: 1950–1970.
10.
HsiaoSTLiuLMehtaCR. Optimal promising zone designs. Biom J2019; 61: 1175–1186.
11.
EdwardsJMWaltersSJKunzC, et al.A systematic review of the “promising zone” design. Trials2020; 21: 1000.
12.
ShanG. Promising zone two-stage design for a single-arm study with binary outcome. Stat Methods Med Res2023; 32: 1159–1168.
13.
SimonR. Optimal two-stage designs for phase II clinical trials. Control Clin Trials1989; 10: 1–10.
StorerBEKimC. Exact properties of some exact test statistics for comparing two binomial proportions. J Am Stat Assoc1990; 85: 146–155.
16.
BoschlooRD. Raised conditional level of significance for the 2 2-table when testing the equality of two probabilities. Stat Neerl1970; 24: 1–9.
17.
ChenJTakanamiYJanssonJ, et al.Practical considerations of promising zone design for interim sample size re-estimation: an application to graphite for graft vs host disease. Contemp Clin Trials2025; 148: 107765.
18.
BlumKAJohnsonJLJungSH, et al.Serious pulmonary toxicity with SGN-30 and gemcitabine, vinorelbine, and liposomal doxorubicin in patients with relapsed/refractory hodgkin lymphoma (HL): cancer and leukemia group B (CALGB) 50502. Blood2008; 112: 232–232.
19.
ShanGDodge FrancisCLiuJ, et al.Application of adaptive designs in clinical research. In: Modern inference based on health-related markers: biomarkers and statistical decision making. Academic Press, 2024, pp.229–243.
20.
LinYShihWJ. Adaptive two-stage designs for single-arm phase IIA cancer clinical trials. Biometrics2004; 60: 482–490.
21.
ShanG. Exact confidence limits for the response rate in two-stage designs with over- or under-enrollment in the second stage. Stat Methods Med Res2018; 27: 1045–1055.
22.
JennisonCTurnbullBW. Group sequential methods with applications to clinical trials. New York: CRC Press, 1999.