Abstract
Despite the best efforts of investigators, problems forcing design changes can occur in clinical trials. Changes are usually relatively minor, but sometimes not. The primary endpoint or analysis may need to be revised, for example. It is common to regard any conclusion from such a tarnished trial as hypothesis-generating rather than definitive. This article reviews a very useful technique, re-randomization tests, for dealing with such anomalies. Re-randomization tests remain valid for testing a strong null hypothesis that treatment has no effect on the data that led to design changes. Another way of expressing this is that the data used to inform a design change must give no information about the treatment labels. This restriction has implications for limiting the amount of information examined by a committee deciding whether to make design alterations. While nothing can eliminate the pall cast by breaches of protocol, re-randomization tests following blinded and limited data examination go a long way toward amelioration.
Keywords
Introduction
This article reviews the role of re-randomization tests when unplanned changes are made before breaking the treatment blind in clinical trials. Are conclusions from such a trial only hypothesis-generating, or are they definitive? We will see that the truth lies somewhere between these two extremes.
Randomization, blinding, and pre-specification of outcomes and analysis methods are cornerstones of clinical trials. Randomization prevents biased assignment of patients to treatments, tends to balance groups with respect to important prognostic factors, and can form the basis for testing the null hypothesis of no effect of a treatment. Regarding the last point, a re-randomization test, also called a randomization test or permutation test, can be applied whereby the null distribution is approximated by fixing the data at their observed values, re-randomizing according to the randomization method (e.g. permuted blocks), computing the value of the test statistic under these re-randomized treatment labels, and repeating this process a huge number of times. This idea was first used by Fisher in connection with two experiments, one to confirm or refute a woman’s claim that she could taste whether milk or tea was added first to a cup, and the other involving growth rates of plants. 1 He recognized that with large sample sizes, re-randomization tests give nearly the same answer as their parametric counterparts, a fact that has been corroborated by several authors.2–4 Permutation tests have been applied in simple settings such as those considered by Fisher, and to more complicated settings involving ordering p-values from multiple comparisons and testing them in a way that protects the family-wise error rate. 5 Section 1.1.1 of Good 6 presents an A–Z list of the multitude of applications of these tests in different fields.
Blinding of treatment assignments tends to balance the groups with respect to the power of positive thinking, and prevents patients or doctors in the control arm from supplementing their therapy to make up for the fact that they did not receive the experimental treatment, for example. Blinding may also lessen dropout, as patients in an unblinded trial who know they are receiving placebo may discontinue participation.
Pre-specification of analyses is essential to avoid inflation of the type-1 error rate due to multiplicity. Imagine, for example, switching from a signed rank test to a sign test after observing that all paired differences have positive signs. The actual type-1 error rate is impossible to even quantify because if we had seen data that were favorable for a different test (e.g. the paired t-test), we might have switched to that test. For this reason, data-driven changes in clinical trials are rightly criticized.
Despite the fact that unplanned changes should be avoided whenever possible, they are sometimes unavoidable. It is not unusual to see a primary event rate lower than anticipated. In extreme cases, power may be dismal. This might lead to increasing the sample size, changing recruitment to target a higher risk population, or changing the primary endpoint. The first two options are less extreme and more common than the third, but changing the primary outcome can be appealing in certain situations. For example, suppose that, a priori, there are two nearly equally compelling endpoints A and B, but A is selected as primary. If the overall event rate for A turns out to be much lower than for B, a switch from A to B may be in order. Alternatively, in a trial of lung disease, A and B may be outcomes defined in terms of imaging using a new technique, and blinded evaluation of scans reveals that A is much more difficult than B to measure reliably. Again, a change in primary endpoint may be in order.
This unplanned change setting and the closely related article of Posch and Proschan 7 is somewhat different from the long literature on pre-planned adaptations based on blinded data. For example, Hogg et al. 8 examined rank statistics from lumped data from both groups to deduce the heaviness of the tails of the distribution and thereby choose a powerful rank test. Gould 9 and Gould and Shih 10 showed how to recalculate sample size using blinded data for binary and normal outcomes, respectively. Edwards 11 used blinded data to select from a set of pre-specified models, and then applied a re-randomization test. Proschan et al. 12 considered a similar adaptive regression setting. These methods exploit the fact that if the alternative hypothesis is true, we can gain an advantage by tailoring the analysis to what is seen in blinded data, while if the null hypothesis is true, we seem to lose nothing. We will see that this is not quite correct; there can be difficulties in interpretation. In the setting we consider, something untoward has caused us to change the outcome or analysis plan. The same reasoning used for pre-planned analyses can be used to justify re-randomization tests for unplanned changes, and the same limitations apply. The difference is that when we apply these methods out of necessity, the advantages far outweigh the disadvantages.
The outline for the remainder of this article is as follows. We first review re-randomization tests in a static setting with no changes. We show that they control the conditional type-1 error rate, given the observed outcome data
Re-randomization tests and conditioning
This section summarizes the well-known results that (1) re-randomization tests are conditional tests given a set of data and (2) controlling the conditional type-1 error rate controls the unconditional error rate as well. Lehmann
13
recognized these facts in connection with a two-sample shift alternative setting with unknown and arbitrary continuous density
Throughout this article, we consider the setting of rejecting a null hypothesis of no treatment effect for small values of a test statistic
To understand why re-randomization tests work, imagine a time reversal in which we observe data
where
In reality, we randomize
The formal short proof is as follows
because
The critical value
This follows immediately from the fact that unconditional type-1 error rate is simply the expected value of conditional type-1 error rate. Theorem 2 tells us that the type-1 error rate of a re-randomization test is controlled at level
Unplanned change setting
The preceding section showed that a re-randomization test applied to a pre-specified outcome variable
As before, it is helpful when trying to assess validity to picture the time reversal scenario in which the data
Before making precise the notion in the preceding paragraph, we need some notation. Denote the new test statistic and outcome variable selected after examining
The proof is as follows. The Borel assumption clearly implies that
When we use Theorem 3 in an unplanned change setting, we are implicitly assuming that our actions can be represented by some function
A crucial part of Theorem 3 is that it tests the strong null hypothesis that
Examples and limitations
We begin with a fairly innocuous example to illustrate the usefulness of re-randomization tests in an unplanned change setting. We then move to murkier examples and illustrate why there are down sides of the procedure.
The solution to this problem in the clinical trial is simple: switch to a Wilcoxon rank-sum test and use its re-randomization distribution to test for statistical significance. By Theorem 3, this controls the conditional and unconditional type-1 error rate.□
The following examples illustrate the restriction on conclusions that can be made from re-randomization tests in an unplanned change setting.
To understand why Example 2 does not contradict Theorem 3, consider the following remark.
We did not make a type-1 error in the apparent paradox of Example 2 because the joint distribution of
The fact that the re-randomization test allows rejection of only the strong null hypothesis is problematic because we are always looking at more data than we think. When we look at any outcome variable, we see not only the blinded data, but also the amount of missing data. Therefore,
Although Example 2 showed that one can technically conclude only that treatment had some impact on
Remember that the methods of this article are to be used in emergency settings only. When something untoward occurs, the onus is on the trial investigators to argue, based on the facts of the particular trial, that results were not just a consequence of inadvertent unblinding. For instance, suppose there were very few missing observations, and only two outcomes were examined, either of which could reasonably have been selected a priori as the primary endpoint. If, after a forced change of endpoints based on blinded data, the ultimately selected outcome showed a statistically significant result, one could make a good case that treatment benefitted that endpoint.
Mid-course changes
Up to now, we have considered a simple situation with no interim monitoring, and any change is made at the end of the trial. A slightly different approach is to examine blinded data at a single intermediate time point, make a change in primary outcome or analysis, and then use a re-randomization test on the selected outcome at the end of the trial. The proper analysis at the end of the trial is a stratified re-randomization test. For instance, suppose that we are conditioning on per-arm sample sizes, as in Remark 2. Then we construct the re-randomization distribution by considering only those re-randomizations that match the per-arm sample sizes at both the time of the change and the end of the trial. For example, if the numbers assigned to treatment and control are (206, 197) at the time the change was made and (400, 395) at the end of the trial, then one would consider only those re-randomizations yielding these same per-arm sample sizes at the time of change and end of study.
A much more complicated setting is when there is interim monitoring in addition to an unplanned change. For example, monitoring may have occurred with the original primary endpoint, then a change is made, and monitoring proceeds with the new primary outcome. In most cases, this presents enormous problems with both statistics and interpretation. Monitoring is often based on a Brownian motion approximation to the joint distribution of a standardized statistic over time (Lan and Zucker,
20
chapter 2 of Proschan et al.
21
). It would be very odd and difficult to calculate boundaries using one endpoint for some time points and another for other time points. One potential option would be to subtract the amount of alpha used by the time of change. For instance, suppose that the change in endpoint is made after using cumulative alpha
Our simple setting does not cover response-adaptive randomization, which continually changes treatment assignment probabilities in response to observed outcome data on a short-term binary endpoint. Interestingly, a re-randomization test in that setting is also valid under a very commonly satisfied condition given by Simon and Simon. 23
Discussion
Re-randomization tests are invaluable tools when something goes awry in a clinical trial and forces an unplanned design change before breaking the treatment blind, though the methodology is not a panacea. A major drawback is that, technically, the conclusion one can draw from a statistically significant result is that the joint distribution of all data examined differs by arm. The more data one uses to make a change, the weaker the conclusion must be when rejecting this null hypothesis. Moreover, we are always looking at more than we think, such as the amount of missing data.
Another problem with unplanned changes is that they do not lend themselves to estimation. It is easy to compute the null distribution of the test statistic because the strong null hypothesis implies that
One word of warning is in order. Sometimes data and safety monitoring boards review data with labels “Arm A” and “Arm B,” without knowing which arm corresponds to the experimental treatment. They argue that keeping the arms “blinded” in this way until they are ready to make a definitive conclusions prevents bias. This does not constitute blinding in the sense of this article. The only valid re-randomization distribution if one sees data by arm, but without knowing which arm is the experimental treatment, would assign probability 1/2 to arm A being treatment and 1/2 to control. In other words, the re-randomization distribution takes on only two possible values. The one-tailed re-randomization p-value can be either
The methods described in this article are intended for emergency use. Besides the limitations mentioned above, there is a very practical concern that results may not be accepted. Many clinical trialists are steadfast about following the protocol. They believe that whenever there is a major change, such as a new primary endpoint or analysis method, results can only be hypothesis-generating. This skepticism is understandable. Nonetheless, I believe that when the change is made before breaking the treatment blind, a re-randomization test provides evidence that is closer to definitive than to hypothesis-generating.
Footnotes
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
