Abstract

I really enjoyed Dr Bretz’s analogy of the Swiss Army knife. I am surrounded by a family of Boy Scouts, and they love their Swiss Army knives. The big Boy Scout motto is “Be prepared,” which I think is a perfect analogy for the discussion of adaptive designs. After several years of studying these designs, we all know that the key step, reinforced by Dr Bretz, is the importance of simulations. The dilemma that we still need to address is the infrastructure needed to support the conduct of these simulations. Where is the infrastructure to do these simulation studies? These do not just appear at “a push of the button.” It can take months, sometimes years to figure out the questions we want, or need, to ask. Then, we need to simulate the design(s) to really understand the operating characteristics. There are very few research settings, particularly in academics, where there is a free-standing think tank endowed with the financial coverage, computers, and knowledge to do the more complex simulations. There are basic simulations. For example, if we’re doing a group sequential design to stop early for efficacy and/or futility, we want to look at the probability of stopping early under various stopping rules/boundaries and so forth, so maybe those are the more basic ones that can be done with minimal infrastructure support. However, when you start to mix response-adaptive randomization, early stopping rules, blinded sample size re-estimation into a design, the number of different scenarios that could be anticipated becomes huge. We need to consider how the infrastructure for trial planning will be supported and funded.
We also need to consider the continuing education of our collaborators. Our clinician colleagues want “better” study designs. They have heard about adaptive designs and they really want to push the envelope, which is great, but we need to ensure access to the necessary resources to determine what the infrastructure is going to look like.
Whether you are in an industry or academic setting you have to have affordable and user-friendly tool boxes. It shouldn’t require a special group to do these designs. I hope that one day we can go back to planning grant mechanisms (R34 or the U34, still available from some NIH institutes) that provide funding for administrative planning for a clinical trial. What I would really like to see are planning grants for developing a response-adaptive randomization design. These grants could provide a year of funding to conduct the simulation studies and work through the issues prior to the start of the trial. Continued funding of methodology grants is also necessary in order to further the science because there are still a lot of questions to be answered in this area. Several universities, including mine, have Clinical and Translational Science Award (CTSA) programs that do allow us to think about clinical trial designs and ways to build an infrastructure within our CTSA to support our clinical researchers.
Dr Bretz et al. 1 enumerate factors that are not necessarily considered at the design phase—the implementation factors. Potential awardees focus primarily on the study design and may give less attention to the implementation details until they know about funding. If we actually get the study funded, then there’s that “oh my” moment, where you say to yourself, “How am I going to do it within this budgeted amount?” A perfect example from my own experience is implementing this novel response-adaptive randomization design from your grant proposal so that you could check off your “innovation” component. But now you have to do it, and there are things that none of the investigators considered. For example: if there are expiration dates on your study drug, how are you going to get that drug to the various sites, and how are you going to get shipment of potentially more drug to the site without breaking blinds or revealing potential efficacy information. Another thing to consider with response-adaptive randomization is the updating frequency. After the “burn-in” period [the number of subjects observed before adjusting the randomization ratio], how often do you update the allocation ratio, because that’s obviously going to impact the variability and the power of the study. If you are updating frequently, then what happens with the data validation process? Do you have the infrastructure to make sure that the data validation is being done in a timely manner? For these updates, do you do an “intention to treat” or an “as treated” analysis? You may have treatment crossovers, and when we think about ethical benefits of response-adaptive randomization, is it more ethical to use the data from the “as treated” or “intention to treat” population? And then, of course, you need to consider how you will handle missing data. These are issues that not everyone thinks about when trying so hard to get the funding for the trial. But you are going to be faced with these questions once you actually conduct the study.
This is a good segue to Dr Thall’s 2 presentation. An earlier manuscript from his group 3 reported on a similar simulation study, comparing different response-adaptive randomizations under the scenarios of early stopping versus not early stopping. He shows the clear trade-off between the proposed ethical benefit of response-adaptive randomization and what could actually be happening, depending on which type of response-adaptive randomization you are using and how you are implementing it. Researchers are developing new methods to address these questions, as well as to address the bias in the treatment estimate, as Dr Thall points out. Bowden and Trippa 4 have developed a method to correct for the bias. That’s one way of handling it, but this goes back to a more basic question: should we always use such a correction factor, or should we be using these designs at all? What are the gains/losses when using adaptive designs? Maybe you should be using adaptive designs only in certain scenarios. We also have to think about the impact on power. As Dr Thall showed, there is an impact on power when using response-adaptive randomization, and that impact is going to be based on various aspects of the design including the variability in the allocation ratio, and potentially the updating frequency. Dr Thall presented a scenario with a burn-in period of 10 subjects per arm. We tend to make our burn-in much larger than that, especially in our larger trials. You need simulation studies to determine the right properties of the response-adaptive randomization that you propose.
So when you think about all these complexities, you have to take a step back and say, Okay, yes I probably can do it and I would work very hard to get it done, but why am I doing it? Is it really necessary to do it to answer the clinical question that I have facing me right now?
In some scenarios it is. Maybe it is appropriate for a multi-arm trial versus a two-arm trial, but again, this is where we still have a long way to go of knowing exactly what we should put in our tool box and how we should use it.
Dr Parmar et al. 5 focused on the analysis aspects and the unexpected changes in adaptive designs. We all know that there are potential unexpected changes in any trial. I appreciate the literature on permutation tests, but in general, when we think about everyone who tries to propose an adaptive study, they sometimes forget about the use of the permutation test and how important it really is if you make some changes. Even listening to the speaker’s great examples, there is that gray area when considering whether to change your primary study outcome, and who has input into that decision. For example, the trial sponsor or the Food and Drug Administration (FDA) might have to weigh in—and the trial data and safety monitoring board may also have a say. What does this mean for your pre-specified Statistical Analysis Plan?
A few years ago, I was fortunate enough to be a co-investigator of a grant that was jointly funded by NIH and FDA to advance regulatory science. Our proposal was focused on how you actually go about designing a confirmatory adaptive design. I lead the statistical and data management center for an NIH-funded network called the Neurological Emergency Treatment Trial Network (NETT) that conducts phase III confirmatory studies. We used the NETT as our laboratory. We ran focus groups of the key stakeholders—not only statisticians and clinicians, but regulatory and peer review participants—to learn what is important when they are reviewing these grants, or helping design these studies. We were able to examine what really goes on in an adaptive study. We developed four confirmatory trials during the grant period and were able to submit several grant applications for their conduct. We considered blinded sample size re-estimation if the clinicians weren’t sure about the assumptions underlying their study design, and other approaches such as response-adaptive randomization with covariate balancing. We also developed “shadow studies” that will allow us to examine what would have occurred if we used an alternative design. The Stroke Hyperglycemia Insulin Network Effort (SHINE) trial 6 is currently being conducted by the NETT and was funded as a group sequential design with blinded sample size re-estimation and response-adaptive randomization. The SHINE shadow study is an alternative Bayesian design of the SHINE trial. Once the SHINE trial is completed and submitted for publication, the team will reanalyze the data as if it had been conducted under the alternative Bayesian design. Another NETT trial that was examined is Established Status Epilepticus Treatment Trial (ESETT), 7 which is a comparative effectiveness trial in status epilepticus. These two studies are going on right now and once we’re done with the studies we’ll be able to describe what we have learned about implementing these adaptive designs. In summary, I believe we are on to something but we have a way to go with learning to use all of these things we are adding to our tool box.
I’m a bioethicist and pediatric oncologist here at Penn, and I’m going to be talking about some of the ethical considerations that arise with respect to response-adaptive randomization. Let me just start by saying a couple of words about Dr Bretz’s presentation. He asked about the criteria that define good trial designs. The criteria might be validity and precision of the estimates that you get from your trials, trial efficiency, and third, simplicity. Dr Bretz has a Swiss bias, but we are here in the United States, and I just wanted to introduce him to the American equivalent of the Swiss Army knife, which is the Leatherman® (Figure 1), which I think is much cooler than the Swiss Army knife. The next time he gives this talk, he can say that the adaptive trial is the Leatherman of statistical design!

Leatherman® SurgeTM (reprinted with approval).
Speakers have raised the question of the ethical benefits of using outcome-adaptive randomization, and in particular, the question of whether we can provide more benefits to the trial participants and/or decrease the harm. Professor Thall said the motivation for response-adaptive randomization, at least as reported in the literature, is that it’s more ethical than using fixed ratios since if B is really better than A, then on average, more patients on the trial will be treated with B rather than with A. However, if I understood his remarks correctly, adaptive randomization, at least in two-arm trials, is associated with a meaningful likelihood of actually randomizing more patients to the inferior arm (which is not what you want), biased mean estimates (which is also not want you want), inflated type I error, and/or reduced power. Presumably, if you want to avoid these paradoxical results, then you need much bigger sample sizes. I also took away the message that in most multi-arm designs adaptive randomization may be problematic as well.
So let me back up and share a framework for thinking about these issues. This framework comes from a classic paper by Professor Emanuel et al. 8 (who is now my department chair), and his colleagues at the NIH at the time, laying out seven criteria for ethical research. The first of these is that the question that you’re asking (or the answer to the question) must have social or scientific value. Other criteria: the methods that you’re using must be likely to give you a scientifically valid answer, subject selection must be fair, informed consent must be obtained, the risk–benefit ratio of the trial must be favorable (and here, importantly to point out, it’s not just the risk–benefit for the individual participants but also the benefits from what we learned that we carry forward to future patients), and then a number of ways in which we show respect for potential and enrolled subjects.
The primary considerations that are at stake as we think about adaptive randomization are scientific validity, favorable risk–benefit ratios, informed consent (in ways that I will clarify), and then another value, which was not in the Emanuel framework, which is efficiency. Let’s work through each of these. With respect to scientific validity, I took away from Professor Thall’s talk and also other things in the literature that there are concerns about biased mean estimates associated with response-adaptive randomization. There is of course the possibility that if there are temporal trends, for example, in the prognosis or other concurrent treatments or aspects of the population that are changing over the course of your trial, then that might present a threat to the validity of your trial. Furthermore, the threat multiplies when we’re talking about open-label trials, because both investigators and participants may behave in different ways as they start to see randomizations favoring one or the other treatment.
On the question of efficiency, the overall sample size is generally minimized by one-to-one randomization and is further minimized by appropriate stopping rules for efficacy and futility. But response-adaptive randomization may require increased sample sizes (sometimes even substantially increased sample sizes) to achieve the same alpha and beta errors, and the question is whether this might actually overwhelm the effect of reducing the proportion randomized to the inferior arm. You might end up with a smaller percentage of your participants randomized to the inferior arm of the trial, but a larger absolute number on trial. I want to point to a paper I co-authored with Susan Ellenberg 9 in Clinical Trials last year laying out some of these concerns.
Relevant to the balance of benefit and harm, it is of course true that at any given moment the randomization probability in an adaptive trial is going to reflect your expectations of relative benefit, given what you know at that time. But as I mentioned a moment ago, it does not necessarily follow that at the end of the day, you’re going to have fewer participants randomized to the inferior arm than you would have had in your best fixed ratio alternative, and importantly, it also doesn’t follow that more patients outside the trial will benefit because you’re going to declare efficacy earlier if, in fact, there is efficacy. So, although from the perspective of that individual participant about to be randomized, response-adaptive randomization might be of benefit, it’s not clear, to me at least, that there’s a benefit to the larger population, either of participants in the trial, or equally or even more important, the population of patients that you hope will benefit from what you learn.
Let me say a word about informed consent. If you think about what you might disclose to your participants as they are considering joining the trial, what might you say if you’re using response-adaptive randomization? Well you might say, assignment probabilities will vary according to the interim data. If it’s looking like one treatment is better than the other, we will adjust our randomization probabilities in that direction. Or you might theoretically go further and tell them, if you join the trial, right now the randomization probability is 2:1 or 3:1, or whatever it is, and we’ll actually tell you what your randomization probability is at this moment. Of course, this is particularly problematic in open-label trials. Why is that the case? Well, of course, if both the investigators and the participants know what those randomization probabilities are, it might affect the willingness of potential participants to enroll in the trial, or the willingness of investigators to recruit more participants. Second, it might lead to differential dropout, particularly in open-label trials. If I find out that the randomization probability is 3:1, and I’m randomized to the 1, I’m not so happy about it. I drop out. And then there’s the problem of potential distress if I find out that I’m randomized to the worst performing arm. So from the point of view of informed consent, and particularly in open-label trials, I think there are problems with response-adaptive randomization.
Finally, let me address the question of professional integrity with respect to the clinician investigators who are taking part in these trials. If one treatment is looking better than another and yet you’re randomizing 1:1, that feels bad. You really want to maximize the chances of benefit for your patients. But imagine for the moment that our randomization probabilities tilt all the way to 10:1, because that’s how the data are looking at the moment, you’re still randomizing 1 out of 10 or 1 out of 11 participants to the less preferred arm, and so the question is what do you say to that participant? What do you say to them, in fact, if it is an open-label trial and they know that almost everyone else is assigned to the other treatment? If it’s a blinded trial, the problem is hidden but hasn’t disappeared; how do you justify it to yourself if you know you may still be randomizing participants to something that has already been determined to be very likely inferior.
In summary, whether response-adaptive randomization has what you might call consequentialist advantages is an empirical question and that’s where all of your expertise comes in. Are fewer participants assigned to the inferior arm? Are resources used more efficiently? Is there a more rapid time to a decision about whether a treatment works or doesn’t work? What I want to suggest and what I hope I can add to this conversation is even if the answer to all those questions is yes, that does not necessarily resolve the ethical tension that I think motivates adaptive randomization designs in the first place.
Even though adaptive randomization is a very large field, I’m going to limit my comments to response-adaptive randomization. My first exposure to response-adaptive randomization was at my very first Society for Clinical Trials meeting in 1990, where the ECMO (Extracorporeal Membrane Oxygenation) trial 10 was a matter of great debate and identified as one of the two “most contentious adaptive trials ever to be conducted.” 11 In ECMO, the first infant was assigned to the control treatment and died, the next 11 infants were assigned to the experimental treatment and lived. These 11 weren’t randomized; they were assigned by the Play the Winner rule. 12 I’ve heard many arguments that it was an inconclusive trial, and my colleague, Dr KyungMann Kim (quoted with permission), says that this is where statisticians betrayed science by allowing this approach. The numerous design, ethical, and inferential issues related to ECMO have been amply covered elsewhere.10,13 I’d like to opine that the debate is a modern version of the older debate over randomization. There have always been (and still are) people who have denigrated randomization and say that it’s not needed or it’s less important. They are fortunately in the minority. In one of the comments on the ECMO trial, Berry 14 stated that “Randomization is not essential for scientific inference.” So I think that we can interpret response-adaptive randomization in that context.
I’m grateful to Frank Bretz for his exhaustive presentation and summary of all the references. That was quite impressive. I agree with him that response-adaptive randomization is a part of a multi-faceted approach for designing early phase trials. He didn’t restrict himself to early phase trials, but since I agree with him on early phase trials, suppose we are now considering a randomized trial. Because this is an example taken from current work I’m doing with industry, I’ll keep the description vague. We have a two-arm trial with a certain randomization ratio, call it 1:1, between a high dose of a drug and placebo; phase IIB is dose ranging, and we have a different randomization ratio between the same treatments, call it 2:1, so you have 1:1 in IIA and 2:1 in IIB. Now, the bias-proof way of combining those is via analysis of variance (ANOVA); you do a stratified test where you have one block with a certain randomization ratio, another block with the other randomization ratio. The cost of doing that is lower efficiency. This analysis is less powerful than ignoring blocks, a fact that can easily be shown either analytically or by simulation. Also, since the randomization ratios in some of the proposed scenarios are more extreme than 2:1, the costs are also more extreme and so we can’t afford to block. We just combine them, which violates the randomization, but since it’s a phase II trial that seems acceptable. But we need to be aware of both the benefit, that we can directly compare outcomes from IIA and IIB to increase power, and the price, which forces the assumption that patients and treatments are the same in IIA and IIB. We are thus making inference by combining randomized and nonrandomized comparisons: the most extreme case, where all patients in IIA receive one treatment and all in IIB receive another, would be strictly nonrandomized. I am not averse to sacrificing bias for the sake of power in phase II trials as long the source of bias is acknowledged. However, for confirmatory trials, I strongly affirm that the point of randomization is to get rid of the bias, and you should not take such liberties with them. For these, unblocked response-adaptive randomization is not appropriate.
Dr Proschan’s presentation was very succinct and informative. Leaving out early results is a pretty old issue. For example, in the old clot-busting trials with aspirin or other agents, the proposal was frequently made that we should leave out early events because everybody knows that if you take aspirin today, it won’t stop a heart attack tomorrow. The standard response to that is that poisoning people and getting rid of the frail ones early on may make your drug end up looking good. That’s not new, but I would like to warn you against what is a more subtle but still dangerous strategy, and that is of using rank tests with increasing weights. We have traditionally used constant weights via the log-rank test or Peto’s version of the Wilcoxon rank test with decreasing weights 15 —either is fine. For example, for prostate cancer, as time goes on you’d expect older men to die of heart attacks and so on, and so you prefer a decrease in weights. That means that later deaths are less important. That’s always been my view. However, what you don’t want is for early deaths to constitute an advantage—for example, by using a rank test with increasing weights. Suppose somebody asks you to analyze data using such a test but then, after you have done so, reports a data error: instead of dying at 10 years, a patient died at 1 year. You find that because his death was weighted less, that improves the results for his treatment. Why should that be? And yet with increasing-weight rank tests it can be. I’ve been told by a reliable source that the Women’s Health Initiative was initially designed with such an increasing weighted rank statistic and it was then rejected. They used the log-rank statistic instead. The reason for the initial choice was that hormones were not thought to influence women’s risks of various diseases immediately. The reason for the change is the same as above.
Simon and Simon 16 point out that “Under some conditions in which the prognosis of future patients is determined by knowledge of the current randomization rates, the type I error is not strictly protected.” That means that blinding, as has been pointed out, is important. If I know what treatment I want and I wait to get it, then of course that’s not a randomized trial; even the nicest permutation test presented by the Simons cannot correct for that, so we still have to be very careful. The only foolproof bias–proof method is blocked randomization, and the first place I can see that referenced is in Jennison and Turnbull’s book. 17 Karrison et al. 18 give a simple example.
With regard to Peter Thall’s presentation, he is preaching to the choir. The problem with a time trend can be summarized by what I heard at my first Society for Clinical Trials meeting about the ECMO trial. 13 Paul Meier, whose opinion I then and always have since adhered to, was sitting next to me. He leaned over and he whispered to me “The real problem is the time trend.” I have been whispering it (or perhaps yelling it) ever since and wondering why has it taken so long to be observed?
I have gone through many papers on response-adaptive randomization, and I reviewed their simulations. I looked for scenarios in these simulations in which there were time trends, but there were almost none. In fact, major papers (e.g., on the BATTLE study, a large trial on personalized therapy for lung cancer (ClinicalTrials.gov numbers: NCT00409968, NCT00411671, NCT00411632, NCT00410059, and NCT00410189)) by distinguished statisticians do not include time trends in the simulations. 19 But, based on my experience (and speculation):
The disease itself can change over time, sometimes radically (e.g. AIDS in the early 1990s).
The definition of the disease can change, due to new scientific discoveries of diagnostic methods (e.g. stage migration).
The trial inclusion criteria can change, either formally (in which case we can stratify analysis on before versus after the change) or informally due to “recruiting zeal” or other issues (in which case we cannot).
Centers can change, such as when VA centers enter subjects into a trial earlier or later than academic institutions.
Subjects within centers can change, especially but not only with chronic diseases, due to the phenomenon of a “queue of desperate patients lining up at the door.”
In addition to these examples, an investigator who wants to game the system could cross his or her fingers that his favored treatment arm is ahead, then progressively enroll better prognosis patients over time.
Finally, in case you thought that the problems with response-adaptive randomization due to time trend demonstrated by Peter Thall were due to the nefarious mechanisms of Bayesian statistics, there’s a very simple example (first contrived by Karrison et al. 18 ) where the randomization ratio changes and patients somehow do worse later or better later, and therefore, you come to completely false conclusions.
In conclusion, I agree that the Swiss Army knife analogy is a good metaphor for response-adaptive randomization. It’s useful in some cases, it’s useful in small trials, but for confirmatory trials, I think we had better stick to a fixed randomization ratio.
As I understand it, there’s a rule of thumb that once you go beyond 3:1 or 5:1 randomization in response-adaptive randomization, diminishing returns set in. Should adaptive designs take that into account?
Beyond a ratio of 5:1, there is not much of a gain in power. To me the issue is not the randomization ratio itself, but varying the randomization ratio. If it varies among centers, that’s potentially dangerous unless you account for it.
I wanted to begin by completely agreeing with Peter Thall about adaptive randomization. If you use really foolish response-adaptive randomization rules and don’t vary them over a wide range, you get bad results. If I wanted to illustrate that response-adaptive randomization is always a bad idea, I’d have done precisely what Dr Thall did. He started adapting the randomization ratio very early and very aggressively. This means that if the better arm starts out worse, as sometimes happens due to natural variability, you never go back to it; there is no second chance. When using response-adaptive randomization you really have to consider the right time to turn it on. Dr Durkalski described the NETT earlier; we spent hours with one of their studies deciding when to start adaptive randomization and decided to start only after 300 patients were enrolled.
Also, we almost never lower the randomization probability in the control arm; we just let the randomization rates change over the various experimental arms (which might be different doses). The major bias issue Dr Thall mentions is a result of allowing the proportion assigned to the control arm to change. When you allow the control arm’s probability to be adaptive, if the control looks bad, you randomize away from it, so when the probability gets low, you get fewer data and it has no chance to regress to the mean. So now you have large bias. Instead, if you take the arms that aren’t doing well and reassign those patients to the control as well, a lot of that bias goes away. In the ESETT 7 paper (a three-arm comparative effectiveness trial in status epilepticus), we actually show power is higher, sample size is lower, and more patients are randomized to the better arm. I agree that there are some cases where it’s bad, but there are cases where you get a win-win-win with better power, lower sample size, and more patients on the best treatment.
I agree with Dr Chappell that drift is a key issue. We discuss this in a recent paper 20 where we re-analyzed the Antihypertensive and Lipid Lowering Treatment to Prevent Heart Attack Trial (ALLHAT)21,22 from the 1990s as if it were a Bayesian adaptive trial with response-adaptive randomization. We did this without knowing the results. There was enormous drift in ALLHAT. We got the real data from Dr Barry Davis, the lead statistician on ALLHAT and one of our coauthors, and the first thing I did was look at trend overall and saw enormous trends. I thought that the paper would be awful, because we’re all taught that response-adaptive randomization breaks down with extensive drift. But our estimates were very similar to ALLHAT’s.
Also, Dr Thall was critical of the idea of identifying areas of potential weaknesses of response-adaptive randomization and then making little adjustments to make it work better. To me, finding what doesn’t work well and making improvements is how science is supposed to work.
I’m looking forward to reading the ALLHAT paper. I have been hearing presentations and reading papers for 25 years on response-adaptive randomization, and (with the exception of Dr Thall and his group’s recent ones2,3) almost never have they included time trends in their simulations.
Dr Durkalski, in discussing Dr Proschan’s presentation you mentioned the regulatory considerations around changing a primary outcome. Rather than formally requesting an outcome change, wouldn’t it be just as beneficial to do a post hoc analysis with this new outcome? I’d guess that the FDA might require a new trial anyway with the new outcome that you’ve created.
In the trials I have been involved in, even though some unplanned things occurred, they have not necessarily been related to changing the outcome variable. I do worry about making such a change, especially in a confirmatory setting. By the time you start a confirmatory trial you should have information from previous studies to really know whether this is a reliable outcome and whether you can actually measure it. However, things happen, so although your suggestion of doing the post hoc analysis may be acceptable in some settings, it may not be if you are trying to get this treatment approved. I cannot comment on whether FDA would require you to do a new trial.
There is a perception issue of non-validity whenever you change something in the middle of a trial. I would argue that when you do it in a completely blinded way, it is as if there were pre-specification. But since the perception exists, I would do it only in emergency situations.
Back to the issue about response-adaptive randomization and drift. That can happen even with covariate-adaptive randomization, not just response-adaptive randomization; we talked about a particular case in a paper with Proschan et al. 23 The beautiful thing about re-randomization tests is that they automatically take drift into account and provide a valid p-value. Whereas, in simple settings, the t-test and the re-randomization test are asymptotically equivalent, that’s not true in these more complicated settings and that’s why you really need to do a re-randomization test to protect against potential temporal trends.
Most of the treatments that we study in rehabilitation are very hard to do in a blinded fashion. They are either experience-based interventions like gait training or exercise, or they’re device treatments that have visible attributes. I understand the problems associated with open-label studies and response-adaptive randomization. Would you say one shouldn’t use it for unblinded studies or are there any scenarios in which unblinded studies could safely be conducted this way?
From an ethical perspective, even if there are these consequentialist benefits from doing adaptive randomization, it has not solved all the ethical problems. But I also acknowledge that the question of whether there are, in fact, efficiency benefits or validity or precision or early stopping or those sorts of benefits is an empirical question. I find it difficult to imagine in an open-label trial, how you wouldn’t in essence be running a very grave risk of sabotaging the trial through bias if you were to adaptively randomize.
I agree. I think there would be a risk of selection bias. Maybe if you are doing a very large study with many participating centers, it wouldn’t be as obvious, but again, that depends on the treatment effect and how much you are varying your allocation ratios. You would have to tread carefully in an open-label trial.
I participated in a trial of adolescent idiopathic scoliosis, which is a back deformity in children, testing use of a special brace. 25 The trial had a randomized and a “participant preference” cohort. The primary analysis of the control versus the brace approach in the randomized cohort was intention to treat and the reported odds ratio (OR) was 1.93. Also presented was an as-treated analysis using the combined randomized and participant preference cohorts (reported OR: 4.11). While the OR for the preference cohort alone was not reported, these observed differences argue in favor of true randomization and echo Dr Durkalski’s comments. I don’t see how you could prevent bias in response-adaptive randomization unless you block randomize with the consequential sacrifice in power.
I think that everybody enjoyed Dr Bretz’s analogy of the Swiss Army knife. Listening to the talks this morning, I thought about another analogy. “When you have a hammer, everything looks like a nail.” It looks like for response-adaptive randomization one expects that method to be the best design for any, or many, objectives. I think the consideration should be reversed. You look at the objective of the study and then find the optimal design in order to address the objectives. If there are a multitude of objectives, you would need to specify the optimal design under a compound objective.
My second comment about having a method that we want to apply to solve all our problems is about bias. I’m not talking about the operational bias that we just discussed, but rather statistical bias. I think in the original and in almost all definitions of adaptive designs, there is always a requirement that the analysis should be valid. If you use an adaptive design, it’s well known that the naive estimate of the parameter will be statistically biased. But that doesn’t mean that the adaptive design is not valid and should not be used. There are methods for adjusting for the bias; I think Dr Thall mentioned methods that reduce bias, but there are even unbiased adaptive estimates. I think that we need to be careful about making strong statements about response-adaptive randomization in general based on just one such method.
Dr Joffe raised an interesting point about the potential for patient dropout because of being unhappy to have been randomized into a less efficacious arm. I wonder whether missing data techniques can be used in adaptive clinical trials to address this question.
In open-label adaptive randomization trials with potential drop out and so on, if drop out were at random, missing data techniques would be useful to combat selection bias. Since the question is more likely non-random missingness, are there any techniques for imputing data missing not at random that are available in this context? I’m skeptical, but do others want to address the possibilities here?
As somebody who works in both clinical trials and on analytical methods for non-ignorable missing data, I can offer a comment. There are, of course, lots of different techniques that we can use to attempt to impute, but the validity of our result depends heavily on all the assumptions we make about the non-ignorable mechanism. So, to the extent that we’re happy with our assumptions then yes, technically we could apply some of those techniques. But in some sense I think we’re just adding to assumptions based on our ill-founded beliefs where we think something’s going on and then all too frequently it turns out that something else is going on. So application of these techniques could be very risky, in my view.
I think the link made between missing data and response-adaptive randomization is an important one. A method that could be used is principal stratification, as suggested by the National Research Council 26 report, with similar cautions as indicated by Dr Troxel.
Dr Durkalski has talked about missing data being a particular issue in adaptive trials. Can you provide some more detail on what the differences are? Let’s assume we’ve got a blinded trial, so we don’t necessarily have the selection bias issue, but what are the special issues of dealing with missing data in an adaptive trial versus, say, a sequential design trial?
Missing data gives us enough headaches; we don’t like missing data in any setting. My specific comment was based on an experience we had when we were setting up our response-adaptive algorithm. We had the “oh my” moment with several challenges to solve when setting it up. If you are using response-adaptive randomization and you have missing outcome data, you can’t throw those people out, but what is the best way to impute outcomes? Multiple imputation is highly recommended, but there are specific parameters that you have to set when you develop your multiple imputation: how many imputations you do before you get consistency in your result; the covariates you will use to impute those data. These decisions could affect the performance of your response-adaptive randomization so you need to consider this during the design stage. Usually when we’re dealing with multiple imputation, or any kind of imputation method, it is at the end of the trial or during an interim analysis and we have the opportunity to do sensitivity analyses. We may look at worst-case or complete-case scenarios, just to make sure our results aren’t changing based on how we impute the data. If you are trying to do this in a response-adaptive setting to update your allocation ratio and if you have frequent updates to your data, or even if you don’t have frequent updates, you have to think about these issues and how to incorporate and defend them at the end of the day when you’re reporting your results.
Dr Joffe asked if you have response-adaptive randomization and end up in a 10:1 scenario, what do we tell the singleton in an ongoing trial? I would argue that there is little difference between his question and the scenario in a fixed randomization trial: before the results are in, but in the face of accumulating data, what do we tell the 50% who have to get the inferior treatment? To me, fixed randomization is not better, or more ethical, just because we tell more people we gave suboptimal therapy.
I would agree with that. I wouldn’t want to argue that it’s better. I think in either case, you have data at a point where you’re at or close to an answer. There might be a 10% or a 50% chance that you’re randomizing someone to a treatment where we’re getting pretty confident that it is less good than the other treatment. The issue here is you set out your rules for when you stop the trial, either at the end of the trial or at some interim analysis, and you have to be comfortable with those stopping rules. You have to be able to defend those stopping rules, and I don’t think there’s any way to make this problem go away. You’re doing a randomized trial, there’s always going to be some last patient who’s randomized to the less preferred arm of a trial. It’s a generic problem that you can’t make go away with outcome-adaptive randomization. The question of whether there are going to be fewer patients at the end of the day who are randomized to the less effective treatment is an empirical one, which is for you all to work out. But there’s just no way, with or without outcome-adaptive randomization, to make this issue of the last participant in the drug trial go away.
There is a trade-off between individual ethics (the patients in your trial) and collective ethics (the good of science). If you are getting close to an answer about the various arms, you might consider doing something like switching the randomization probabilities, or you might tell the next patient which arm is doing better. But this will compromise the results and then the medical community at large may not accept the results, and so you’re doing harm that way. I do agree that the primary responsibility is individual ethics owed to the patients in the trial.
For the clinician investigators who sit across the table from the patients, it’s probably true that the primary responsibility is individual ethics, but if you’re the principal investigator or the statistician who’s designing the trial before there’s ever an individual participant, I would actually argue that the primary responsibility is to get to a valid answer to the question that you’re asking. When I presented the Emanuel et al. 8 framework, the first substantive question is, does it answer a question of social value. The second question (before we get to risk–benefit for participants, informed consent, or anything else) is scientific validity. It was very clear that if you answer those questions in order and if the answer is we’re having concerns about scientific validity, then don’t move on to the other things. Fix the scientific validity concern.
As far as the blinding, I don’t think the selection bias goes away just because you have blinding. For instance, I was involved in one trial where we found that a number of our patients were getting their pills analyzed chemically. We didn’t know this for quite a while, so blinding is not perfect!
If you’ve got a disease where you have some control over when you present for randomization, there’s likely to be advantages to waiting until later in the trial and that may change the characteristics of the patients, even if it’s a blinded trial and you don’t know what the randomization probabilities are, or which treatment you’re going to get.
There are many situations where much missing endpoint data is due to delayed responses, and so you have trouble getting information for adaptation. At the same time, the enrollment of trials tends to accelerate over time. The group sequential approach still gets you early decision making with more flexibility for delayed responses, but doesn’t allow you to adapt.
If I understand Dr Anderson’s point about delayed responses correctly, an example of this situation would be a trial in which there is central pathology review. You might have a site-based diagnosis of a heart attack, and the electrocardiograms (EKGs) are sent for central review. Some months later, you have the adjudicated answer for your primary end point. Meanwhile, the Data and Safety Monitoring Board is wondering what’s going on. First of all, you have preliminary data on the accuracy of the locals’ diagnoses. If they’re 100%, then you can keep going. If it’s very low, you wait for the adjudicated data, but in between there’s some interesting statistics to be done; the issues and assumptions are similar to any other kind of missing data, and if the assumption of missing at random make any sense, you may be in the clear.
I think for the majority of studies, it is common practice to incorporate some sort of early stopping rules, especially in the confirmatory setting. As far as including other adaptations beyond those interim looks, I guess it goes back to all the issues discussed today that we still need to iron out. If you have a delayed response, is there a surrogate response that can be used for the response-adaptive randomization, or should we really be waiting for the primary outcome? Surrogates are tough, and if a surrogate is an early time point of the true outcome, you have to have some data to justify that you are truly using your most reliable outcome.
Participants
Footnotes
Acknowledgements
The authors thank The Center for Clinical Epidemiology and Biostatistics in the Perelman School of Medicine at the University of Pennsylvania, Genentech (A Member of the Roche Group), Johnson & Johnson (Janssen R&D), and Merck.
Declaration of conflicting interests
Dr Jason T Connor is an employee of Berry Consultants, LLC which specializes in methods of innovative clinical trial design, including adaptive designs. No other conflicts were reported.
Funding
External funding for this conference was provided by Genentech (A Member of the Roche Group), Johnson & Johnson (Janssen R&D), and Merck.
