Abstract

Our three speakers this afternoon covered a wide variety of adaptive designs. One topic that came up in all three talks, as well as in the talks this morning, is that adaptive trials need high-quality data during the course of the study. I will address this issue in my closing remarks.
From Max Parmar, we heard about multi-arm, multi-stage studies, which are really exciting. He addressed the important topic of concurrent controls. One feature that I’ve seen in certain adaptive designs is the failure to use concurrent controls.
Bruce Turnbull’s intriguing slide listed controversies (Table 1). I suspect we each placed ourselves in either the left or the right column. Dr Turnbull talked about Data and Safety Monitoring Boards (DSMBs) in the context of adaptive designs and gave examples where the design didn’t work as expected. Those of us who sit on DSMBs know that in many trials (maybe even in most trials), things don’t proceed as expected. Dr LaVange described the kinds of adaptive designs the Food and Drug Administration (FDA) has seen. She discussed the differences in philosophy and experience among the FDA Centers (Biologics, Drugs, and Devices). Her description of what the FDA means by “better understood” and “less well understood” designs was very helpful to those of us who have puzzled over that language in the FDA guidance on adaptive designs. 1
The current controversies.
Many adaptive trials base the adaptation on a surrogate endpoint. If that surrogate shows a promising result, the study continues to the end. On the other hand, if the surrogate is not indicative of likely clinical benefit, the trial may stop for futility. Many people have pointed out that surrogates may seem promising at an interim analysis, allowing the study to continue, but the intervention shows no benefit on the clinical outcome at the end of the trial. Less common is the case when the surrogate shows no benefit, but if the trial continued, the clinical outcome would show benefit.
Of course, if the trial stops when the surrogate shows no benefit, one would never know whether the intervention would provide benefit on the clinical outcome. An interesting example of this possibility comes from a paper that was published online in the Lancet a few days ago 2 describing a randomized double-blind trial of a cell therapy for ischemic heart failure. The trial could have been perfect for adaptation. Because cell therapy is such a new type of intervention, not much is known about it, and to do a large clinical outcome trial is expensive. The study was a Phase IIB trial, so it’s not formally “confirmatory.” The intervention was designed to remodel the heart in order to make it pump more efficiently so that, theoretically, the patient would be less likely to experience a clinical cardiac event. The investigators designed a standard trial, continuing to the end to assess the effect on the occurrence of clinical outcomes. But they could have designed this trial in an adaptive way. They could have looked at the extent of remodeling partway through the trial. If the drug did not remodel the heart sufficiently, the protocol could have specified how to change the dose, or whether to declare futility and stop the trial. In such an adaptive design, only if the drug performed as expected on the surrogate endpoints (i.e. those measuring remodeling) would the trial continue to the end and be evaluated on the basis of the primary clinical outcome. What would have happened if the investigators had designed the trial to adapt to interim assessments of remodeling? In fact, at the end of the trial, the data showed no evidence of remodeling. Neither the ejection fraction nor the measures of dimensions of the heart changed, so the trial would likely have been declared futile (or the dose too low). Instead, the investigators followed the participants to the planned end of the trial. Interestingly, the trial showed benefit on the composite clinical outcome in spite of the fact that the intervention appeared to have essentially no effect on the surrogates. This is an example of a trial that would have gotten the wrong answer had the investigators done what I probably would have recommended that they do: look at the surrogate and then pull the plug if there is no evidence of benefit on the surrogate. Had they done that, they probably would have killed this product and may have dampened enthusiasm about cell therapy in general. I don’t think we know why the surrogates failed, in this case, to predict what was going to happen to the clinical outcomes. Maybe the putative mechanism was incorrect, maybe the measurements of ejection fraction were inaccurate (which they often are), who knows?
I do want to make a general point that all of today’s speakers addressed implicitly or explicitly: interim data are notoriously inaccurate. There is a good reason for that inaccuracy: it is really hard to collect accurate data in real time. There is also a bad reason: many groups who actually do data collection and are in charge of putting interim data together don’t really know how. I will give some examples. They are not exactly true; the drugs and disease are somewhat changed, but they retain the structure of true situations.
For the first example, imagine that the outcome of a trial in pulmonary arterial hypertension was worsening lung function. And imagine that someone was to be declared to have experienced worsening lung function only if the worsening was confirmed 4 weeks after the initial measured worsening. Finally, imagine that a Clinical Endpoint Committee was responsible for confirming the outcomes declared by the investigators. Our group was asked to look at the data before the sponsor unblinded the study. We noticed that the committee called 100% of the events it received positive. We found that surprising; usually such committees reject some events they review. So we looked through the data and found some other cases we would have sent to the committee for adjudication. For example, some patients had dramatic declines in lung function but died before they were scheduled to have their lung function confirmed. That was not called a possible event because the decline had not been confirmed. Because the trial had already ended, we were able to send these cases to the committee. They adjudicated most, but not all, of these cases, as events. Had this been an adaptive trial, the adaptation would have been based on an incorrect number of cases. Thus, in an adaptive trial, it is really important to count the number of cases precisely during the course of the trial.
Here is another example, this one from an adaptive design. We, the DSMB, on looking at the report we received, couldn’t figure out how many deaths had occurred. In all, there seemed to be fewer than 20, but the number of deaths seemed to depend on what page in the report one looked at. We as the DSMB had excruciatingly specific rules of what we were supposed to do to follow the adaptive plan. We told the sponsor that if we couldn’t even count the deaths, we couldn’t trust the data enough to follow their rules for adaptation.
Here’s another slightly modified example from a study of pain. The outcome was a visual analog scale. Respondents had to report a number on an 11-point numeric scale that ran from “no pain at all” (0) to “excruciatingly painful” (10). The form was translated into 60 languages, and in some of the languages, the translation was backward. Thus, a 10 would mean “excruciating” in most languages but “no pain” in others. This was correctable because we had the words and the score; we picked up the problem because the responses were wildly different in different countries. Again, these errors can be identified and corrected at the end of the trial, but it is much harder to correct them during the trial, which is when the adaptations must be done.
And these examples don’t even touch issues of faulty randomization, the use of the wrong randomization code for interim reports, or labs with different units. As far as the labs, it’s often easy to tell there are problems when the different units produce non-overlapping distributions, but it can be impossible to disentangle results when the units are overlapping. Other issues are under-reporting of adverse events in certain countries and non-proportional hazards when the adaptation is to be based on a Cox model. Clear understanding of these types of data is often part of what the adaptation requires.
What should one do in planning to design an adaptive trial? First, as Max Parmar has stressed, be reasonably certain that the gain adaptation is likely to bring is worth the machinery that is necessary to make sure the adaptation will be done correctly. Second, as Dr Turnbull has warned, ensure that everyone who needs to be part of the adaptation is on board with, and understands, the design. Third, as Dr LaVange has advised, check that the regulators are going to accept the method. Fourth, be confident that whoever is presenting the data is meticulous in preparing the report and truly understands the data because once a decision is made, there is no going back. Finally, if the adaptation is based on a surrogate outcome, the underlying relationship between the surrogate and the clinical outcome should be well-understood.
I want to tell you about I-SPY 23,4 and use this trial to illustrate some of the issues. I-SPY 2 is an adaptively randomized Phase II trial that was initiated in 2010; it is designed to identify active drugs for breast cancer in women who are early in the disease process and to speed drug development in this area. The idea is that if you test drugs at the time a woman is diagnosed, before she goes to surgery, you can see the effect of the drug on the tumor, but you also get important information at the time of surgery. You hope to see a pathologic complete response, which has been debated as a surrogate for long-term outcomes such as recurrence and event-free survival and has been the subject of much discussion with FDA leadership. Our trial endpoint is pathologic complete response at the time of surgery.
From Figure 1 you can see that patients are randomized to receive a standard treatment, typically paclitaxel, combined with one of multiple other agents. Another drug, trastuzumab, is added for patients who have the human epidermal growth factor receptor 2 positive (HER2+) subtype. Numerous arms are being studied at the same time. Patients are biopsied before beginning treatment, and again 3 weeks after beginning treatment, and then, we collect tissue at the time of surgery, so we have lots of biomarker data. We have the opportunity to see who’s responding, how well they are responding, and what biomarkers might predict response. We also do magnetic resonance imaging (MRI) scans to noninvasively monitor tumor status over time; that’s important because those data are used to adjust the randomization probabilities. We worked closely with the FDA to design the trial. Our endpoint of pathologic complete response has been supported by a meta-analysis done in collaboration with the FDA. In regard to the issue that came up earlier, we have a very active DSMB that meets monthly. It’s not that common for a Phase II oncology trial to have a DSMB, particularly one that’s so intensely engaged, but because of the adaptive nature of the trial, it’s very important for our DSMB to be involved in making decisions about what happens to the study arms.

(a) The steps in the adaptive randomization process used in this trial. The longitudinal model refers to the course of the patient through the neoadjuvant therapy, as measured by serial magnetic resonance imaging (MRI) scans. (b) The schema for the experimental-therapy group that received neratinib and for the control group. After screening, patients with human epidermal growth factor receptor 2 (HER2)–positive cancer were eligible to undergo adaptive randomization to receive neratinib plus paclitaxel. The control was trastuzumab plus paclitaxel. Patients with HER2-negative cancer were eligible to be randomly assigned to receive neratinib plus paclitaxel; the control was paclitaxel alone. Patients with HER2-positive cancer or HER2-negative cancer then received standard treatment with doxorubicin and cyclophosphamide to complete their neoadjuvant therapy. (c) The details regarding the screening, randomization, and treatment of the patients. Patients were categorized according to whether they received no experimental therapy or at least one dose of experimental therapy. 3 (Permission for reprinting requested.)
It’s also important for the DSMB to be involved because of safety issues. These are potentially curable patients getting new drugs and we have to be very careful about that. Since 2010, we have introduced 11 different drugs into the trial from six different drug companies. We’ve been able to look at multiple drugs from different companies simultaneously, which isn’t something that you could typically do outside of this kind of a platform trial, and we’ve “graduated” or found efficacy in five of those arms with biomarker signatures that help us understand who would respond and how. The study works by using the biomarker, imaging, and pathological response data that we collect on each patient as they go through their randomized treatment course to adapt the randomization probabilities for the next patient; it is a continual process of learning, a continual process of adapting the probabilities for the next patient who is coming into the trial to determine what arm they will go onto.
Over time, different agents will graduate or leave the trial for other reasons and new drugs can come on, so this platform approach is very efficient because we don’t have to start a new trial every time we want to test a new drug. We have an Agents Committee that selects and prioritizes these drugs, and we have an external review board that oversees the process. We’ve published on the safety criteria that we use in considering drugs for inclusion into the trial. 5 This is a complex process with a lot of moving parts; using adaptive designs in the real world requires addressing and managing a lot of logistical issues. We affectionately call our standing platform the I-SPY 2 Randomization Engine. As patients enter the trial, we do have continual update, continual application of that data into the longitudinal model, and continual update of predicted probabilities for each experimental arm.
As these predicted probabilities are updated, we have rules that govern the disposition of a treatment arm. A treatment arm will “graduate,” meaning it will have demonstrated efficacy, if it meets the criterion of having a greater than 85% predicted probability of success in a subsequent confirmatory randomized Phase III trial. Now you can imagine that that is an endpoint that most oncologists are not used to seeing and that can be tricky because we’re trying to convince people that we have really learned something meaningful about this drug that will make it worth taking forward to Phase III. We have spent a lot of time educating drug companies about the efficiencies of using this type of design with just a fraction of the patients that they would normally need to use in Phase II to make a “go, no-go” decision. It’s also possible that an arm will be stopped for futility, so if after a period of time, there’s a less than 10% predicted probability of success in a randomized Phase III trial, that drug or arm will come out of the trial for futility. Now, of course, toxicity is being analyzed as we go through this process and we have, in fact, had a drug that came out of the trial because of excess toxicity.
The other issue in response-adaptive randomization touched on earlier today is that you will continue to go back to certain drugs over and over again, and what if you’re not getting an answer? We have placed caps on the number of patients that will be accrued before we will say we’re not going to continue to accrue to this particular arm of the trial. So, in addition to stopping for futility, graduating, or stopping for toxicity, it’s possible that an arm will just reach its limit, which is for the overall trial, 150 patients, and it will stop without having reached any of these criteria because it simply isn’t going to and enrolling more patients will not likely give us that answer.
There are practical challenges in the operational aspects of this trial at many different levels. We spend a lot of time engaging with pharmaceutical companies to think about whether response-adaptive randomization is going to be a good way for their drug to be tested. Companies need reassurance that their drug won’t be compared to another drug in the trial; every arm is only compared to the control arm. It is also essential that the companies understand that we are not sharing data between companies. Such arrangements between the trial and the companies who entrust us to evaluate their drugs are very important. They require that we set limits on what kind of comparison we’ll make and what we set as the threshold for success. If we miss an effect, that has the potential to derail a drug, and that possibility has, at times, made it difficult for companies to decide whether they want to put their drug in the trial.
There’s also the issue of who knows what and when they know it. Using response-adaptive randomization means that the arms that are performing better accrue more patients, so if an investigator knew how many patients were on a particular arm, there would be great potential for bias. In I-SPY 2, I can’t even know when a drug has completed study until every patient who’s been assigned to that drug has gone all the way through the trial and completed it. Maintaining firewalls between who knows what and when they’re allowed to know it is really a logistical challenge, particularly when a drug is graduated and some people know that and other people don’t. We have data release guidelines that help define when we will tell the company, when we will tell the investigators, and when we will tell the world.
In this kind of trial where the investigators are limited in what they’re allowed to know, and there are clear rules about what will happen to a particular arm, the DSMB is essential to making sure that we are following those rules and making those critical decisions. The DSMB is also monitoring our toxicity on a monthly basis and they can also recommend terminating an arm based on toxicity. Engaging the DSMB and having the DSMB really guide what we do is something that has never been questioned by the investigators in I-SPY 2, but the disconnect between the discussions that the DSMB is having and what the investigators are seeing on the ground can make things difficult. Dr Wittes argued that we want to collect the data in the best possible way for them to be making these decisions, but there are subtleties, as you who’ve worked with clinical trial data know, that may not become apparent and can’t be necessarily discussed because of the firewall between the DSMB and the investigators on the trial.
There are also practical challenges at the site level, both for site investigators and for the site staff who are managing the patients, relating to the issue of bringing in data in real time. This is not the typical way clinical trials are conducted, particularly in oncology, where data usually come in quite late. There are many data queries; it’s not until the end of the trial that the data are fully assembled. We needed a change of culture at the 20 sites that are participating in this trial, to really collect the data in real time so it can feed the randomization engine, and to have real time data on toxicity. In most trials, we collect an enormous amount of data that we never use. It’s incredibly expensive to collect and clean all of this data, so responsive-adaptive randomization has actually made us very lean because if we’re going to get this data in real time, it’s got to be limited to just the essential data elements. We’ve got to be able to source verify it and clean it and be comfortable that it is absolutely accurate, in real time. The process of doing this in an adaptive way has had a downstream impact on how the trial is conducted at a site. I think that this is really good for the data because we’re not waiting 2–3 years to go back and try to figure out what happened to a patient. We’re figuring it out right now because we’ve got to know in order to use that data.
But again, there is a certain amount of information that site investigators and staff just can’t know, so a drug will reach that point of graduation but there will still be a group of patients randomized to that arm who are continuing to be treated, and site investigators don’t know that. They don’t know that a drug has graduated because if they did, that would bias them and we have to then update our predicted probabilities when the very last patient has gone to surgery, so there can be a disconnect in terms of how many patients are in the pipeline at the time that the drug reaches the graduation threshold. Those patients need to move through all of the treatment and go to surgery, and at that time then, we take that data and we update those probabilities; sometimes, those probabilities fall and that’s a risk that we take, but the modeling has really helped to minimize that issue.
And finally I want to address the issue of safety data and sharing those data with the sites. If you can’t know how many patients are on an arm, then you can never know the denominator for the number of events that are occurring for different toxicities, so we can only share aggregate data and percentages. We can share if there’s been a serious adverse event. The data do, of course, go to the DSMB, and if they see a safety issue, they’ll alert us and we can share that with investigators when a drug is withdrawn from the trial for toxicity. In my role in the trial, I was the first to hear about a severe toxicity because investigators would call me and say, I’m seeing something I wasn’t expecting to see, I’m seeing something that I’m really worried about. That underscores the importance of the real-time data collection as the sites are completing the case report forms. The site staff are trained to look for the toxicity because they know they’re going to have to record it immediately so that safety issues can be identified very quickly.
I also treat patients on this trial and at any given time, I don’t necessarily know what arms are active and what arms aren’t active, which can be awkward when I’m talking to a patient about the trial and they say to me, well what drugs might I get? This is an open-label study. They will know what they’ve been randomized to, but I can’t accurately tell them what the possibilities are. I can tell them what drugs are active in the trial and I can try to explain the possibilities. They are reassured by the lack of blinding—they will know what they are getting—but it can be difficult both to explain adaptive randomization to patients and also to not be able to tell them precisely the drugs to which they could potentially be randomized. We have engaged a group of about 70 patient advocates to help us make sure that patients understand what we’re asking them to engage in and they are very involved in our monthly calls. When issues come up, particularly regarding toxicity or when an arm graduates and we need to go back and provide that information to patients, our patient advocates have been very important in guiding how we convey that information in a meaningful and understandable way.
It has been an incredible experience to work on I-SPY 2 and I do believe response-adaptive randomization is the right way forward. We’ve had a very productive relationship with the FDA as we now start to design confirmatory trials for these drugs. But adaptive randomization is not easy; it’s not for the faint of heart and there are lots of lessons yet to be learned.
Let me start by noting the great irony in the field of biostatistics: our job is to identify the newest, greatest technologies or identify whether the newest, greatest technologies are, in fact, great. We are asked to identify and quantify how well they work, are they safe, are they efficacious? But at the same time, there are many statisticians who believe the technologies we use—meaning the statistical methodologies—were as good as they could get by 1933, or maybe up to 1977 with the development of group sequential designs, and they adamantly refuse to believe that there is much room for improvement in statistical or clinical trial technologies. Why don’t we experiment with how we design, perform, and analyze experiments?
To this point, I will summarize professor Donald Berry’s remarks at the kickoff meeting of an adaptive trial for glioblastoma multiforme in November 2015:
The first randomized clinical trial (RCT) was conducted 70 years ago in England. It was revolutionary. It changed medical research from case study and anecdote into a real science. The RCT itself has remained unchanged since that time. This resiliency is a testament to the esteem in which it’s held in the research community. But such stasis indicates complacency that is unusual in science and technology in the modern era, which is all about change and advancement. What other technology hasn’t changed in 70 years? Meanwhile, cancer biology is advancing with lightning speed. Biologists are providing a burgeoning pipeline of potential cancer therapies and these must wait their turn for evaluation in a clinical trial in an ever-lengthening queue.
Consider the constraints we faced when the statistical techniques we use were invented, such as the lack of computing power. Those constraints aren’t there anymore, but often we think that those well-understood techniques are all we can use, or we’re afraid to use new things.
Let’s revisit the topic of this meeting: “Adaptive trials, where are we?” The fact is that industry is way ahead of academia. In part, that’s due to the rules of the game: industry cares about results and also opportunity costs, meaning if a trial is going to fail, a company would rather it fail early. A sponsor can take that time and that money and reinvest it in a new trial or on a different drug. Many of our biggest efficiencies from adaptive trials have been in futility stopping. Private companies, at least ones with more than one compound, understand this and therefore realize it’s worth their while to think really hard about futility stopping rules in the design stage. As a result, I think industry is tremendously ahead in terms of adaptive trials.
In academia (and I don’t mean to be too critical here), a key consideration is getting funding. Many times when working with academic collaborators we suggest futility analyses, and the principal investigator does not want to incorporate such analyses. What’s unspoken is that if their trial stops for futility, they have to give the money back. All these people, site investigators, study coordinators are operating on “soft money.” This isn’t a criticism—it is merely a result of the different incentive structures between industry-sponsored trials and academic trials. Part of the game is to get grants and keep the very good people you work with funded. As discussed by DeMets and Califf, 6 in academic trials, there is much less incentive to find the answer rapidly. I agree with Dr Turnbull that it is critical for data monitoring committees (DMCs) for adaptive trials to have a statistician who really understands adaptive trials; someone who understands and buys into the fundamentals of adaptive trials, what is different about them and how they need to be executed. It’s also critical that they remember that the number 1 goal, just like in any clinical trial, is not only to protect patients and ensure their safety, but also to ensure that the protocol is followed, which means ensuring that pre-defined adaptations take place if they’re supposed to take place.
Just because the trial is adaptive does not mean that the DMC gets to make decisions that are not pre-specified. Sometimes collaborators have erroneously thought that running an adaptive trial means that then they or their DMC gets to figure out what to do next as the data emerge. An adaptive trial should be prospectively adaptive, with pre-specification of all the decision rules, timing of interim analyses, and decision thresholds. This is a critical point because there is sometimes a misperception that DMCs for adaptive trials have more flexibility to change the design—and that’s just wrong. In a fixed design trial, we would never just let the DMC members look at data and recommend what to do next based on their own intuition. We allow DSMBs to see completely unblinded data to protect patients and the integrity of the trial. By letting a completely unblinded DMC then decide what to do with a trial completely contradicts why we actually conduct blinded trials—any ad hoc change to the design, as well thought out by experts as it may be, introduces unquantifiable operational bias and also means that it’s impossible to calculate Type I and Type II error rates. The only way to understand error rates is to have prespecified rules and understand the probabilities of those rules under various real-life situations.
I somewhat disagree with Dr Turnbull that the DMC should have input into the study design if it is supposed to be independent of the study implementation. It seems that if the DMC plays a big role, it’s now a collaborator with the sponsor, whether that sponsor is National Institutes of Health (NIH) or an industry group.
Dr DeMichele presented I-SPY 2, a complex trial. We sometimes hear that trials should be simpler. I’ve heard the opinion that adaptive trials are just too complicated, and I believe that is a bit lazy. I completely understand how my bicycle works. I have no idea how an airplane works, but I flew here yesterday. I flew here because it’s safer, more efficacious, and more efficient. But of course it is dramatically more complicated. Airplanes are complicated, but we agree they’re the best way to get anywhere more than a couple of hours away.
The other concern about trial complexity is infrastructure. Again, the way trials are often conducted today is not fundamentally different from the way they were conducted 30 or 40 years ago. But that approach fails to take advantage of data technology in the 21st century. I can go on my phone right now and see real-time basketball scores in Greece. There is no reason not to use all the information we can as soon as we know it. We live in a world where it is possible that if an event happens in California today, we can get it adjudicated within a week and get it quickly to a data coordinating center. The infrastructure exists. If we’re not using it, that’s our fault. None of us would want our doctor to ignore valuable information when he or she is treating us—but we design trials to ignore data all the time.
Turning to the FDA, the perception of the FDA’s conservatism in regard to adaptive trials is dramatically overblown. I have so many examples of people saying, “oh, FDA would never permit that,” and I can show them a trial in which I have done—with FDA’s blessing—precisely what they claim FDA would never allow. We hear “the FDA doesn’t allow simulations that demonstrate type I error.” But they do in all three centers—Drugs, Devices, and Biologics—and not just for early-stage trials. The FDA is not monolithic, and there’s no such thing as “the FDA.” Different centers and different review groups often react differently. And most are more innovative than people expect them to be.
One key lesson I have learned when interacting with FDA and other key stakeholders is to show examples of trials. Examples let key stakeholders see how a trial will unfold—what data lead to which decisions; for example, what effect sizes would lead to earlier stopping, or what lack of effect would lead to futility stopping.
I think one reason industry is ahead of academia in implementing new methods is the pre-submission meeting at which a company can discuss its proposed trial design with the FDA review group. In NIH-funded trials, principal investigators are understandably hesitant to use new statistical methodology because the grant review group has to agree with the research question, research plan, and also this new methodology that they might not understand or might disagree with. If a single reviewer doesn’t like it, the grant may not get funded. Academics with a low score have to go back in another review cycle—where they’re likely to have new reviewers with perhaps new and different concerns.
In the FDA process, you design the trial, submit it to the FDA for feedback, and then you have a conversation with FDA to reach an understanding of the regulatory concerns. You may be able to address the concerns with a conversation, or you can come back later after making changes FDA suggests.
My favorite line of Mahesh Parmar’s talk was when he said that if you’re going to spend years (and implicitly millions of dollars) doing a trial, you should spend many months designing it, not draft a design at the last minute because a grant is due the next week. I understand that this kind of situation can be hard to avoid in academia and, again, I think it takes us back to the need for change in the incentive structure. We need more planning grants and we need to understand that there’s a huge benefit to thinking hard and long about trial design during the planning process. Imagine if NASA had put out a request for proposals for getting to the moon, letting any group describe their idea—most with little resources for the development of their idea—and then NASA had chosen one group’s design and gave them money to execute it. That’s a horrible way to get to the moon—but that’s how we are trying to cure cancer.
Now, I would like to turn to the idea of a “prospective postmortem.” (My collaborator, Dr Roger Lewis calls this “anticipated regret.”) During the design stage, all collaborators, from the doctors to the statisticians to the study coordinators who really understand how things work, to representatives of the patients who might enroll, should be asked, “Imagine if, once we’re at the end of this trial, the primary endpoint just barely fails to document efficacy. What might have gone wrong?” Then the group would brainstorm and write down everything they could think of. A lot of items on that list could probably be fixed upfront, but if there are still some uncertainties, particularly relating to the statistics, you may be able to build in adaptations and then monitor those concerns throughout the trial, and let the design adapt accordingly.
Mahesh Parmar talked about the “two-arm culture.” Siddhartha Mukherjee quotes Gracia Buffleben in The Emperor of All Maladies, “Dying people don’t have the time or the energy. We can’t keep doing this one woman, one drug, one company at a time.” 7 The solution is platform trials that allow us to study many drugs at once.8,9 For example, just a few months ago, the Wall Street Journal published an article noting that when multiple companies are researching treatments for a rare disease simultaneously, that can delay advances because they’re competing against one another for patients. 10 The concern is that even if one of the drugs is beneficial, it’s hard to show it because the trial can’t accrue enough patients when other trials with poorer drugs are competing for patients. By doing a platform trial in which many or all the drugs are included in the same trial, we can address this problem; using response-adaptive randomization, patients will tend to be assigned to the drugs that are working the best.
I will close with a message to students. As biostatisticians, we are public health officials and our goal should be addressing disease and producing optimal processes to address patient care. As Dr Joffe said earlier, that means balancing in-trial care with care of patients on the horizon—and to be good public health officials, we have to think about both the patients in the trial and the patients on the horizon. So I challenge the students here not to just think of error rates in a single trial. Perform simulations and calculate power and Type I error. But also, for example, for an infectious disease trial, calculate how many people are going to die both in the trial and on the horizon over a wide range of scenarios and potential trial designs. Calculate how many might get suboptimal treatments using a sequence of two-arm trials versus a strategy of one multi-arm trial. And consider what really is most important for public health. Is it Type I error? It doesn’t have to be. I challenge you to leave here knowing it is OK to experiment with the way we conduct experiments.
Dr DeMichele provided a description of the I-SPY 2 trial and pointed out that FDA has been a major player in that consortium from the early planning stages through today. Our Center Director, Dr Woodcock, co-authored a paper with the I-SPY 2 principal investigator on the trial design and importance of the collaboration the trial represents. 11 I do think that platform trials are a very exciting innovation in cancer research and have a lot of potential for getting safe and effective therapies to those patients who will most benefit from them in an efficient manner.
My favorite comment from the third speaker, Dr Connor, was that FDA is not monolithic. We joke internally about hearing that the FDA said this or the FDA said that and wonder who exactly is being referenced. To put this in perspective, the Office of Biostatistics in the Center for Drug Evaluation and Research includes over 200 statisticians spread across eight divisions, and we are involved in everything from generic drugs to biosimilars to pharmacologic and toxicology studies to drug abuse studies, and that’s over and above our traditional role of reviewing new drugs for cardio-renal, psychiatry, neurology, endocrinology, and many other therapeutic areas. It should come as no surprise that we do not speak with One Voice, although it has been a key initiative of mine to try and get clearer messages out to sponsors and other researchers about statistical policy at the FDA, that is, what statistical principles we rely on in evaluating certain trial designs and analysis methods. We try for clarity and transparency in our review work, so when you hear someone say what FDA will or will not accept, it is probably worth your while to make sure you are talking to the right people, namely, the reviewers who are actually involved in your drug development program.
The other comment I’ll make is that when you’re adapting, my own sense is to make big adaptations, don’t aim to make small adaptations. They can easily not come into fruition, and you can spend a lot of time doing quite small things in terms of your protocol, which are of no real value, albeit requiring massive calculations. That’s partly why we went for stop-go type decisions, either a new arm added or an arm to be stopped in STAMPEDE 12 and in FOCUS4. 13
My final point is that if we are assuming that an effect is going to be subtle, then we may also be assuming that if there is little or no effect on an intermediate outcome measure, then there will be little or no effect on the final outcome. However, just because there is an effect on the intermediate doesn’t mean there’s going to be an effect on the final. As I see it, the response-adaptive randomization would use this equivocal intermediate response to help determine the randomization ratio and whether you should take things through to the final outcome measure. I don’t think we usually know enough about the relationship of the intermediate and the final outcome measures to be able to reliably judge that, in terms of the magnitudes of those benefits, and whether, even if we based the value of an intermediate outcome on past clinical trials, we still wouldn’t know that magnitude of relationship between the intermediate and the final would hold for a new treatment that we are evaluating.
Another comment has to do with how DSMBs were run 25 years ago when I was at the National Heart, Lung, and Blood Institute. The committee started its life as a policy advisory board; it made comments on the protocol and the operations. It then became a DSMB. Once it became a DSMB, it no longer changed the design. I think that model was very useful because it ensured that the DSMB had sufficient expertise and knowledge to review the interim data. I think the worry expressed is that the DSMB might become too invested in the design and the trial to be independent, but that’s not what I observed. In my experience, these committees retained scientific objectivity.
Participants
Footnotes
Acknowledgements
The authors thank the Center for Clinical Epidemiology and Biostatistics in the Perelman School of Medicine at the University of Pennsylvania, Genentech (a member of the Roche Group), Johnson & Johnson (Janssen R&D), and Merck.
Declaration of conflicting interests
Dr Jason T Connor is an employee of Berry Consultants, LLC, which specializes in methods of innovative clinical trial design. No other conflicts were reported.
Funding
External funding for this conference was provided by Genentech (a member of the Roche Group), Johnson & Johnson (Janssen R&D), and Merck.
