
Editorial
Select search scope: search across all journals or within the current journal


One key objective of a multi-regional clinical trial (MRCT) is to use the trial results to ‘bridge’ from the global level to local region in support of local registrations. However, data from each individual country are typically limited and the large number of countries will increase the chance of false positive findings.
Graphical tools to facilitate identification of potential outlying countries could be useful for country-level assessment. Existing methods such as funnel plot and expected range of treatment effect can substantially increase the false positive rate. The expected range approach can also have a very low power when there are a large number of small countries, which is typical in a MRCT.
In this article, we apply normal probability plots, commonly used as a diagnostic tool in linear regression analysis, to assess the differences among countries. Evidence of possible inconsistency, which incorporates both the estimated treatment effect and sample size, is plotted against its expected order statistic.
A simulation study is conducted to assess the impact of the negative correlation among residuals due to unequal sample sizes among countries and the performance of the proposed methods compared to existing approaches. The proposed methods tend to have a balanced consideration with substantially smaller false positive rate and reasonable probability to identify outlying countries in realistic scenarios.
While much lower than that of commonly used methods, the false positive rates of the proposed methods are not strictly controlled. This may be acceptable for these graphical tools with intention to flag potential outliers for investigation.
We recommend routine use of normal probability plots in MRCTs as a tool to identify potential outliers. If the normal probability plot is approximately linear but has heavy tails with a few outlying countries, these potential outliers should be examined carefully to understand the possible reasons.
In the planning of a dose finding study, a primary design objective is to maintain high accuracy in terms of the probability of selecting the maximum tolerated dose. While numerous dose finding methods have been proposed in the literature, concrete guidance on sample size determination is lacking.
With a motivation to provide quick and easy calculations during trial planning, we present closed form formulae for sample size determination associated with the use of the Bayesian continual reassessment method (CRM).
We examine the sampling distribution of a nonparametric optimal design and exploit it as a proxy to empirically derive an accuracy index of the CRM using linear regression.
We apply the formulae to determine the sample size of a phase I trial of PTEN-long in pancreatic cancer patients and demonstrate that the formulae give results very similar to simulation. The formulae are implemented by an R function ‘getn’ in the package ‘dfcrm’.
The results are developed for the Bayesian CRM and should be validated by simulation when used for other dose finding methods.
The analytical formulae we propose give quick and accurate approximation of the required sample size for the CRM. The approach used to derive the formulae can be applied to obtain sample size formulae for other dose finding methods.
The two-stage, likelihood-based continual reassessment method (CRM-L) entails the specification of a set of design parameters prior to the beginning of its use in a study. The impression of clinicians is that the success of model-based designs, such as CRM-L, depends upon some of the choices made with regard to these specifications, such as the choice of parametric dose–toxicity model and the initial guess of toxicity probabilities.
In studying the efficiency and comparative performance of competing dose-finding designs for finite (typically small) samples, the nonparametric optimal benchmark is a useful tool. When comparing a dose-finding design to the optimal design, we are able to assess how much room there is for potential improvement.
The optimal method, based only on an assumption of monotonicity of the dose–toxicity function, is a valuable theoretical construct serving as a benchmark in theoretical studies, similar to that of a Cramér–Rao bound. We consider the performance of CRM-L under various design specifications and how it compares to the optimal design across a range of practical situations.
Using simple recommendations for design specifications, the CRM-L will produce performances, in terms of identifying doses at and around the maximum tolerated dose (MTD), that are close to the optimal method on average over a broad group of dose–toxicity scenarios.
Although the simulation settings vary in the number of doses considered, the target toxicity rate, and the sample size, the results here are presented for a small, though widely used, set of two-stage CRM designs.
Based on simulations here, and many others not shown, CRM-L is almost as accurate, in many scenarios, as the nonparametric optimal design. On average, there appears to be very little margin for improvement. Even if a finely tuned skeleton offers some improvement over a simple skeleton, the improvement is necessarily very small.
Evaluation of anti-infective drugs for licensure often relies on a noninferiority (NI) design where new drug B is noninferior to comparator drug A if the difference in success rates is reliably not worse than some fixed margin. The margin must be based on historical studies that estimate the magnitude of the benefit of drug A over placebo. This approach hampers drug development because the obligatory evidence for margin determination is often nonexistent.
To develop a new method for licensure of anti-infective drugs when there is no historical evidence for reliable construction of the NI margin.
The minimum inhibitory concentration (MIC) measures the minimum amount of drug that it takes to inhibit growth of bacteria in vitro. Patients who are infected with bacteria that have a low MIC to a given drug are expected to have good outcomes when treated with that drug. Thus, a differential effect of drug B versus drug A, if it exists, is likely to occur in patients whose pathogens have discordant MICs (e.g., low MIC for drug B, high MIC for drug A, or vice versa). A new paradigm for licensure of anti-infective drugs is proposed where a clinically acceptable NI margin is selected and licensure supported if the NI margin is met and B is reliably demonstrated superior to A in a subset of patients whose paired MICs favor B. The requirement for some evidence of superiority encourages a study that is carefully designed and executed.
Simulations indicate that the approach shows promise in realistic settings, provided adequate data are available. A simulated example illustrates use of the methods.
If the data have small sample size, weak MIC/success relationship, or high correlation between MIC-A, MIC-B, this procedure will have poor power.
Discordant MIC analysis may offer a novel path to licensure for certain anti-infective drugs.
Clinical validation of a predictive biomarker is especially difficult when the biomarker cannot be assessed retrospectively. A cost-effective, prospective multicenter replication study with rapid accrual is warranted prior to further validation studies such as a marker-based strategy for treatment selection. However, it is often unknown how measurement error and bias in a multicenter trial will differ from that in single-institution studies.
Power calculations using simulated data may inform the efficient design of a multicenter study to replicate single-institution findings. This case study used serial standardized uptake value (SUV) measures from 18F-fluorodeoxyglucose (FDG) positron emission tomography (PET) to predict early response to breast cancer neoadjuvant chemotherapy. We examined the impact of accelerating accrual through increased inclusion of secondary sites with greater levels of measurement error and bias. We also examined whether enrichment designs based on breast cancer initial uptake could increase the study power for a fixed budget (200 total scans).
Reference FDG PET SUV data were selected with replacement from a single-institution trial; pathologic complete response (pCR) data were simulated using a logistic regression model predicting response by mid-therapy percent change in SUV. The impact of increased error for SUV measurements in multicenter trials was simulated by sampling from error and bias distributions: 20%−40% measurement error, 0%−40% bias, and fixed error/bias values. The proportion of patients recruited from secondary sites (with higher additional error/bias compared to primary sites) varied from 25% to 75%.
Reference power (from source data with no added error) was 0.92 for N = 100 to detect an association between percentage change in SUV and response. With moderate (20%) simulated measurement error for 3/4, 1/2, and 1/4 of measurements and 40% for the remainder, power was 0.70, 0.61, and 0.53, respectively. Reduction of study power was similar for other manifestations of measurement error (bias as a percentage of true value, absolute error, and absolute bias). Enrichment designs, which recruit additional patients by not conducting a second scan in patients with unsuitable pre-therapy uptake (low baseline SUV), did not lead to greater power for studies constrained to the same total cost.
Simulation parameters could be incorrect, or not generalizable. Under a different logistic regression model relating mid-therapy percent change in SUV to pCR (with no relationship for patients with low baseline SUV, rather than the modest point estimate from reference data), the enrichment design did have somewhat greater power than the unselected design.
Even moderate additional measurement error substantially reduced study power under both unselected and enrichment designs.
Despite the proliferation of health information technology (IT) interventions, descriptions of the unique considerations for conducting randomized trials of health IT interventions intended for patient use are lacking.
Our purpose is to describe the protocol to evaluate Pocket PATH® (Personal Assistant for Tracking Health), a novel health IT intervention, as an exemplar of how to address issues that may be unique to a randomized controlled trial (RCT) to evaluate health IT intended for patient use.
An overview of the study protocol is presented. Unique considerations for health IT intervention trials and strategies are described to maintain equipoise, to monitor data safety and intervention fidelity, and to keep pace with changing technology during such trials.
The sovereignty granted to technology, the rapid pace of changes in technology, ubiquitous use in health care, and obligation to maintain the safety of research participants challenge researchers to address these issues in ways that maintain the integrity of intervention trials designed to evaluate the impact of health IT interventions intended for patient use.
Our experience evaluating the efficacy of Pocket PATH may provide practical guidance to investigators about how to comply with established procedures for conducting RCTs and include strategies to address the unique issues associated with the evaluation of health IT for patient use.
The Prostate Cancer Intervention Versus Observation Trial (PIVOT) randomized 731 men with localized prostate cancer to radical prostatectomy or observation.
We describe the methods and results for cause-of-death assignments in PIVOT, and compare them to alternative strategies for ascertaining prostate cancer–specific mortality, as well as to the methods and results in the similar Scandinavian Prostate Cancer Group Study 4 (SPCG-4) trial.
Three PIVOT Endpoints Committee members, blinded to randomized treatment assignments, reviewed medical records and death certificates when available to assign a cause of death using a primary and a secondary adjudication question. Initial disagreements were resolved through discussion. The level of initial agreement among committee members was examined, as well as guesses at randomized treatment assignments for a convenience sample of cases. Final cause of death determinations were compared to death certificates.
Complete agreement on cause of death by all three committee members before any discussion was achieved in 200/354 (56%) cases on the primary and 209/354 (59%) cases on the secondary. However, complete agreement on the primary rose to 306/354 (86%) when ‘definite’ and ‘probably’ categories were collapsed, as planned a priori. The three committee members’ proportions of correct guesses of randomized treatment assignment were 82/121 (68%), 113/148 (76%), and 99/134 (74%). Using the committee’s final adjudications as a gold standard, death certificates had suboptimal sensitivities, specificities, or predictive values depending on how they were used to determine cause of death.
There was no separate ‘gold standard’ by which to judge the accuracy of the final endpoints committee adjudications, and useful death certificates could not be obtained on about a third of PIVOT participants who died.
The low level of initial agreement on cause of death among endpoint committee members and the potential for biased determinations due to partial unblinding to treatment assignment raise methodologic concerns about using prostate cancer mortality as an endpoint in clinical trials like PIVOT.
Results from clinical trials are often slowly implemented. We studied whether participation in multicenter clinical trials improves reported dissemination, convincement, and subsequent implementation of its results.
We sent a web-based questionnaire to gynecologists, residents, nurses, and midwives in all obstetrics and gynecology departments in the Netherlands. For nine trials in perinatology, reproductive medicine, and gynecologic oncology, we asked the respondents whether they had knowledge of the results, were convinced by the results, and what percentage of their patients were treated according to the results of these trials. We compared the level of knowledge, convincement, and reported implementation of results in practice for the nine trials for respondents who worked in hospitals that had recruited for a trial with respondents who worked in a hospital that had not recruited for that trial. The reported implementation was restricted to six trials that showed decisive results.
We analyzed 202 questionnaires from 83 departments in obstetrics and gynecology in the Netherlands (93% of all departments). The percentage of respondents who had worked in a hospital that recruited for a specific study varied between 8% and 71% per study and was 28% on average. The relative risk (RR) for knowledge of the study result for respondents who had worked in a recruiting hospital was for all studies positive and varied between 1.1 and 3.3 (pooled RR: 1.8, 95% confidence interval (CI): 1.7–1.9). In general, health-care workers were convinced of trial results, independent of whether they had worked in a hospital that recruited for a trial or not (pooled RR: 1.02, 95% CI: 0.99–1.05). Reported implementation of trial’s results, that is, less than 20% were treated with unfavorable treatment according to study results, was better in hospitals that had recruited for those trials (pooled RR: 1.1, 95% CI: 1.02–1.19).
Participation in these multicenter clinical trials was associated with better knowledge about the trial’s results, with a minor improvement of the reported implementation of the study results.
Recent research has proposed a new method for defining a favorable outcome in traumatic brain injury and stroke research.
This new method is called the sliding dichotomy, and it is suggested as a potential solution to the problem of underpowered clinical trials.
We present a brief simulation study and graphical comparison of the power of each method to detect varying treatment effect sizes.
Simulations of a patient population similar to the National Acute Brain Injury Study: Hypothermia (NABISH) study indicate that the sliding dichotomy method does not result in higher power than traditional methods.
Although the sliding dichotomy may present gains in power in some cases, several aspects of the patient population need to be considered in choosing between sliding dichotomy and traditional definitions of favorable outcomes.
Subjects who enroll in multiple studies have been found to use deception at times to overcome restrictive screening criteria. Deception undermines subject safety as well as study integrity. Little is known about the extent to which experienced research subjects use deception and what type of information is concealed, withheld, or distorted.
This study examined the prevalence of deception and types of deception used by subjects enrolling in multiple studies.
Self-report of deceptive behavior used to gain entry into clinical trials was measured among a sample of 100 subjects who had participated in at least two studies in the past year.
Three quarters of subjects reported concealing some health information from researchers in their lifetime to avoid exclusion from enrollment in a study. Health problems were concealed by 32% of the sample, use of prescribed medications by 28%, and recreational drug use by 20% of the sample. One quarter of subjects reported exaggerating symptoms in order to qualify for a study and 14% reported pretending to have a health condition in order to qualify.
Although this study finds high rates of lifetime deceptive behavior, the frequency and context of this behavior is unknown. Understanding the context and frequency of deception will inform the extent to which it jeopardizes study integrity and safety.
The use of deception threatens both participant safety and the integrity of research findings. Deception may be fueled in part by undue inducements, overly restrictive criteria for entry, and increased demand for healthy controls. Screening measures designed to detect deception among study subjects would aid in both protecting subjects and ensuring the quality of research findings.
Children living in nonmetropolitan communities are underserved by evidence-based mental health care and are underrepresented in clinical trials.
In this article, we describe lessons learned in conducting the Children’s Attention-Deficit Hyperactivity Disorder (ADHD) Telemental Health (TMH) Treatment Study (CATTS), a randomized controlled trial testing the effectiveness of TMH in improving outcomes of children with ADHD living in underserved communities.
Children were referred by primary care providers (PCPs). The test intervention group received six telepsychiatry sessions with each session followed by an caregiver behavior training session delivered in-person by a local therapist. A secure website was used to support decision making by the telepsychiatrists and to facilitate real-time collaboration between the telepsychiatrists and community therapists. The control group received a single telepsychiatry consultation. Questionnaires tapping ADHD symptoms and other outcomes were administered to parents and teachers online through a secure portal from personal computers.
total of 88 PCPs in seven communities referred the 223 children who participated in the trial. Attrition in treatment sessions and research assessments was very low.
TMH proved to be a viable means of providing evidence-based pharmacological services to children and training to local therapists. Recruitment was enhanced by offering the control group a telepsychiatry consultation. Site-specific strategies were needed to meet recruitment targets.
The CATTS trial used methods designed to optimize inclusion of children living in multiple dispersed and underserved areas. The study will serve as a model for other research projects aiming at reducing geographic disparities in access to quality mental health care.
Participation in an exercise trial is a major commitment for cancer survivors, but few exercise trials have evaluated patient satisfaction with trial participation.
To examine patient satisfaction with participation in the Healthy Exercise for Lymphoma Patients (HELP) Trial and to explore possible determinants.
The HELP Trial randomized 122 lymphoma patients to 12 weeks of supervised aerobic exercise training (AET; n = 60) or to usual care (UC; n = 62), with the option of participating in a 4-week posttrial exercise program. At the 6-month follow-up assessment, participants evaluated their overall trial satisfaction.
Personal satisfaction with trial participation was strongly influenced by group assignment with participants randomized to AET reporting participation to be more rewarding (p < 0.001) and personally useful (p < 0.001) than participants randomized to UC. UC participants who completed the optional 4-week posttrial exercise program reported participation to be more rewarding (p = 0.008) and personally useful (p < 0.001) than UC participants who declined the program.
The study is limited by the lack of a validated measure of participant satisfaction, and the fact that the offer of participation in the posttrial exercise program to the UC group was not randomized.
Lymphoma patients randomized to UC viewed it as less rewarding and personally useful despite being offered a 4-week posttrial exercise program. UC participants who completed the 4-week program reported personal satisfaction levels similar to the AET group; however, the causal direction of this association is unknown. Researchers should continue to evaluate participant satisfaction in exercise trials.
The use of decision-support interventions in the context of decisions about trial participation is an emergent field. There is a lack of evidence about what information is deemed important to support decisions about informed consent for clinical trials, and whether different groups agree on the information for inclusion.
The overall objective was to determine the items which different stakeholder groups viewed to be important for inclusion in a decision-support tool when making decisions about clinical trial participation, with a view to use these as a framework for developing decision-support tools in this context. This is the first study to have addressed this issue.
A modified Delphi method was used to determine agreement on importance of items. The ‘stakeholder’ panel was made up of 49 individuals from 5 groups: 11 trialists, 6 research nurses, 7 ethics committee chairs, 9 decision-support experts, and 16 patients (9 trial experienced and 7 trial non-experienced). Two rounds of rating were completed. Items with a median of 7–10 with ≥65% of any one group (from aggregate ratings) in agreement were considered important for inclusion.
The stakeholder panel achieved consensus on the majority of items included (60/66), agreeing that these were important for inclusion in a decision-support tool for trial participation. These included items covering information about trial participation and standard care, information on the likelihood of receiving different treatments, information to help patients determine what matters most to them, ensuring that the information is balanced, guidance on how to make a decision, disclosure of any conflicts of interest, using plain language in the tool, and guidance on the decision-support development process. Some areas of divergence among the panel were also identified relating to the use of patient stories.
Selection bias may be a limitation in this study due to the manner in which the participants were invited to take part, and therefore, the representativeness, and reproducibility with another group of stakeholders, may differ.
Agreement was obtained on a number of items, which we recommend should be used as a framework to develop useful tools to support decision-making about participation in clinical trials.
There are many benefits of data sharing, including the promotion of new research from effective use of existing data, replication of findings through re-analysis of pooled data files, meta-analysis using individual patient data, and reinforcement of open scientific inquiry. A randomized controlled trial is considered as the ‘gold standard’ for establishing treatment effectiveness, but clinical trial research is very costly, and sharing data is an opportunity to expand the investment of the clinical trial beyond its original goals at minimal costs.
We describe the goals, developments, and usage of the Data Share website (http://www.ctndatashare.org) for the National Drug Abuse Treatment Clinical Trials Network (CTN) in the United States, including lessons learned, limitations, and major revisions, and considerations for future directions to improve data sharing.
Data management and programming procedures were conducted to produce uniform and Health Insurance Portability and Accountability Act (HIPAA)-compliant de-identified research data files from the completed trials of the CTN for archiving, managing, and sharing on the Data Share website.
Since its inception in 2006 and through October 2012, nearly 1700 downloads from 27 clinical trials have been accessed from the Data Share website, with the use increasing over the years. Individuals from 31 countries have downloaded data from the website, and there have been at least 13 publications derived from analyzing data through the public Data Share website.
Minimal control over data requests and usage has resulted in little information and lack of control regarding how the data from the website are used. Lack of uniformity in data elements collected across CTN trials has limited cross-study analyses.
The Data Share website offers researchers easy access to de-identified data files with the goal to promote additional research and identify new findings from completed CTN studies. To maximize the utility of the website, ongoing collaborative efforts are needed to standardize the core measures used for data collection in the CTN studies with the goal to increase their comparability and to facilitate the ability to pool data files for cross-study analyses.


