
Other
Select search scope: search across all journals or within the current journal

The International Conference on Harmonization (ICH) guideline “Statistical Principles for Clinical Trials” was adopted by the Committee for Proprietary Medicinal Products in March 1998, and consequently is operational in Europe. In October 1998 a one-day discussion forum was held in London by Statisticians in the Pharmaceutical Industry (PSI). The aim of the meeting was to discuss how statisticians were responding to some of the issues in the guideline, and to document consensus views where they existed. The forum was attended by industrial, academic, and regulatory statisticians. This paper outlines the questions raised, resulting discussions, and consensus views reached.
What do we mean by therapeutic equivalence and how do we prove that two drugs are therapeutically equivalent? We suggest that therapeutic equivalence should be defined and demonstrated using predefined limits on the dose scale rather than on the effect scale. A simple design of an equivalence study is discussed, including statistical analysis and the power calculation.
In this paper we introduce a fully Bayesian approach to sample size determination in clinical trials. In contrast to the usual Bayesian decision theoretic methodology, which assumes a single decision maker, our approach recognizes the existence of three decision makers, namely: the pharmaceutical company conducting the trial, which decides on its size; the regulator, whose approval is necessary for the drug to be licensed for sale; and the public at large, who determine ultimate usage. Moreover, we model the subsequent usage by plausible assumptions for actual behavior, rather than assuming that it represents decisions which are in some sense optimal.
The results, not surprisingly, show that the optimal sample size depends strongly on the expected benefit from a conclusively favorable outcome, and on the strength of the evidence required by the regulator.
In this paper we present a decision analytic approach to determining sample size in a set of Phase III drug efficacy trials sponsored by a pharmaceutical company. In this approach we built a model to predict the expected net present value of the drug at varying sample sizes. We then chose the sample size that maximizes that value. We took into consideration effects that sample size had on the probability of approval, the cost of the studies, the time it takes the drug to reach the market, and a variety of other factors. We found that increasing the sample size increases the chance that the drug will be approved. On the other hand increasing the sample size increases the cost of the studies and the time it will take for the drug to reach the market. The model weighs these factors to produce the expected net present value.
In the pharmaceutical industry it is often important to ensure fast patient enrollment. A common practice to speed up patient enrollment is to use more clinical centers. One of the concerns regarding this practice is that increasing the total number of centers to a certain degree may decrease the statistical efficiency of treatment comparisons. This is mainly because a study with too many clinical centers usually includes quite a few small centers and very often they do not have enough patients to represent all treatment groups. These small centers often carry little information on treatment differences. This paper utilizes a statistical model to quantify the relationship between statistical efficiency and the number of centers under typical clinical trial settings. Results in this paper provide useful statistical knowledge in choosing the number of clinical centers when planning multicenter studies.
This paper argues that randomization should more often be considered for early trials of experimental cancer treatments. Uncontrolled as well as randomized designs for Phase II trials are reviewed. The severe limitations of uncontrolled designs are illustrated, and possible difficulties with randomized designs discussed. It is argued that in randomized Phase II trials, the treatment allocation method should be chosen to achieve good balance with respect to the most important prognostic factors.
Assessment of child growth is problematic: growth is nonlinear in the long-term, and unpredictable in the short-term; growth is subject to a number of environmental as well as genetic influences; and growth is difficult to measure reliably. The potential for growth delay as an effect of asthma is established, although it has proved difficult to quantify how great an impact this has on height, growth velocity, or final attained height. In the treatment of asthmatic children, there remain uncertainties as to the effect of inhaled corticosteroids on growth, given the great number of factors affecting growth. In this paper we present recommendations for the design and analysis of trials to assess the effect of regular treatment with inhaled corticosteroids on growth in asthmatic children. Design recommendations are articulated for study duration, entry criteria, other factors that may affect growth, measuring height, measuring growth, study objectives, and considerations relating to confounding between treatment allocation and the effect of the disease on growth. Special attention is given to analyses that address both the intra-subject correlation arising from multiple measurements in longitudinal studies of growth and the potential bias in treatment comparisons due to dropouts, especially those due to treatment failure.
Bilateral pharmacokinetic interaction studies are performed to test if the coadministration of the two drugs alters the kinetics of either one. For this purpose, standard designs are crossovers in which the two drugs are successively given alone and together. Because the systemic concentration of one drug, A, will be zero when the other drug, B, is given alone and vice versa, two separate analyses for Drugs A and B are required. Two symmetrical designs which generate identical analyses for Drugs A and B are reviewed: the dual balanced two-period and the balanced Latin Square three-period designs. Their efficiency in accurately and precisely estimating direct treatment differences (AB-B) or (AB-A) is compared, when analyzed with a standard linear mixed model that either does or does not include a first-order carry-over effect. From this evaluation, the balanced three-period design is preferred because of its 25% superior efficiency when no carryover effect is present. In the presence of carry-over, the three-period design also produces unbiased estimators for the direct treatment and carry-over effects whose variances are proportional to the within-subject variability.
The American College of Rheumatology Responder Index (ACR index) is commonly used to measure treatment efficacy in rheumatoid arthritis trials. This composite index measures the patients' responses as improved or not improved. We use a logit model to describe the dose-response relationship and construct dual-objective optimal designs using prior information. The specific objectives of our designs are to simultaneously estimate the parameters in the dose-response curve and the minimum dose that reaches clinically meaningful significance.
Screening designs are frequently used in the pharmaceutical industry to efficiently study the importance of variables during the formulation and process development phase of a new pharmaceutical compound. This paper discusses the construction of an orthogonal 16-run hierarchical screening design for a 3 × 28 experiment, based on a fold-over Hadamard matrix, under conditions which define a new class of designs. The extension to larger designs of this type are described and the analysis is given in terms of sums and differences. This new class of designs will provide the applied statistician with a greater ability to develop efficient screening designs. A detailed example of such a design is given.
Periodic reevaluation of human and financial resource use efficiency in clinical drug development is essential to continued profitability. Starting with a clean slate, knowing what
The key project milestone dates were met, and the final statistical report was delivered on time at a cost one-third less than anticipated. The quality and usefulness of the report and the process by which it was produced were eminently satisfactory to the customer. Many areas for process improvement were identified. The biggest savings were achieved by reducing rework and review, with no sacrifice in the quality or integrity of the final report. Anticipating and planning for a CRO's needs present opportunities for improving the efficiency of interactions with CROs.
Prentice proposed a definition of surrogate endpoints and operational criteria for their validation. Freedman supplemented these criteria with the proportion explained, which is supposed to represent the proportion of the treatment effect upon the true endpoint which is mediated by the surrogate endpoint. In this paper, we argue that the proportion explained should be replaced by other quantities, such as the relative effect linking the effects of treatment on both endpoints and an individual-level measure of agreement between both endpoints. The latter quantity carries over naturally when data are available on several randomized experiments, while the former can be extended to be a trial-level measure of agreement between the effects of treatment of both endpoints. This approach suggests a new definition of surrogacy, and allows one to predict the effect of treatment upon the true endpoint, given its observed effect upon the surrogate endpoint.
For fixed sample size designs, there is a risk that expected trial outcomes may not obtain adequate power because there is usually some uncertainty about the variance of the outcome variable in the planning phase. This deficit can be remedied by reestimating the variance during the ongoing trial and modifying the initially planned sample size if necessary. From a regulatory view, any adjustment to sample size should occur without unblinding the treatment group membership of the data. Corresponding methods proposed until now are restricted to studies comparing two treatment groups. We present sample size recalculation procedures for multiarmed trials and compare their performance by Monte Carlo simulations.
Internal pilot studies used for establishing the required sample size for a clinical trial part way through that trial are attractive when one is faced with difficulties in determining underlying variances or underlying event rates. By extrapolating a known curio of
The questions of
Objective: The objective of this study was to derive a measure of quality-time to be used in treatment evaluations for gastroesophageal reflux disease (reflux disease). Specifically, we sought to refine the Extended Quality-adjusted Time without Symptoms and Toxicities (Q-TWiST) approach and propose a new measure, called Quality-Days Incrementally Gained (QDIG).
Design: One hundred and sixty-seven patients with reflux disease randomized to one of two treatments completed a health-related quality-of-life questionnaire at baseline and after four and eight weeks of treatment. The questionnaire contains generic and reflux disease-specific measures and corresponding importance items. The weighted assessment score, fundamental to calculating both the Extended Q-TWiST and the QDIG, was computed using several weighting schemes to determine the most robust and appropriate statistic.
Main Outcome Measures and Results: The Extended Q-TWiST was affected by the relative weighting of baseline and follow-up weighted assessment score. The QDIG directly discounts the baseline score and showed the greatest sensitivity to treatment differences. The variance in the Extended Q-TWiST and QDIG were both reduced by using importance weights carried forward from baseline rather than time-varying importance weights, and by using population-weights rather than individual person weights.
Conclusions: Careful consideration should be made when deciding to use the Extended Q-TWiST or QDIG approach. Our data suggest the QDIG approach is superior in studies of short duration with heterogeneous populations.
This paper describes several statistical methods for analyzing medical device reports received by the Food and Drug Administration. The nonparametric regressions (polynomial, loess, kernel smooth, and cubic spline smoothing) are used as exploratory tools to evaluate trends in adverse event reports. Several statistical models, including simple Poisson, mixed binomial/Poisson, zero-truncated Poisson, and the negative binomial (compound Poisson), are used to monitor and to determine the upper 95% threshold values for reported adverse events during the study period. The hip implant injury data and the intravenous tube total (sum of death, device malfunction, and injury) data are used to illustrate our model fitting procedures. In this paper, only the numerator data (medical device adverse events), not the denominator data (medical device usage), are available in our statistical analysis. The possible effects of marketing time and other factors, which may be available in adverse drug reactions, are not available and are not considered for medical device adverse events in this paper.
New Drug Applications require the formation of an integrated safety summary as part of the license submission. The objective of this summary is to identify important serious adverse events and to characterize more common, nonserious adverse events. This review typically involves pooling the data across different studies and comparing the event rates on the drug and a common comparator (placebo or active). The most common approach is to collapse data from all studies and summarize them as if they came from a single study. From a scientific viewpoint this is not optimal as between study variability is not accounted for. The dangers of simple pooling are demonstrated with a real example. It is argued that where appropriate, more use should be made of statistical techniques that account for between-study variability; this will be particularly relevant when the treatment allocation ratio is unbalanced across studies.
Statistical hypothesis tests are used as a flagging device to highlight differences worth further attention in the evaluation of repeated dose toxicity studies in the rat. Raw data of quantitative parameters of 19 regulatory toxicity studies were collected with their final interpretation of each study. An investigation was done on the consistency between flagging by statistical tests and biological significance by final interpretation. Williams's test at 2.5% of the significance level showed as much accuracy (correct results compared with the sum of false negative and false positive results) as the rate of Dunnett's test at the 5% significance level. Since a monotonic dose-response relationship is usually assumed in selection of dose levels, Williams's test with ordered alternative hypotheses is recommended as a routine procedure instead of the currently used Dunnett's test. A supplementary procedure, using Steel's test, was shown to be effective for flagging unexpected ‘downturn’ dose response.
Covariates that affect the outcome of a disease are often incorporated into the design and analysis of clinical trials. This serves two main purposes: 1. To improve the credibility of the trial results by demonstrating that any observed treatment effect is not accounted for by an imbalance in patient characteristics, and 2. To improve statistical efficiency. In this paper, we review procedures for the adjustment of treatment effects for the influence of covariates and discuss some statistical and regulatory issues on the applications of these procedures.
A major problem in the analysis of clinical trials is missing data caused by patients dropping out of the study before completion. This problem can result in biased treatment comparisons and also impact the overall statistical power of the study. This paper discusses some basic issues about missing data as well as potential “watch outs.” The topic of missing data is often not a major concern until it is time for data collection and data analysis. This paper provides potential design considerations that should be considered in order to mitigate patients from dropping out of a clinical study. In addition, the concept of the missing-data mechanism is discussed. Five general strategies of handling missing data are presented: 1. Complete-case analysis, 2. “Weighting methods,” 3. Imputation methods, 4. Analyzing data as incomplete, and 5. “Other” methods. Within each strategy, several methods are presented along with advantages and disadvantages. Also briefly discussed is how the International Conference on Harmonization (ICH) addresses the issue of missing data. Finally, several of the methods that are illustrated in the paper are compared using a simulated data set.
Many of the reservations that might attach to the use of meta-analysis generally (for example, regarding publication bias) do not apply in the specific context of drug development. A meta-analysis is, in fact, a highly natural and appropriate way to summarize the results of a drug development program, as has been recognized in the International Conference on Harmonization (ICH) E9 guideline. Since a sponsor will have access to all original data, the data from a set of clinical trials in a drug development program have a very similar (hierarchical) structure to the data from a set of centers in a single multicenter trial. Curiously, however, the controversies over analyzing multicenter trials have often been different from those in the field of meta-analysis. In this paper, the options open to the meta-analyst in drug development are examined and comparisons to approaches used in analyzing multicenter trials are made in an attempt to provide some unifying insights, in particular as regards the handling of models with interactions.
Some developments are described in the case of multidose equivalence studies (multiple doses of the test substance and one dose of the reference). If equivalence cannot be proven for the test substance at one of the doses studied, the procedure uses linear interpolation to examine intermediate doses. This greatly increases the power of the equivalence testing procedure, provided that equivalence for an intermediate dose only is considered an acceptable result of the study.
The closed test procedure for equivalence testing of the doses studied is also described in order to clarify methods presented in other statistical publications. It is often difficult to specify the equivalence criterion (delta) precisely, although it is taken as a known constant in many statistical publications. It is recommended that results be presented in a form that allows them to be checked against several different values of delta.
Equivalence trials have the objective of demonstrating that an investigational treatment, for example, a test drug under development, is not different from a reference treatment by more than a prespecified clinically irrelevant amount. The purpose of this paper is to investigate for the crossover design the situation when equivalence is defined in terms of the ratio of location parameters. An approximate formula for sample size calculation is presented for the case of normally distributed endpoints, and nonparametric methods for testing and calculation of confidence intervals are provided.
Bioequivalence between two treatments or two drugs is often assessed by comparing the two proportions (success rate or eradication rate) of binomial outcomes when the conventional pharmacokinetic parameters are inadequate for the assessment. Setting the equivalence limits can be based on one of the three measures: difference, ratio, or odds ratio between the two binomial probabilities. This paper reviews the existing asymptotic test statistics for comparing two independent binomial probabilities in terms of the three measures in the context of equivalence or noninferiority testing. The actual type I error and power of the asymptotic tests are evaluated by enumerating the exact probabilities in the rejection region. The results show that to establish an equivalence between two treatments with an equivalence limit of 20% in difference, a sample size of at least 50 per treatment is needed. When the sample size is sufficient, the actual type I error rate is close to the nominal level (slightly above the nominal level in several cases) for a test in terms of difference for equivalence limits, and it tends to exceed the nominal level for tests in terms of ratio or odds ratio.
A dose-response study, which is performed to determine whether or not there is any effect of a new drug related to dose, plays a very important role in the clinical development of a drug. Finding evidence of the dose-response relationship is usually done based on hypothesis testing, which has been considered an appropriate way to analyze a dose-response study. Hypothesis testing does not provide information about certain structures of the dose-response relationship, especially the shape and location of the dose-response curve, though the information is most helpful in determining the clinical dose of a drug. In this paper, the model-based approach with data-adaptive distribution is introduced to infer the dose-response relationship. We also introduce the statistical descriptive use of the empirical cumulative distribution function. Furthermore, methods to compare two dose-response curves are considered.
For the analysis of dose-response relationship under the assumption of ordered alternatives several global trend tests are available. Furthermore, there are multiple test procedures which can identify doses as effective or even as minimally effective. In this paper it is shown that the principles of multiple comparisons and interim analyses can be combined inflexible and adaptive strategies for dose-response analyses; these procedures control the experimentwise error rate.
We consider analysis of binary observations from multiple sites of each subject. In this case, observations from the same subject tend to be correlated. In estimating the common response probability in correlated binary data, two weighting systems have been most popular: equal weights to sites, and equal weights to subjects. When the number of sites varies subject by subject, performance of these two weighting systems depends on the extent of correlation among sites within each subject. In this paper, we describe a new weighting method that minimizes the variance of the estimator. We apply these methods to data from a study involving an enzymatic diagnostic test to illustrate the estimation of the sensitivity and the specificity of periodontal diagnostic tests. Simulation studies were conducted to compare the performance of the new estimator with that of other estimators.
When a confirmatory test is completely accurate or has known low error rates, the sensitivity and the specificity of a screening test can be estimated. When the error rates for the confirmatory test are unknown, Hui and Walter (2) presented a method for estimating the sensitivity and specificity of both the screening and the confirmatory tests using the tests on two populations with different prevalence rates of the infection. The method requires that the tests have equal error rates in the two populations. When this requirement is not met, we show that the estimated prevalence rates are robust when the difference in the prevalence rates of the two subpopulations is large. An alternative design, requiring only one population, but other assumptions, is also described.
Statistical methods developed for survival analysis (time to event) were applied to treatments which are potent viral load suppressers in HIV-infected subjects. The methodology was applied to HIV-RNA levels obtained from the clinical study AG1343–511 (nelfinavir [NFV] in combination with zidovudine +lamivudine). Utilizing stringent treatment response definitions based on the limit of quantification (LOQ) of two assays, the analyses established that nelfinavir containing arms were significantly better than the control arm and the 750 mg NFV arm was superior to the 500 mg NFV arm. Analyses of durability (duration) of response revealed that subjects qualified as responders to more sensitive (lower LOQ) assay were associated with longer duration of response.
The statistical methodology utilized and the parameters defined in this article can be applied to the evaluation of anti-HIV treatments in general. For potent anti-HIV treatments, the method can be utilized with the definitions developed here; less potent treatments can be analyzed with proper modifications in definitions of response and relapse.
Consistency in statistical analyses of preclinical studies is a request from authorities worldwide because it facilitates the evaluation and comparison of similar studies. The goal of this paper is to describe mixed-effect analysis of variance models broadly applicable to the analysis of continuous responses from toxicity and safety pharmacology studies. Traditional models are discussed together with more complex models including the split-plot model and related models for repeated measurements. The model fitting process ranges from the choice of model to the final check of model assumptions which we exemplify using case studies analyzed by the SAS® procedure.
The pharmaceutical industry is characterized by mergers and restructuring; extensive use of outsourcing and global development strategies; new regulatory requirements for electronic records; new methodologies for conducting large clinical trials; and the need to effectively exchange information between the components of a corporate information system. Middleware connectivity and the workflow management systems ensure comprehensive integration of a corporate information system. Middleware enables pharmaceutical companies to exchange information between multiple systems or components, and to use and distribute the information on a server using an event-driven mechanism. Workflow technology can streamline and secure the information flow and approval processes, supplying a tool to electronically view, manage, revise, share, and distribute virtually any information or document across the enterprise without paper. With these tools, information management can move from the traditional request/replay model, in which users and applications must ask for status updates on needed information, to an event-driven approach, in which they receive this information automatically when it becomes available. A strategy for an adaptive evolution of the information management process in response to continuous improvements in technology and the evolution of business strategies is presented.