
Research article
Select search scope: search across all journals or within the current journal

A Drug Information Association Workshop on “Statistical Methodology in Non-Clinical and Toxicological Studies” was held at the end of March 1996. The purpose of this meeting was to discuss and to obtain a consensus on the appropriateness of current and new biostatistical methods relevant in this field of drug development. The following summary represents a consensus of the Working Group on Biostatistics in Mutagenicity Studies. The recommendations outline the relevant principles of design and analysis rather than provide detailed specification of statistical methodology, thus providing the possibility of making reasonable choices between alternative approaches.
The Drug Information Association “3rd Annual Biostatistical Meeting” was held on August 27 and 28, 1996 in Tokyo. The purpose of this meeting was to discuss biostatistical recommendations for repeated toxicity studies. The purpose, objective, design, conduct, analysis, and interpretation are described.
This paper deals with the procedures used in the statistical review and evaluation of
Two groups of procedures have been proposed to model the dose-response curve of the number of revenant colonies. The first group includes procedures which are based on biological mechanisms of reverse mutation of bacteria from auxotrophic cells to prototrophic cells. Those procedures consider the toxic effect in addition to the mutagenic effect of the test compound to reflect the possible nonmonotonic or downturn phenomenon in revenant count. They also consider the multigeneration phenomenon of the reverse mutation process. The second group consists of empirically-based procedures. Most procedures in this group also include terms for mutagenic and toxic effects of the test compound.
The unweighted and weighted least squares, the maximum likelihood method, and the quasi likelihood method are used to estimate the parameters in the dose-response curve. Standard procedures based on normal approximation and the likelihood ratio test are used to test the mutagenic and toxic effects.
This paper discusses issues related to the statistical analysis of
The relative power of the different
The purpose of statistical analysis of the data and the relative importance of hypothesis testing and estimation is discussed in the context of the choice of analytical methods, the power of the designs, and the Type 1 and 2 errors associated with the tests. The potential to incorporate historical control data into the assessments is considered.
The statistical test of the traditional hypothesis of no treatment effect is commonly used in toxicological experiments. Failing to reject the hypothesis often leads to the conclusion of evidence in favor of safety. The major drawback of this indirect approach is the fact that what is controlled by a prespecified level is the probability of erroneously concluding hazard (producer risk). The primary concern of safety assessment, however, is the control of the consumer risk, that is, limiting the probability of erroneously concluding safety. In order to restrict this risk, safety has to be formulated as the alternative and hazard, that is, the opposite, has to be formulated as the hypothesis.
Preclinical tests in genetic toxicology represent safety studies. Therefore, the primary concern in the evaluation of Ames test data is the control of the consumer risk, that is, the risk of erroneously concluding safety. Hence, an equivalence test procedure is adequate. This approach is presented beside the two-fold rule and classical tests for differences with respect to an order restricted alternative.
The fixed-dose procedure and the acute-toxic-class method provide alternatives to the LD50 test for the classification of substances by their acute oral toxicity. This paper uses a general mathematical model to explore a wider class of fixed-dose procedures. In these procedures the number of animals included at each dose and the decision criteria regarding the next dose to be used are altered. The use of a number of measures of performance for the procedures, including the expected number of deaths caused and the probability that a substance is given an inappropriate classification, enable comparison of the procedures in this class.
It is found that misclassification is least likely for a test in which the most likely classification depends only on the LD50 of the compound under investigation. Reducing the number of animals used at each dose reduces the expected number of deaths but increases the probability of misclassification. As a compromise, it is proposed that a procedure with six animals tested at each dose be used. The decision as to whether to continue at a higher or lower dose would be based on whether three or more of these animals die. Assuming that the 51 compounds from the Health and Safety Executive database are representative of those tested, such a procedure would give the correct classification for approximately 90% of compounds.
An application of the dose-response curve in the case of dicothomous responses according to a Bayesian point of view is presented in this paper. A dose response curve in cases of dichotomous data describes the relationship between the probability of observing an occurence and a dose level. A researcher may sometimes have some ideas of the shape of such a curve. This curve may be regarded as an “initial dose response curve.” Having observed (mi, ni, xi) i = 1 … k, that is, m, events out of n, proofs at dose xi, a new curve can be obtained. This curve depends on the initial curve stated by the researcher and the data observed. This curve may be regarded as a final dose response curve. It can be obtained from the initial curve via the Bayes formula. The application will be carried out considering a logistic dose response curve.
The current diversity in statistical methodology for even quite simple situations is emphasized. The reasons for this diversity can be grouped into two broad categories: Cultural, including national preference, historical precedent, and resource availability; and statistical, including inadequacies in statistical theory, inadequate knowledge of the biological context, and other practical constraints.
Although criteria for an ideal methodology could be defined, no current methods satisfy them and are therefore suboptimal. It is inferred that to attempt worldwide standardization on a particular statistical test in a particular situation (“Standardization by Method”) is generally not desirable. A proposal will be made, however, that it should be possible to achieve standardization by the setting of “performance criteria” in order to classify proposed methods as “acceptable” or “unacceptable” (“Standardization by Performance”).
A Drug Information Association Workshop on “Statistical Methodology on Non-Clinical and Toxicological Studies” was held on March, 25-27, 1996. The purpose of this meeting was to discuss the appropriateness of current and new biostatistical methods in this field of drug development. This paper proposes a simple closure test for dose-response relationships under real data conditions. This approach takes into account deviation from variance homogeneity and monotonicity assumptions. Moreover, this approach can be easily adapted for “any” kind of two-sample tests, for example, nonparametric, dichotomous, censored, and so forth. Power was compared by a simulation study for selected conditions.
A study was planned which evaluated the consistency between the judgments of toxicologists and well-established statistical methods to detect dose-dependency in repeated toxicity studies. Eighteen bits of repeated toxicity study data for both male and female rats were collected, which included hematology, blood chemistry, and organ weight (absolute and relative) data, totaling 2,001 items. Three or four toxicologists who had more than 10 years of experience judged dose-dependency and change at each dose for each item. The consistency between statistical methods and toxicologists was not as high as was expected. One of the major sources of this inconsistency resulted from the large amount of variation in the recognition among toxicologists. The recognition was qualitatively different, for instance, when there was a difference only between the highest dose and the control group, even if it was very evident, some toxicologists did not regard it as a dose-dependent relationship. It was shown that the consistency of change at each dose was moderately high. It is also noted that toxicologists tend to judge more conservatively than statistical tests at the 5% significance level. Considering the result of this study, a maximum contrast method is recommended to examine the dose-response shape, while avoiding the multiplicity of statistical tests.
This paper introduces the notion of a maximum contrast method for data analysis on experiments in a one-way layout with levels corresponding to doses of a test chemical. It is defined as the method which uses the maximum of a set of contrast statistics to make decisions for objectives, such as determining the minimum effective dose (MED) in clinical trials, the minimum toxic dose (MTD) in chronic toxicity studies, or a plausible dose-response pattern in dose-response studies, where a contrast statistic is defined as the ratio of a contrast of mean response for dose groups and an estimator of the standard deviation of the numerator. Along with this notion the authors devised an extended Williams method for unequal group sizes and compared the performance of three methods: the Williams method, a modified Williams method, and the max-t method, all of which are often used to identify the MTD based on a newly introduced loss function. They concluded that the modified Williams method is comparatively better for the purpose of identifying the MTD, although each method has its own advantageous pattern.
Comparisons of several treatments with a control represent a standard situation in preclinical trials. Usually, they are considered with a single variable, resulting in multiple test procedures such as the Dunnett test (1). Here, the multivariate many-to-one problem is considered, where several variables are observed on each individual of the control and treatment groups.
Classical MANOVA tests and their derivatives for the many-to-one problem require large sample sizes in order to be powerful if the dimension is high. In this paper, a new class of stabilized multivariate tests proposed by Läuter (2) and Läuter, Glimm, and Kropf (3) is extended to this special design. The new tests are based on linear scores which are derived in a certain way from the original variables. They utilize factorial relations among the variables.
It is shown here that the procedures keep the multiple level. In simulation experiments several versions of multivariate tests are compared with each other. Standard approaches are included as well as different score versions and a comparison of Dunnett-like procedures with Bonferroni-type procedures. Generally, an improved power of the new tests compared to standard procedures is demonstrated.
Power calculations are provided for many of the variables recorded in general toxicology studies. It is hoped that these will help the statistician and scientist to discuss what differences are biologically important. Hypotheses relating to satellite recovery groups (also called off-dose groups) are also discussed.
It is important that statistical tests should be reliable for small samples and rare events. Permutation tests (1) form an important class of tests that can provide “exact” results for small samples. Permutation tests themselves, however, have limitations which must be realized if they are to be used reliably. The paper shows, by illustrative examples, some uses and limitations of permutation tests.
A trend test for dichotomous endpoints analogous to the nonparametric Jonckheere test is developed. The power of this and all other single trend tests for dichotomous endpoints strongly depends on the shape of the dose response curve. Combined tests which have a stable power over a wide range of the ordered alternative are suggested. One can combine several contrast tests to a so-called adjustive test which is more powerful than a Cochran-Armitage test with equally-spaced scores. The latter was recommended by Armitage (1) in case there is no a priori knowledge of the type of the trend.
The nonparametric age-adjusted test proposed by Kodell and Ahn (1) for assessing dose-related trend with respect to the tumor incidence rate in animal experiments is extended from the case of multiple sacrifices to the case of a single (terminal) sacrifice. The tumor incidence rate is made identifiable for time intervals preceding the final time interval by assuming constant proportionality of the tumor prevalences in live and dead animals. Information on cause of death is not required. A Monte Carlo simulation study is conducted to assess size and power of the test.
A widely used approach to analyzing carcinogenicity assays was established by Peto et al. (1). In addition to assessing time to death and the tumor types present at death for each animal, one needs information on the cause of death for each animal to apply this kind of analysis.
The “cause of death” data appear to be very unreliable, and they tend to vary heavily between different pathologists. In other studies, they are not assessed at all, and one tends to analyze these data as if they had “time to death from tumor” instead of “time to death with tumor” information available.
In this paper, a model (see Groeneboom [2]) to analyze “time to death with tumor” data is presented. An overview of the two available sample tests will be given, and the results of a simulation study will be presented.
The classical chronic carcinogenicity bioassay has served researchers well over the last few decades. Recently, however, several factors have prompted a reevaluation of this approach, including budget pressures, a growing list of agents for study, and a better understanding of carcinogenesis. Herein some emerging strategies for obtaining similar or better information are considered. Rather than basing hazard identification and risk estimation on chronic bioassays of two sexes in two species, the possibility of a “reduced protocol” using a subset of the four assays is considered. The possibility of supplementing reduced protocol results from maximum tolerated dose information is raised. The trend to biologically based dose response models is considered, especially with reference to statistical issues associated with the “two-stage birth-death” model and physiologically-based pharmacokinetic models.
The use of a satellite group for pharmacokinetic evaluations in preclinical toxicokinetic studies is costly and prohibits direct correlations between toxicological findings and drug concentrations within the same animal. In studies for which some prior information about the kinetics of the drug is available and where the statistical hypotheses to be tested are relatively few, a sparse sampling “population kinetics” approach to the study design has many advantages. This paper describes how the sparse sampling method was used prospectively for a 30-day oral toxicokinetic study involving five rats of each sex and two sampling days. Population means of the kinetic parameters for a two-compartment model were estimated using the NONMEM mixed-effects modeling package, and the posthoc routine used to produce individual Bayesian estimates for each animal. From the individual kinetic parameters, the noncompartmental parameters: AUC, Cmax, and Tmax, were calculated using SAS.
A three-stage parametric statistical analysis for destructive toxicokinetic studies is described. It provides an alternative to Bailer's method. In the first stage of the analysis either the log-normal or gamma distribution is fitted to observed concentrations. The analysis of variance (deviance) of concentrations is used to test treatment by time interactions and for preliminary assessment of dose-proportionality. In Stage 2, trapezoidal AUC estimates and their variances are calculated from the corresponding maximum-likelihood estimates of each treatment by time concentration. Stage 3 comprises a weighted regression analysis of the AUCs that give tests of dose-proportionality and experimental factors of interest. The tests of dose-proportionality are not new, but they are presented in a form that facilitates graphical presentation.
The choice of gamma or lognormal distribution is justified by a distributional study of four cases. The defects of the Bailer plus Satterthwaite correction method is discussed with the aid of two Monte-Carlo studies. The proposed parametric methods preserve many of the features of the Bailer method but provide superior statistical inference using standard software tools.
This paper provides a brief review of the properties shared by chemical kinetics, clinical pharmacokinetics, and toxicokinetics and on where they differ. The special situation of toxicokinetics and some methods for estimating secondary pharmacokinetic parameters are discussed. The focus is on the application of resampling techniques to toxicokinetic studies and a comparison to the classical approaches. In many cases resampling techniques enjoy the advantage of being easy to perform and easy to adapt to new designs, and there is generally no need to fit the data to certain distributions. Complicated theoretical calculations can be circumvented and new methods, for example, of area under the data (AUD) calculation, may be incorporated without substantial effort. Two resampling techniques, the pseudo-profile based bootstrap and the pooled data bootstrap, are outlined. One of the appealing features of the pseudo-profile based bootstrap is the possibility of estimating summary statistics on the basis of individual secondary pharmacokinetic parameter estimates.
The inhibitory capacity of a competitive antagonist is measured using the pA2 (the negative logarithm of the antagonist dissociation constant). The conventional two-step procedure, based on Schild's regression, neglects a portion of the errors. New single step methods are proposed based on nonlinear regression on dose-effect. Whereas verification of the competition hypothesis often interferes with pA2 estimation, the development of improved nonlinear models could explain and partially incorporate the reasons for noncompetition. The improvements are of two kinds: 1. Pharmacological, based on the inclusion of suspected pharmacological mechanisms such as spare receptors or transduction into a nonlinear fixed-effects model (with examples from the literature); and 2. Experimental, incorporating biovariability as a random effect into a mixed-effects model.
In stability analysis, the current Food and Drug Administration (FDA) recommended procedure for estimating the expiration dating period (shelf-life) of a drug is limited to a single package, single strength product. Since most drug products are manufactured with more than one strength and are marketed in more than one package, stability analyses must be carried out for every combination of package and/or strength. This paper proposes a generalization of the current FDA procedure to analyze the stability data from a multiple package and/or strength study. Monte Carlo simulation was used to address some issues with the current procedure and evaluate the proposed generalization procedure. The proposed procedure is illustrated by an application to a data set consisting of five batches and two packages. Statistical issues and problems with the current approach of concern to industrial statisticians and the generalization are also discussed.
In order to establish a shelf life claim for the potency of pharmaceutical compounds drug stability studies are conducted by pharmaceutical companies. Frequently, the effects on stability of a variety of fixed factors are investigated, leading to large and expensive study designs. Consequently, the use of fractional factorial or matrix designs have received increasing attention from statisticians, with a view toward managing the size of such studies. Fixed factors of interest have been studied with respect to matrixing; however, matrixing on the choice of time points has not been thoroughly investigated. This paper studies the effect on the expiration dating of various matrix designs specifically on the choice of sampling time points, in relation to analytical variation and degradation rate. Some useful tables are also given comparing expiration dates from matrixing across sampling time points in relation to the analytical variation and degradation rate.
STAVEX enables experimenters in research and development to apply statistical design and analysis of experiments in their routine work independently of a statistician. STAVEX presents expert knowledge in industrial statistics in the language of practitioners. The system can handle process factors, such as temperature or concentration, as well as mixture factors which have to satisfy a summing-up condition, for example, typical in galenic formulations.
The STAVEX philosophy is based on the three typical stages of statistical experimental design: screening, modeling, and optimization. During the first two stages the goal is factor reduction. It is recommended to begin experimentation with a large (say more than eight) number of factors. In the final optimization stage, the number of factors is small (three or less); then, a quadratic model is fitted which should allow for good predictions.
STAVEX guides the users through these different experimentation stages, writes reports, suggests a look at appropriate graphical displays (such as half-normal plots or overlaid contour plots) to help in the interpretation and understanding of the results, and provides extensive model diagnostics concerning outlier detection, transformations, and model deviations. STAVEX runs on PCs under Microsoft Windows 3.1, Windows NT, or Windows 95.
A continuous-time Markov model is used to analyze electrocardiogram data obtained from a preclinical study in rabbits of five antiarrhythmic compounds. The preclinical protocol and data are introduced briefly. Some theoretical background for finite-state continuous-time Markov chain models is presented. The electrocardiogram data are then modeled as a continuous-time Markov process with the states being five categories of arrhythmias. The Markov model used assumes that the dwell times in the states are independent and exponentially distributed according to a parameter which depends on the antiarrhythmic compound and the arrhythmia state. For the five antiarrhythmic compounds the transition probabilities from state-to-state and the limiting distributions of the process are calculated and compared.
A Drug Information Association Workshop on “Statistical Methodology on Non-Clinical and Toxicological Studies” was held March 25–27, 1996. The purpose of this meeting was to discuss the appropriateness of current and new biostatistical methods in this field of drug development. This paper describes the application of adaptive interim analysis in pharmacological studies. Bauer and Köhne (1) published a two-step approach in the case of unknown a priori information. This approach is now widely used for clinical trails. Here, the advantages of use in some pharmacological studies will be discussed.