
Editorial
Select search scope: search across all journals or within the current journal


Statistically based experimental designs have been available for over a century. However, many preclinical researchers are completely unaware of these methods, and the success of experiments is usually equated only with ‘
Animal research often involves experiments in which the effect of several factors on a particular outcome is of scientific interest. Many researchers approach such experiments by varying just one factor at a time. As a consequence, they design and analyze the experiments based on a pairwise comparison between two groups. However, this approach uses unreasonably large numbers of animals and leads to severe limitations in terms of the research questions that can be answered. Factorial designs and analyses offer a more efficient way to perform and assess experiments with multiple factors of interest. We will illustrate the basic principles behind these designs, discussing a simple example with only two factors before suggesting how to design and analyze more complex experiments involving larger numbers of factors based on multiway analysis of variance.
Blinding and randomisation are important methods for increasing the robustness of pre-clinical studies, as incomplete or improper implementation thereof is recognised as a source of bias. Randomisation ensures that any known and unknown covariates introducing bias are randomly distributed over the experimental groups. Thereby, differences between the experimental groups that might otherwise have contributed to false positive or -negative results are diminished. Methods for randomisation range from simple randomisation (e.g. rolling a dice) to advanced randomisation strategies involving the use of specialised software. Blinding on the other hand ensures that researchers are unaware of group allocation during the preparation, execution and acquisition and/or the analysis of the data. This minimises the risk of unintentional influences resulting in bias. Methods for blinding require strong protocols and a team approach. In this review, we outline methods for randomisation and blinding and give practical tips on how to implement them, with a focus on animal studies.
Random treatment assignment is essential in demonstrating a causal relationship between a treatment and the outcome of interest. Randomisation ensures that animals assigned to different treatment groups do not differ from each other systematically, except for the randomly assigned treatment. The randomisation pattern should also dictate the statistical analysis.
The normality assumption postulates that empirical data derives from a normal (Gaussian) population. It is a pillar of inferential statistics that enables the theorization of probability functions and the computation of p-values thereof. The breach of this assumption may not impose a formal mathematical constraint on the computation of inferential outputs (e.g., p-values) but may make them inoperable and possibly lead to unethical waste of laboratory animals. Various methods, including statistical tests and qualitative visual examination, can reveal incompatibility with normality and the choice of a procedure should not be trivialized. The following minireview will provide a brief overview of diagrammatical methods and statistical tests commonly employed to evaluate congruence with normality. Special attention will be given to the potential pitfalls associated with their application. Normality is an unachievable ideal that practically never accurately describes natural variables, and detrimental consequences of non-normality may be safeguarded by using large samples. Therefore, the very concept of preliminary normality testing is also, arguably provocatively, questioned.
Most classical statistical tests assume data are normally distributed. If this assumption is not met, researchers often turn to non-parametric methods. These methods have some drawbacks, and if no suitable non-parametric test exists, a normal distribution may be used inappropriately instead. A better option is to select a distribution appropriate for the data from dozens available in modern software packages. Selecting a distribution that represents the data generating process is a crucial but overlooked step in analysing data. This paper discusses several alternative distributions and the types of data that they are suitable for.
Absence of statistical significance (i.e.,
Variability is inherent in most biological systems due to differences among members of the population. Two types of variation are commonly observed in studies: differences among samples and the “error” in estimating a population parameter (e.g. mean) from a sample. While these concepts are fundamentally very different, the associated variation is often expressed using similar notation—an interval that represents a range of values with a lower and upper bound. In this article we discuss how common intervals are used (and misused).
The purpose of many preclinical studies is to determine whether an experimental intervention affects an outcome through a particular mechanism, but the analytical methods and inferential logic typically used cannot answer this question, leading to erroneous conclusions about causal relationships, which can be highly reproducible. A causal mediation analysis can directly test whether a hypothesised mechanism is partly or completely responsible for a treatment’s effect on an outcome. Such an analysis can be easily implemented with modern statistical software. We show how a mediation analysis can distinguish between three different causal relationships that are indistinguishable when using a standard analysis.
Animal research often involves measuring the outcomes of interest multiple times on the same animal, whether over time or for different exposures. These repeated outcomes measured on the same animal are correlated due to animal-specific characteristics. While this repeated measures data can address more complex research questions than single-outcome data, the statistical analysis must take into account the study design resulting in correlated outcomes, which violate the independence assumption of standard statistical methods (e.g. a two-sample
The theory and practice of statistics comprises two main schools of thought: frequentist statistics and Bayesian statistics. Frequentist methods are most commonly used to analyze animal-based laboratory data, while Bayesian statistical methods have been implemented less widely and may be relatively unfamiliar to practitioners in experimental science. This paper provides a high-level overview of Bayesian statistics and how they compare with frequentist methods. Using examples in rodent toxicity research, we argue that Bayesian methods have much to offer laboratory animal researchers. We advocate for increased attention to and adoption of Bayesian methods in laboratory animal research. Bayesian statistical theory, methods, software, and education have advanced significantly in the last 30 years, making these tools more accessible than ever.
Cage effects: some researchers worry about them, some don’t, and some aren't even aware of them. When statistical analyses do
Pilots are small-scale initial experiments that are intended to guide the design of future, larger studies, with a view to increasing their effectiveness. In this statistical primer we highlight five common mistakes that limit the utility of pilot studies and provide practical guidance to avoid such errors and increase their effectiveness. The common thread connecting these mistakes is insufficient planning and over-interpretation of the results. This approach compromises the ultimate goals of the research programme and the future experimental cascade. In support of our view that over-interpretation is an error, we present a simple simulation to demonstrate that pilots will generally generate an inaccurate estimate of the variability of the biological endpoint under study and that frequent under-estimation will lead to inconclusive and unethical subsequent experiments. We argue that well planned pilots are an important part of the research cascade and still need to be implemented to a high standard.
Null hypothesis significance testing is a statistical tool commonly employed throughout laboratory animal research. When experimental results are reported, the reproducibility of the results is of utmost importance. Establishing standard, robust, and adequately powered statistical methodology in the analysis of laboratory animal data is critical to ensure reproducible and valid results. Simulation studies are a reliable method for assessing the power of statistical tests, however, biologists may not be familiar with simulation studies for power despite their efficacy and accessibility. Through an example of simulated Harlan Sprague-Dawley (HSD) rat organ weight data, we highlight the importance of conducting power analyses in laboratory animal research. Using simulations to determine statistical power prior to an experiment is a financially and ethically sound way to validate statistical tests and to help ensure reproducibility of findings in line with the 4R principles of animal welfare.
Heterogeneity of study samples is ubiquitous in animal experiments. Here, we discuss the different options of how to deal with heterogeneity in the statistical analysis of a single experiment. Specifically, data from different sub-groups (e.g. sex, strain, age cohorts) may be analysed separately, heterogenization factors may be ignored and data pooled for analysis, or heterogenization factors may be included as additional variables in the statistical model. The cost of ignoring a heterogenization factor is an inflated estimate of the variance and a consequent loss of statistical power. Therefore, it is usually preferable to include the heterogenization factor in the statistical model, especially if the heterogenization factor has been introduced intentionally (e.g. using both sexes). If heterogenization factors are included, they can be treated either as fixed factors in an analysis of variance design or sometimes as random effects in mixed effects regression models. Finally, for an appropriate sample size estimation, it is necessary to decide whether to treat heterogenization factors as nuisance variables, or whether the experiment should be powered to be able to detect not only the main effect of the treatment but also interactions between heterogenization factors and the treatment variable.
