Abstract
Toxicologic pathologists assess large data sets from nonclinical studies to identify treatment-related effects to assist in predicting human safety hazards. Statistical testing can facilitate data interpretation by highlighting group differences that have a low probability of random occurrence based on a pre-determined P-value cut-off (eg, P < .05). While this method has been used in the interpretation of pathology data for decades, the appropriateness of utilizing statistical testing in this way has been challenged. Here, we discuss common statistical pitfalls in the analysis of toxicologic pathology data, with emphasis on clinical pathology, reaffirming that appropriate use of statistical analysis requires an understanding of (1) the parameters assessed; (2) the inherent strengths and weaknesses of the statistical method used; and (3) that appropriate interpretation of pathology data is based on the pathologist’s expertise. The presence or absence of statistical significance should not supersede expert judgment but should be one of many tools used to reach a conclusion.
Keywords
*This article is an opinion piece submitted to the Toxicologic Pathology Forum (TPF). This perspective is the particular view of the authors and does not represent an official position of the Society of Toxicologic Pathology (STP), British Society of Toxicological Pathology (BSTP), or European Society of Toxicologic Pathology (ESTP), nor should it be considered to reflect the opinions, policies, or positions of regulatory agencies or the authors’ employer/organization. The TPF is designed to stimulate discussion of topics relevant to regulatory issues in toxicologic pathology. Readers of Toxicologic Pathology are encouraged to send their thoughts on TPF opinion articles or ideas for new discussion topics to
Introduction
General toxicology studies are conducted using the scientific method to test hypotheses about treatment effects and generate reliable information regarding the potential untoward effects of prospective new medicines. Studies are based on standard elements of experimental research, including the use of control groups to hold all variables constant except for that of the treatment, such that changes in the experimental group not found in the control group are generally considered to be caused by the treatment.53,61 Accordingly, current regulatory guidelines on repeated-dose toxicology studies describe the use of a control group to which a vehicle is administered, in line with experimental research using the scientific method (Table 1).
Regulatory Guidance Documents Discussing Use of Controlled Experiments in General Toxicity Studies.
To reliably test hypotheses about causal relationships between independent variables and dependent variables (endpoints), experimental research must adhere to the following elements 64 :
1) manipulation of one or more independent variables: something is purposefully changed by the researcher in the environment;
2) control of extraneous variables to prevent outside factors from influencing the dependent variable;
3) random assignment and random selection to ensure that the groups or treatments are similar at the beginning of the study, to provide confidence that the manipulation (group or treatment) “caused” the outcome.
Null hypothesis significance testing (NHST) is applied to experiments to test a “null hypothesis” (that there is no effect or relationship between the independent and the dependent variables) vs a specified “alternative hypothesis” (that there is a specified effect or relationship). The “null hypothesis” is rejected based on a test statistic (and associated P-value) which reflects how likely it is that the data collected (or an even more extreme data sample) would be observed if the null hypothesis was actually true. 42 The NHST provides a decision framework to test a proposed hypothesis while controlling and balancing the α level (ie, chance of a type I error) and the power (ie, likelihood of detecting an effect; the inverse of a type II error, or 1-β). 45 This approach was originally developed in an industrial quality control setting where factors are more easily controlled and the experiment is likely evaluating only a single endpoint. In an in vivo toxicological research setting, however, effect size cannot be predicted nor is it the same for every endpoint being evaluated. Importantly, in vivo experimental research does not depend on the use of NHST to drive decision-making. Rather, the most rigorous scientific methods utilize many sources of information, including the NHST results, to assist in reaching a decision based on expert judgment.
The statistical analysis of clinical and anatomic pathology data from toxicological safety studies is a long-standing practice that predates the Good Laboratory Practice (GLP) regulations. In fact, GLP regulations were established partially because some statistical procedures used at the time were considered flawed or misleading. 6 However, regulatory guidance documents (summarized in Table 2 and reviewed by Hothorn 39 ) such as GLP regulations (21 CFR Part 5813,14), Food and Drug Administration (FDA) Center for Drug Evaluation and Research (CDER) 23 and Redbook,21,22 European Medicines Agency (EMA),17,18 and Organization for Economic Co-operation and Development (OECD)55,56 do not stipulate specific requirements for statistical analysis performed in routine toxicological safety studies with respect to pathology data. Instead, these guidelines generally recognize the need for flexible statistical approaches that may be required on a case-by-case basis to address specific study designs. 39 These documents do not suggest that data interpretation be based solely on statistical analysis results but rather note the importance of selection of the appropriate statistical method at the time of study design, description of the method used, balance between statistical results and biological relevance, correct use of comparators, and dependencies that must be considered for the statistical evaluation of repeated measures. 39
Regulatory Guidance Documents Discussing Use of Statistical Methods in General Toxicity Studies.
When determining the relationship between treatment and pathology changes observed in general toxicity studies, statistical significance helps to highlight potential trends. However, determining this relationship is not solely based on statistical analysis.29,69,72,74 Animal test systems exhibit high complexity, multidimensionality, and wide variability, both within individuals and among populations. To account for this complexity, interpretation of pathology data for toxic effects requires: (1) the synthesis of other factors that are often more relevant than isolated P-values (ie, weight of evidence [WoE]) and (2) consideration of multiple alternative hypotheses/causes and how much the experimental data support one hypothesis over another hypothesis (eg, ruling out procedural or background findings). 3 Factors that are considered in the WoE approach and consideration of alternative causes are briefly described as follows:
Factors related to the study or data quality include study- or procedure-specific effects that may be either systemic or localized, expected physiologic mechanism related to treatment, artifact introduced via collection technique or timing (especially for clinical pathology data), and methodology (eg, assay, fixation, or tissue staining).
Factors related to physiology of the test system itself may include baseline status of the individuals and populations, patterns of change across time in individuals and groups, observation of corresponding changes in related parameters or organs and/or factors related to physiologic responses to stimuli/stressors/injury. The trained pathologist will be familiar with common “background” or “artifactual” findings that may be present in only a few animals but can introduce bias or the appearance of a false trend, as well as evidence of pathology related to disease conditions that are not expected yet may still be unrelated to treatment (eg, infectious disease affecting an individual or subset of individuals on study).
Assessment of toxicologic pathology data requires an integrated assessment of all study endpoints (eg, exposure and metabolism, pharmacodynamics, in-life observations, clinical pathology, and anatomic pathology).2,3,19,30,62,70 Pathologists use WoE approach based on multiple comparisons (eg, comparison to concurrent controls, pretreatment values, and/or patterns across associated endpoints) to establish pathophysiologically relevant patterns of change. The approach varies depending on the parameter, process, or test system involved and requires an understanding of physiologic, pathologic, and technologic principles. Considerations include, but are not limited to, magnitude/severity, incidence, dose-relationship, known biological variability, half-life in circulation, production dynamics, prevalence as a background or spontaneous finding, potential relationship to experimental procedures, and collection timing and technique. Additional considerations for numeric endpoints include the absolute value, influence from intercurrent and concurrent factors, analytical factors, and time course (eg, timing of collection relative to dosing and other procedures, transient vs persistent throughout the dose period, and analyte kinetics).
Over-reliance and over-confidence in statistical significance are common in toxicologic pathology (and more broadly in the scientific community), despite the recognition that statistical significance cannot address every consideration for the appropriate interpretation of toxicologic pathology data.7,25,65 For example, the Environmental Protection Agency (EPA) weighs statistical significance heavily in the identification of treatment-related effects, requiring that all statistically significant findings be addressed even if not biologically relevant 75 and that discordant interpretations be justified (personal communication). Interpretations based on “star searching” for flagged statistically significant differences (often indicated with an asterisk “*”) can risk missing important treatment-related findings that may not be statistically significant (ie, false negatives [FNs] or type II errors) or may produce findings that are not biologically relevant or meaningful (ie, false positives [FPs] or type I errors). A type I error (ie, FP) occurs if a result is statistically significant, but the null hypothesis is true (eg, the effect size is not biologically relevant), whereas a type II error (ie, FN) might occur because the test is not powerful enough to detect a biologically relevant difference.
When the P-value cut-off (P < .05) was introduced, it was not meant to be a definitive test result; rather, it was proposed as an indicator of the probability that an observed difference between study results could have occurred randomly if no treatment effect is present.15,16 Use of P-value cut-offs was intended to be an expedient and somewhat arbitrary stimulus that would lead to further investigation. 40 However, P-values have inherent variability, and treating them as definitive cutoff values is inappropriate both mathematically and biologically. Repeating an experiment multiple times will yield a range of P-values simply due to random variation, and there is essentially no difference between a P = .049 in one study and a P = .051 in another study. In fact, the use of a P-value cutoff of .05 can sometimes lead to claims of different results from two studies when no real difference exists. As with the previous discussion of type I and II errors, because a P-value is only a probability, a statistically non-significant result does not “prove” the null hypothesis, and a statistically significant result does not disprove it nor does it “prove” the alternative hypothesis.
Some scientists and statisticians have proposed eliminating the use of P-value cut-offs due to the common misuse of statistical significance testing resulting in “hyped claims and the dismissal of possibly crucial effects.”1,43,77,50,65 Many emphasize the crucial role of expert judgment and subjective, evidence-based decision-making.10,58 Over-reliance on statistical significance has also received regulatory attention. The National Institutes of Statistical Sciences (NISS) hosted a Webinar entitled “Alternatives to the Traditional p-value” 52 in response to the ASA’s statement on p-Values: “Context, Process and Purpose.” 77 Questioning the historical reliance on statistical P-value interpretations (significance) becomes even more relevant to toxicologic pathology as global laboratory animal use considerations continue to shrink the laboratory animal group sizes in toxicological studies, thereby reducing the power of statistical evaluation.
It is our opinion that the statistical analysis of toxicology study data is a helpful tool but should not be used to replace biological relevance or the judgment of the toxicologic pathologist or other subject matter expert. Biological relevance carries stronger weight than statistical significance and is sufficient as a stand-alone justification to call a difference as treatment-related. The toxicologic pathologist is uniquely qualified to perform a WoE interpretation of the statistical and biological relevance of pathology findings in toxicology studies because of years of specialized comparative medical training, a deep understanding of pathophysiology and toxicology, and professional expertise and experience gained over the course of a career.
This article aims to provide considerations for the interpretation of statistical significance testing results from anatomic and clinical pathology data in general toxicology studies. Other study types such as carcinogenicity studies and exploratory studies are beyond the scope of this article. The application of statistical analysis to carcinogenicity studies has been extensively explored and reviewed.4,23,35,37,39,49,60 In addition, this article does not intend to make recommendations on what statistical tools should be used in general toxicology studies but communicates a professional opinion on the limitations of statistics in study data interpretation, illustrated by specific examples in clinical and anatomic pathology data.
Key Concepts
Toxicology studies are conducted using the scientific method to test hypotheses about treatment effects and generate reliable information regarding potential untoward effects of (potential) new medicines. While statistical analysis can be a useful tool, scientifically sound methods of experimental research do not require the use of statistical testing (or NHST) to make reliable conclusions regarding toxicologic pathology in general toxicity studies.
Biological relevance, as ascertained by subject matter experts with specific training/credentials, is the key determinant for interpretations in toxicologic pathology. Whether a statistical analysis has been performed or not, the most rigorous scientific methods assess biological relevance and employ expert judgment in decision-making. Statistics cannot replace a thorough evaluation of data and must not be relied upon as the sole or primary indicator in determining biological relevance (eg, whether an effect is treatment-related or not).
Statistical significance does not mean a compound-related effect is present, and the lack of statistical significance does not rule out a compound-related effect.
It is recommended that scientists applying statistical test results to toxicologic pathology endpoints understand the strengths, weaknesses, and limitations of the statistical methods being employed to avoid erroneous conclusions about the presence or absence of real effects.
Statistical Principles
Statistical testing performed in general toxicology studies is either descriptive or inferential. Descriptive statistics characterize data in terms of location (ie, mean, median), dispersion (ie, standard deviation, interquartile range), and sample size (ie, count, n). Inferential statistics use data to make inferences about the general population. In toxicity studies, inferential statistics are based on NHST to compare 2 or more groups, resulting in a P-value between 0 and 1. The goal of NHST is to determine if the observed data would likely be generated by chance alone, given the stated hypothesis. Hypothesis testing does not “prove” that there is a treatment-related effect, nor is its goal to “achieve” statistical significance. Importantly, NHST does not allow for the consideration of multiple alternative hypotheses/causes and how much the experimental data support one hypothesis over another hypothesis (eg, procedural effects or background findings). 43
The choice of a specific statistical test depends on the type of data as well as statistician and/or institutional preference. For inferential tests, most institutions perform parametric testing (most commonly one-way analysis of variance [ANOVA] or t-tests; less commonly Welch’s test, and 2-way ANOVA).40,51 Alternative approaches (eg, nonparametric methods [Mann-Whitney test, Chi-square test, Friedman test, and Kruskal-Wallis test], trend tests, adjustment for multiplicity of tests) can be particularly powerful when properly applied. Each statistical test has advantages and disadvantages and requires that specific assumptions be met. Statistical errors can occur when model assumptions are violated.
Decision trees may be employed to determine the appropriate statistical model for a given set of data (ie, so that model assumptions are appropriate for the data). However, there is disagreement among statisticians regarding this approach. One disadvantage of a decision tree approach is that a given parameter may have different tests performed at different time points. Since the different tests have different powers, this can make interpretation confusing or lead to inaccurate conclusions. Of course, the same concern is present whenever sample sizes change among sampling times (eg, main or dosing phase versus recovery phase). Although these approaches can yield robust results, they generally require statistical expertise or consultation with a statistician, and they do not negate the need for thorough individual animal-based review.
Types of Data
The difference between categorical and numeric data is well-recognized. The type of data collected determines the approach for descriptive and inferential statistics, and the statistical power that can be achieved (Table 3).
Types of Data and Statistical Power That Can be Achieved.
Categorical data
Categorical data can be nominal or ordinal (ie, ranked). Nominal data have no natural rank order (eg, urine color or can be dichotomous (ie, present or absent). Ordinal data are ranked, ordered or graded observations (eg, 1+, 2+, 3+, 4+); however, the difference between adjacent ranks is not a consistent value and may not be meaningful. With nominal data, one can count the number of occurrences and derive the frequency of each value in the sample, whereas with ordinal data, one can count the number of occurrences and derive the frequency, but in this case, the mode (ie, the most common finding) and the median of the values will also have meaning. Many of the commonly employed statistical tests that are appropriate for continuous data are not appropriate for nominal or ordinal data; the required assumptions for these tests are not met.
Most macro/microscopic observations are ordinal categorical data reported as incidence and severity. Statistical analysis of severity grading scores is not appropriate 26 ; the variability between ranks, combined with the small group sizes used in regulatory toxicology studies, yields analyses with low statistical power that provide limited value for the interpretation of study findings. 47 Meaningful changes may be readily apparent even if they occurred in only a few animals or a single animal, regardless of the result of the statistical analysis. In addition, statistical analysis may not flag a treatment-related exacerbation of spontaneous findings. For example, many chemicals are known to exacerbate spontaneous chronic progressive nephropathy in rats. 32 Severity as well as incidence and dose-response help to facilitate differentiation of a treatment-related increase in findings from background changes that are commonly present in controls. It has been recommended that severity grades be assigned based on the extent of the morphologic change and that they should be clearly defined in the report narrative for lesions impacting study interpretation.26,67 However, criteria for histopathology severity grading are generally not pre-defined for regulatory toxicology studies, unlike bespoke criteria used in pharmacologic efficacy or animal model studies, or diagnostic criteria used in clinical cancer grading in human patients.
All tests measured using reagent pads (including ketones, proteins, and glucose) are semi-quantitative (categorical ordinal), even if results are reported as numbers and concentrations (Table 4). Importantly, the difference between increments detected by reagent pads varies. For example, pH (eg, urine) is the negative logarithm of the hydrogen ion concentration ([H+]). Therefore, a pH change from 6 to 7 translates to a 9 × 10−7 mol/L change in [H+], while a pH change from 7 to 8 translates to a 9 × 10−8 mol/L change in [H+] (Table 5). This relationship means that statistical analysis of pH requires unique methods. Reagent strips subdivide pH into categories (eg, ≤.g., 6.0, 6.5, 7.0, 8.0, and ≥9.0), thus converting the continuous biologic measure to a categorical ordinal one. Not only should parametric tests not be used to assess pH, but also summary descriptive statistics should report median or incidence median rather than mean and fold change.
Reportable Dipstick Urinalysis Values Demonstrate Their Categorical Ordinal Nature.
Na: not applicable.
Nonlinear Change in Hydrogen Ion Concentration With Linear Increases in pH.
Numeric data
Numeric data may be discrete integers or continuous, wherein they have infinite precision. With discrete numeric data, one can define incidence and derive the frequency, mode, median, and range, and can quantify differences by adding and subtracting values, or define percent change by dividing values. Continuous numeric data with values that are far from zero (eg, organ weights and most serum chemistry endpoints) most effectively leverage the power of statistical models.
Analysis of organ weight data usually includes three different endpoints, namely absolute organ weights and organ weights adjusted for body and brain weight. These are continuous non-zero numeric data that are generally normally distributed and, as such, are readily amenable to statistical analysis.
The assessment of organ weights can be highly sensitive to treatment-induced changes, even in the absence of microscopic changes. When interpreting organ weight data, pathologists generally refer to the magnitude of changes, the presence of correlates with gross pathology, histopathology, and clinical pathology, and the possibility of individual versus group effects to build a WoE in determining the test article relationship of changes. Statistical significance may be leveraged as an additional piece of evidence to support the interpretation of a group effect. However, statistical significance, whether it is in only one or more than one of the three analyses (absolute value, relative to body weight, or relative to brain weight) is not necessary or sufficient to conclude that there is a test article effect. 5 Although statistics are commonly utilized in the evaluation of organ weight data in general toxicology studies, organ weights may also be interpreted with only descriptive statistics (eg, individual animal data, number of animals evaluated, mean, standard deviation) in conjunction with other study data. 68
Statistical Assumptions and Issues
Assumption of Independent Random Sampling
Commonly used statistical tests require that samples be independent of one another and randomly selected. Differences between groups must be attributable to chance alone, each member of the population must have an equal chance of being selected, and the values from one animal should be independent of values from another animal. In toxicology studies, animals are generally randomly selected. Attempts are also made to randomize group allocation, eg, based on age, body weight, or an endpoint of interest such as baseline glucose for a glucose tolerance test. However, the assumption of independent random sampling is sometimes not met, or values may be missing in a non-random manner (see Table 6). For example, animals with abnormal baseline (pretreatment) values may intentionally be placed in control groups. Although lack of independence can be addressed using more sophisticated statistical methods, these are generally not performed. Independence and random sampling are best addressed during study design.
Common Sampling Errors.
Assumption of Normality
Many commonly used statistical tests, such as the ANOVA and t-tests, assume that the data are normally distributed. Organ weight data are usually normally distributed, but clinical pathology results often are not. Even for endpoints with normally distributed historical control data, study results may not reflect this since the sample size is too small. It takes 30 to 120 samples to achieve a Gaussian distribution of random independent samples. 48 In these cases, nonparametric tests or data transformations can be used. However, these alternative statistical methods are not always applied.
Assumption of Homogeneity of Variance
Another common statistical model assumption is homogeneity of variance, ie, comparability of the variability, or spread of the data, across all groups including control groups. Homogeneity of variance may be lost when there is a treatment effect; as treatment effects increase, it is likely that only a portion of the treated animals are affected, and the magnitude of effects in the affected animals increases, both of which cause the spread of the data to increase. A variety of tests can be used to confirm homogeneity of variance. If this assumption is not met, there are statistical tests that account for heterogeneity of variance, or the data can be transformed and appropriate statistical models applied. However, these alternative statistical methods are not always applied. Approaches to statistical evaluation of toxicologic pathology data are not uniform across organizations, and there are no standardized or generally accepted approaches. In one survey of laboratories performing preclinical/nonclinical toxicology studies in the United States and Europe, 76 less than half of the respondents queried tested for homogeneity of variance and none of the organizations tested for normality of distribution, yet most of the organizations used parametric statistical tests consistent with the assumption that the variance of their data was homogeneous.
Outliers
A discussion on guidelines for eliminating outliers from a statistical analysis is beyond the scope of this article; however, it should be mentioned that large outliers can have a large, unwanted effect on the results of statistical analysis. A large outlier, particularly when combined with small sample sizes, can have a sizable effect on the mean and standard deviation values for that group. Also, a large outlier may affect the sensitivity of some commonly used statistical tests such as ANOVA, as the tests use the overall variability of the data in the calculations and a large outlier will inflate this value, thereby making the test less sensitive. Most nonparametric approaches, because they use ranks of the data instead of the observed data value, are not seriously impacted by outliers.
Statistical Power
Statistical power analysis is performed to estimate the smallest sample size needed for an experiment, given (1) a required significance level (eg, P-value cutoff < .05); (2) statistical power (usually 80% or 90%); and (3) effect size regarding a single endpoint of interest with a unique degree of variability. Power and significance level for rejecting the null hypothesis are chosen based on the acceptability of type I and II errors. 44 In general toxicity studies, the sample size is predetermined and constant across multiple different endpoints, each with different characteristics (eg, direction of change, variability, and dynamic range). Therefore, the size of an effect that can be detected may not be aligned with biologic relevance, possibly producing both type I and type II errors for different parameters within a given study.
For example, a power analysis using a two-tailed t-test, a power of 80%, and a significance cutoff (ie, P-value cutoff) of 0.05 yields the number of animals needed in each group to detect the desired signal for selected endpoints (Table 7). Endpoints with a high signal-to-noise ratio require fewer animals (eg, creatinine, alanine aminotransferase, alkaline phosphatase, chloride), and those with a low signal-to-noise ratio require more animals (eg., testicular weight, neutrophil counts) to achieve the same statistical power. In toxicity studies, a constant sample size is applied to all endpoints in the study, and the signal that is detectable using statistical significance testing may be smaller or larger than the desired effect size. The results in Table 7 demonstrate that statistical significance operates on a sliding scale that is dependent on characteristics unique to each endpoint, such as biologic variance and biologically relevant effect size. 8
Power Analysis Example.
Signal (effect size) shown as % difference from controls. Assuming a two-tailed t-test, a power of 80%, and a significance cut-off of .05.
Noise (CV%, coefficient of variability) = approximation of background variability unique to each endpoint, adapted from Har et al, 2013. 34
Desired signal = minimal size of a biologically relevant effect.
Italics: detectable signal is smaller than desired effect size (type I error, false positive).
Bold: detectable signal is larger than the desired effect size (type II error, false negative).
ALT: alanine aminotransferase; MCV: mean cell volume; HCT: hematocrit.
If the detectable signal is smaller than the desired effect size, statistical significance testing will flag differences that are not biologically relevant (ie., type I error, FP). This occurs for endpoints with a high signal-to-noise ratio when group size is too large. Studies may be overpowered for endpoints with a high signal-to-noise ratio, resulting in statistically significant findings that are not biologically relevant. Instances with low background variability (noise) can also yield FP results due to overpowering relative to biologically relevant effect sizes (see italicized results in Table 7) since the detectable signal is smaller than the desired effect size. This phenomenon is rarely recognized but is a common source of statistically significant differences that have no biologic relevance.9,12,31,38,46
If the detectable signal is larger than the desired effect size, small but biologically relevant differences may not be flagged, resulting in a statistical type II error (ie, FN). This occurs for endpoints with a low signal-to-noise ratio when the group size is too small. General toxicity studies (particularly nonrodent studies) are underpowered for many endpoints, and statistical significance testing may not be able to detect biologically relevant differences. An example is large interindividual differences in testicular weight due to variation in age at puberty in peripubertal monkeys and dogs used in toxicity studies. 34 Additionally, for nonrodent studies with small group sizes, interpretations may be based on changes from baseline values in addition to comparison to concurrent controls. Although statistical tests that compare with baseline values may be used, these can be complicated by multiple and/or staggered baseline values. Ultimately, the final interpretation requires the combined use of multiple comparators (eg, individual baseline, timecourse, and changes in other animals), an examination of patterns of change across multiple endpoints in an animal, and knowledge of expected variability in the context of the study. 3 As such, statistical analysis of clinical pathology results may not be useful or even necessary in nonrodent studies.
Use of P-values for interpretation
There has been extensive discussion in the literature over the past number of years about the use and misuse of P-values (see the special edition of the American Statistician entitled “Statistical Inference in the 21st Century: A World Beyond P < .05). 71 This discussion has been supported by the statistical profession and applies to all disciplines, not just biological applications. Unfortunately, this discussion is sometimes misinterpreted to mean that statistical analysis of the data should be abandoned, which is certainly not the intended message. The concern expressed by the statistician community and others is that the current use of P-values, particularly the black/white reliance on P-value cutoffs, is flawed.
The criticism is not new and has been repeatedly expressed over the last 50 years. In the absence of agreed alternatives and insufficient consideration of the many different applications of the concept of statistical significance, we consider the demand for its abolition to be exaggerated. A more pragmatic approach to the issue, supported by targeted instructions for scientists and reviewers, seems to be a more appropriate way forward. 63
In the immediate future, we encourage implementation of a more appropriate utilization of P-values. One critical aspect of this approach would be to stop flagging significant results and instead report the actual P-values. The reader can then interpret exactly how likely or unlikely the tested event would be under the null hypothesis of no difference. A P-value of .045 and a P-value of .0001 do not convey the same information about the likelihood of an event; however, if they are both flagged as being < 0.05, important information is lost. Also, as has been mentioned previously, the statistical analysis should be just one of many pieces of information that the pathologist utilizes to construct their interpretation using a WoE approach and should not be relied on as the sole decision point.
In the longer term, we encourage the development of the next generation of statistical tools that more closely mirror the pathologist’s approach to interpretation. This includes such approaches as a more formal use of historical control data in the analysis, use of multivariate analysis approaches, use of Bayesian analysis for analysis of rare events, and implementation of an estimation of the effect level to help determine if a statistically significant result is biologically meaningful, as well as other new ideas.11,20,36,79 Many new methods are currently in development; however, implementation of new approaches is often slow because it requires the support of regulatory agencies for broad acceptance.
Discussion
“Everybody believes in the exponential law of errors: the experimenters, because they think it can be proved by mathematics; and the mathematicians, because they believe it has been established by observation” (quote attributed to physicist Jonas Ferdinand Gabriel Lippmann). 78 When misunderstood or misused, significance testing in general, and NHST in particular, can be an interpretive crutch that leads to erroneous conclusions.25,27 Statistical significance does not mean that an experiment worked, or that a change is test article-related. It does not indicate that a difference is big, small, or even meaningful. Misuse of statistical significance may even cause pathologists to spend precious effort to write lengthy and unnecessary justifications for interpretations to explain discordance of their conclusions with statistical significance results.
In general toxicity studies, investigators must make clear decisions regarding the treatment relationship and biological meaning of findings, and statistical significance is often used to aid in flagging potentially test article-related trends.24,28,66 There are many statistical approaches available, each with its strengths and weaknesses. Consequently, any given statistical approach may produce anomalous results in any given scenario. When statistics are used, no matter the approach, scientists and reviewers should understand their appropriateness for the data set in question as well as their strengths and limitations. A prerequisite to their proper use is an understanding of statistical logic, principles of uncertainty assessment and P-value interpretation, and limitations introduced due to the statistical model assumptions.41,73 Therefore, investigators should work closely with statisticians to effectively apply statistics to toxicologic pathology data.
Statistical analyses are intended to allow scientists to quantify uncertainty. 59 They are based on mathematical probabilities calculated based on a set of foundational assumptions, which may or may not reflect the complexity, multidimensionality, and wide variability of data produced from biologic systems. To account for this complexity, toxicologic pathologists use a WoE approach that considers alternative causes and integrates biological relevance and other factors that are more important than P-values, such as study procedures, test article class and mechanism of action, and background variability/findings that are unrelated to treatment.
The pathologist brings a unique set of training credentials that allows for medical, physiological, and pathological context to be fully considered when determining biological significance. General toxicologic pathology data analyses and interpretation must remain grounded in biologic and pathophysiologic principles even as new techniques emerge, such as digital pathology, machine learning, high-throughput screening, in vitro toxicology, -OMICs, computational biology, and molecular pathology. Therefore, pathologist participation in the development of these tools is critical to ensure the proper design and implementation for success in meeting scientific and pragmatic needs.
Conclusion
Biological relevance, as determined by toxicologic pathologists, is a key factor when assessing potential treatment-related pathology changes in general toxicity studies. The assessment of biological relevance requires (1) an integrated assessment and WoE approach of multiple endpoints and factors that are more important than P-values and (2) consideration of potential alternative causes such as procedural effects or background findings. Statistical analysis inevitably produces type I and II errors (ie, FP and FN results) because of its mathematical nature and logical assumptions. A statistically significant change may not be biologically relevant, and a biologically relevant change may not be statistically significant. Therefore, statistical significance helps to highlight potential trends but cannot replace a WoE approach where all factors are considered. Statistical significance must not be relied upon as the sole or primary indicator in determining treatment effects.
Footnotes
Author Contributions
The analysis, conclusions, and opinions expressed in this article are solely those of the authors (LR, TA, GT, NS, MS, CS, SB).
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
