Abstract
Background:
There is currently much interest in generating more individualized estimates of treatment effects. However, traditional statistical methods are not well suited to this task. Post hoc subgroup analyses of clinical trials are fraught with methodological problems. We suggest that the alternative research paradigm of predictive analytics, widely used in many business contexts, can be adapted to help.
Methods:
We compare the statistical and analytics perspectives and suggest that predictive modeling should often replace subgroup analysis. We then introduce a new approach, cadit modeling, that can be useful to identify and test individualized causal effects.
Results:
The cadit technique is particularly useful in the context of selecting from among a large number of potential predictors. We describe a new variable-selection algorithm that has been applied in conjunction with cadit. The cadit approach is illustrated through a reanalysis of data from the Randomized Aldactone Evaluation Study trial, which studied the efficacy of spironolactone in heart-failure patients. The trial was successful, but a serious adverse effect (hyperkalemia) was subsequently discovered. Our reanalysis suggests that it may be possible to predict the degree of hyperkalemia based on a logistic model and to identify a subgroup in which the effect is negligible.
Conclusion:
Cadit modeling is a promising alternative to subgroup analyses. Cadit regression is relatively straightforward to implement, generates results that are easy to present and explain, and can mesh straightforwardly with many variable-selection algorithms.
Keywords
Introduction
When companies mount studies of possible alternative marketing approaches, they often apply the principles of statistical design of experiments. In particular, they may test alternative treatment modalities in two or more randomized groups. Similarly, in medical research for pharmaceutical development, a new product is compared against a placebo (or a currently approved drug). There are obviously many differences between the business and medical contexts in how experiments are conducted. These differences have led to alternative methodologies related to data analysis.
Currently, there is a great deal of creative activity in the development of new analytic techniques to exploit so-called Big Data arising mainly in the business context. In particular, predictive modeling (analytics) methods are being derived to address the increasing awareness that one size may not fit all. The challenge is to “personalize” the choice of treatment, so that it is most appropriate for each individual. We believe that with appropriate precautions, both the spirit and content of these recent developments can play an important role in medical research.
Alternative research paradigms
The classical approach to the analysis of clinical trials can be characterized as rigorous, confirmatory, probabilistic, and conservative. The prime directive is to attain reasonable certainty that the treatment is effective and safe enough in relation to its benefits. To help achieve this goal, traditional statistical methods produce an “average” estimate of causal effect presumed to be globally applicable. However, these methods are of limited use in dealing with possible individual variability of the treatment effect. Because of perceived methodological obstacles to post hoc subgroup analyses, a cautious attitude has become the norm.
A typical approach to subgroup analysis can be summarized as follows:
Conduct a few univariate (single covariate) analyses;
Pre-select the modifiers based on a biological rationale;
Define subjective cut-points for continuous predictors;
Fit a linear model with an interaction between treatment and covariate;
Adjust for multiplicity in computing p values;
Regard significant results of non-pre-specified tests as hypothesis-generating.
Such an approach is unlikely to yield a false positive, but stacks the odds against discovering a true causal modifier.
As practiced outside the medical research context, predictive analytics (also known as data mining) effectively reverses the priorities. The primary emphasis is on improving performance of business processes in a cost-effective way (increase sales, reduce errors, decrease fraud, etc.). Discovery of causal relationships is emphasized more, and ethical issues pose far less of a concern. Of course, considerations of patient rights and safety place important restrictions on medical researchers. However, we believe that a somewhat more discovery-oriented stance is both possible and warranted.
To understand how aspects of predictive analytics might properly shift the balance between discovery and confirmation, consider the problem of data analysis from this alternative perspective. Predictive analytics can be characterized as pragmatic, flexible, exploratory, and empirical. The main steps in building a predictive model are as follows:
Develop a multivariate predictive model or algorithm;
Choose from a diverse set of possible algorithmic techniques;
Consider a wide range of potential predictive features (variables);
Expect to over-fit the data;
Routinely validate on a hold-out or independent sample.
Perhaps the most critical difference between the two perspectives pertains to their conventional standards of evidence.
Traditional statistical methods focus mainly on the level of statistical significance (p value). The primary objective is to be sure that a causal effect truly exists and is not merely a statistical artifact. Consequently, careful experimental design to assure a valid hypothesis test is essential. Post hoc subgroup analyses are subject to various potential biases, and therefore generally mistrusted by biostatisticians.
In contrast, practitioners of predictive analytics often deal with very large samples, so statistical significance may be virtually guaranteed. Moreover, because of multiplicity and over-fitting, conventional significance levels are rarely relied upon anyway. Rather, the touchstone for ruling out chance is successful replication. Accordingly, predictive analysts generally place more emphasis than biostatisticians on validation in a new set of data, such as a hold-out or independent sample.
Many of the recent methodological developments in business analytics pertain to situations that do not entail randomized experimentation. These “machine-learning” algorithms aim to accurately predict an outcome based on various features. Therefore, the techniques do not apply directly to the problem of subgroup analysis in clinical trials. However, some recent developments under the rubric of uplift modeling attempt to adapt standard techniques for essentially this purpose.1,2 We will briefly review these methods and propose a novel approach that may overcome some of their limitations.
Subgroups or models?
The problem of personalized, or precision, medicine is usually couched as a matter of subgroup analysis. However, we can imagine that an intervention’s effect may depend on multiple patient characteristics or circumstances. We will define the individualized causal effect (ICE) as the true average causal effect for an individual who has certain specified characteristics. A variable that affects the value of an individual’s causal effect is called a causal modifier. 3 Subgroup analysis represents a rather crude method of identifying causal modifiers. A more efficient approach in the presence of multiple causal modifiers is to build a model to predict each individual’s causal effect as a function of measured variables or covariates.
The main approach currently used for deriving such functional relationships can be characterized as indirect modeling. Suppose we have collected data from a randomized study comparing Treatment A and Treatment B, including the data on a set of variables (covariates) that may help to predict their probability of a “response.” Then, two separate statistical models can be estimated from the data; one of these is based on the data for individuals exposed to Treatment A and one based on data for those who get Treatment B.
For any individual, two corresponding probabilities,
The most common way to derive an estimate of
Then, two logistic models can be estimated separately from the groups that receive Treatment A and Treatment B, respectively. Thus, estimates of
A variation on the two-model approach to estimating the ICE is to derive a single model that incorporates interaction effects. Such an interactive model accomplishes essentially the same objective as the two-model approach. This interactive model also provides a way to perform a statistical test of significance. The standard procedure for testing whether a factor (e.g. sex, age, etc.) is truly related to the causal effect is to test the significance of its interaction with the treatment variable.
The main problem in using an indirect approach pertains to the difficulty of selecting truly predictive independent variables when there are many potential predictors from which to choose. Most standard approaches aim to choose predictors of the outcome variable, not the ICE. Because they are optimized for that specific purpose, the resulting models may not be satisfactory for the purpose of also predicting the ICE. As a result, the selected predictive variables and their weighting in the statistical model can be misleading.
In the parlance of clinical trials, the objective is to obtain accurate “predictive” classification of patients (e.g. identify who will be cured by the drug), which may differ from good “prognostic” classification (e.g. identify who will recover from the illness).
4
For example, a particular variable might strongly increase the probability of a response (e.g. complete recovery), but have no impact on the true causal effect, because it increases both
Theoretically, a method that could directly estimate the ICE based on relevant variables would be a substantial improvement. Radcliffe and Surry 5 initiated recent attempts to develop such a direct methodology. 6 Their approach is an adaptation of computer-intensive data-mining techniques, such as Classification and Regression Trees (CART) and similar methods. 7 These methods “partition” the datasets into various “nodes” that have different rates of events. The adaptation consists of replacing the event rates in the nodes by the difference in event rates under the two treatments as the criterion used to partition the data.
This modified decision-tree approach makes intuitive sense, especially when the ICE depends on complex interactions of several factors. However, the algorithms needed to implement this approach are very complicated. If the ICE can be represented approximately as a weighted sum of several individual factors, some type of regression model could be much easier to implement computationally and to understand intuitively.
A proposed new approach: cadit modeling
In order to estimate the ICE for a particular person, we would seem to need observations of this same individual under conditions that are identical except for the particular treatment modality. Of course, we cannot hold the situation fixed as we vary the treatment in repeated trials for the same individual. So how can we possibly construct a model to predict the ICE?
One answer is to create an observable variable that is mathematically related to the unobservable ICE in a known way. Specifically, the expectation of this new outcome variable is a known mathematical function of the ICE. We call this variable the causal difference or cadit. We will consider first the version of cadit modeling with a binary dependent variable. The idea of the binary cadit was derived independently by the current authors 8 and by Jaskowski and Jaroszewicz. 9
Suppose that we have data from a clinical trial designed to compare the effects of two alternative treatments on the probability of a particular response. Let T be an indicator variable that takes the value 1 if Treatment A is assigned and the value 0 if Treatment B. Assume further that the treatment groups are designed to be of equal size, so that
Let R be an indicator of whether or not the individual responds. Then, in a randomized controlled trial (RCT)
Our aim is to estimate Δ, where
Now, we have defined a binary cadit variable
Let us denote the cadit
Therefore
An estimate
The real value of the approach described above becomes clear when we consider the possibility that Δ may be a function of certain covariates
For any given value of
Then the conditional ICE is
Now let
So, equation (2) implies
Therefore, in principle, we can condition on the covariates to obtain an estimate of
The general approach outlined above can be implemented in various ways. One method is to estimate
The coefficients in the model can be estimated by standard algorithms and implemented using commonly available statistical software. The resulting model can then be used to score each individual based on his or her values of covariates. To find the estimate
In addition, the statistical significance of the overall model, or of individual coefficients, can be tested in the standard way. These tests then indicate whether or not there are significant effects of the covariates on Δ.
As an alternative to logistic regression, we could use
The cadit methodology for continuous outcomes
We have applied the same basic ideas that drive cadit modeling for a binary outcome to deal with the situation of a continuous outcome. Suppose that
Of course, we cannot observe this ICE directly because each individual is observed under only one of the two treatments.
It is conventional to assume that this ICE is the same for all individuals in the population of interest. Under this assumption, it is a standard practice to estimate this presumed uniform effect Δ by utilizing the following relationship
Here
Suppose, however, that Δ can vary across individuals in accordance with the values of one or more covariates,
It is possible to estimate
Suppose there is a causal effect such that Treatment A is superior to Treatment B (e.g. Drug-X is better than placebo). Then, as noted previously, the standard estimate of this effect
Then the continuous cadit variable is defined as the numerical variable
This can be expressed succinctly as
Note that
Let μ represent the overall true mean in the study population, so that
Then, from equation (4), the expectation of
So,
Note that instead of subtracting
Now, suppose that we want to estimate the conditional effect
Thus, we can estimate
For example, we could frame the problem as a linear regression
Once the coefficients are estimated, the estimate of
Alternatively, we could utilize some other reasonable method for estimating a regression model. For example, we could adapt a data-partitioning method like CART by applying it with
Variable selection
When a small number of covariates
Consider, for example, the use of an interactive logistic regression model. Using standard methods such as stepwise search procedures, we might sort through a large number of variable sets to maximize the value of the AUC. However, this AUC would be a measure of accuracy in predicting the outcome, not the ICE. In general, there is no widely accepted measure of the accuracy in predicting the ICE. Cadit modeling could provide a possible solution to this problem. For example, the AUC for the logistic regression model using cadit as the dependent variable is a direct measure of the accuracy in predicting
Of course, optimizing the accuracy criterion within a training sample often involves serious over-fitting. Therefore, the model may not validate well in external samples or in practice. Many different methods have been proposed to reduce such over-fitting, and there is no widely accepted approach. However, most such techniques effectively trade goodness-of-fit in the training sample for somewhat more robust modeling. As a result, such approaches tend to be conservative. To address this conundrum, we have recently developed an algorithm called Highly Accurate and Robust Variable Evaluation Selection and Testing (HARVEST).
Suppose there are N available variables. To select relevant predictors, HARVEST begins with a large set of randomly selected, equal-sized “batches” of variables. For example, in several applications, we have chosen batches consisting of 10 variables. We then typically construct
To accomplish this objective, we shift our focus from the models to the individual variables. For each of the N variables, the algorithm evaluates its overall performance across all the batches of which it is a member. Note that in our typical implementation, the number of such batches for each variable will have approximately a Poisson distribution with an expectation of 100 and a standard deviation of 10. Therefore, for N very large, each variable is likely to be included in at least 50 batches with virtual certainty and is much more likely to appear in closer to 100.
The evaluation process entails several steps, with the aim of identifying a manageable subset (usually less than 20 variables) consisting of only the most relevant predictors. By relevance, we mean the variable’s ability to improve a model’s performance regardless of which other variables are also included in the model. A more complete description of the HARVEST algorithm and how it can be applied will be presented in a subsequent paper.
Once the subset of final variables is determined, a predictive model can be obtained using standard statistical methodology. This model can be fit on the training sample using any appropriate technique. However, if sufficient data are available, it will be preferable to fit the final model on a second training set. In either case, a final validation will involve estimating the accuracy using data that were not included in the training sample. This validation may include the use of techniques for internal validation (hold-out sample, cross-validation, bootstrapping) or external validation on observations from a completely different source. 12
An example: reanalysis of the Randomized Aldactone Evaluation Study
The Randomized Aldactone Evaluation Study (RALES) clinical trial was designed to test whether spironolactone can reduce mortality in patients who suffer from severe heart failure. Results of this study published in 1999 indicated that the drug was highly effective, with a significant reduction in risk (RR = 0.70; p < 0.001). 13 However, the investigators were aware that spironolactone could increase potassium to unsafe levels (hyperkalemia) in some patients. Consequently, potassium levels for patients in the study were carefully monitored and dose reductions were allowed when deemed necessary. Unfortunately, in a routine care setting, a high incidence of hyperkalemia became evident, discouraging physicians from using spironolactone for heart-failure treatment. 14
We conducted a reanalysis of the RALES data using the cadit approach described above. Our aim was to determine whether we could identify a subgroup of severe heart-failure patients for whom this adverse effect was virtually eliminated. As the safety endpoint we chose the maximum potassium level within the first 12 weeks of treatment. This decision was based on the rationale that the extent of dose modulation to forestall hyperkalemia was minimal during this time frame, thus eliminating the “bias” that affected the original study. Because this outcome was numerical, we applied the cadit method in the continuous version.
After preliminary screening, there were 63 variables as potential predictors of the causal effect. These variables were related to patient demographics, history, and concomitant medications. Unfortunately, we did not have data on genomic or other biomarkers. We derived a predictive model based on a random 80% sample of the total 1632 patients for whom data were available. The remaining 20% were held out to use for model validation. We used our HARVEST variable-selection algorithm to identify 6 of the 63 variables for inclusion in a final regression model, which is summarized in Table 1.
Cadit regression model to predict treatment effect score (training sample).
SE: standard error.
The resulting regression model was highly significant, although the
The model’s predictive validity was tested in the 20% hold-out sample of 332 patients. For each patient, we created a score (predicted value of the treatment effect) calculated using the model described above. This score was then entered in a conventional interactive model that included the score, a treatment indicator, and the interaction between these. The results are summarized in Table 2.
Interactive regression model to predict potassium level (hold-out sample).
SE: standard error.
It is noteworthy that the treatment effect coefficient, which was highly significant in the original RALES study, is no longer significant and is much smaller than the interaction term coefficient. This p value for the interaction term coefficient is 0.088, providing some evidence that the treatment effect may indeed depend on individual characteristics as reflected in the derived score. However, additional data would be needed in this particular case to reach a definitive conclusion. Nonetheless, this example illustrates how a post hoc predictive analytics could potentially yield strong evidence.
Conclusion
Statistical methods for estimating and testing a causal effect depend on comparisons between groups. When the causal effect can be modified by various factors, it may be possible to individualize the estimate of effect by taking account of these factors. The potential value of such personalized or precision medicine is now widely recognized. However, the classical statistical approach is not tailored to meet the challenge of personalization. The primary focus is on main effects applicable to broad populations and on testing of pre-specified hypotheses to effectively rule out chance as an alternative explanation. Discovery of causal modification based on measurable factors is downplayed.
In contrast, predictive analytics focuses primarily on exploration and discovery and is generally optimistic that meaningful individual differences can be found. Therefore, potentially important treatment-by-covariate interactions that may possibly be attributable to chance are not necessarily ruled out of consideration. Rather, they may be tested in a hold-out or independent sample. Carefully done, such validation provides direct evidence that an apparent causal effect, even when discovered serendipitously, may be real and likely to persist in the future.
Furthermore, it may be possible to apply a formal statistical test in the validation sample. Such a test may seem to lack some of the statistical power obtained using all study data together for testing. However, compensation may exist in the form of larger effects in more homogeneous subgroups. Moreover, significance testing in an external sample provides some evidence of generalizability.
A recent important step in this direction is the development of adaptive signature enrichment trials.15,16 These innovative research designs pre-specify a division of the total set of observations into a training and validation sample. The training sample is used to develop a biomarker signature (predictive model). Based on the derived signature, a “sensitive” subgroup within the validation sample is defined. Then, the treatment effect is estimated based on this subgroup. This sort of design can be engineered to preserve complete statistical rigor while allowing for some degree of personalization.
The cadit approach could certainly be applied within the framework of adaptive enrichment. However, we believe that it can also be valuable for post hoc exploratory analysis, perhaps even years after the original study was completed. As long as all the exploratory modeling is conducted only on a random training subsample, significance tests in a hold-out subsample would remain valid. Of course, the statistical analysts should ideally be kept “blinded” to the results of any other subgroup analyses that were previously performed.
We have suggested cadit modeling as one promising alternative to subgroup analyses. Cadit regression is relatively straightforward to implement, generates results that are easy to present and explain, and can mesh straightforwardly with many variable-selection algorithms. In particular, we have successfully applied our HARVEST variable-selection algorithm in this way on several research projects. Cadit and HARVEST are offered as examples of innovative thinking drawing on both statistical inference and predictive analytics. Creatively merging these sister disciplines may be the key to unlock the potential of personalized medicine.
Footnotes
Declaration of conflicting interests
The cadit methodology is protected under US Patent no. 8,688,610 B1, granted 1 April 2014. A patent application for the HARVEST algorithm has been submitted to the US Patent and Trademark Office. The purpose of Causalytics, LLC, in seeking these patents is twofold. First, the authors wish to preclude the possibility that these techniques will be patented by other parties. Second, they seek to recoup the substantial investment of effort expended by Causalytics by receiving a license fee if the methods prove instrumental in developing and marketing a commercial product. However, they do not wish to restrict their use in any other way and plan to freely license these techniques for other research purposes. Any questions regarding this policy should be directed to the first author.
Funding
This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors.
