Abstract
Beger, Morgan, and Ward (BM&W) call into question the results of our article on forecasting civil wars. They claim that our theoretically-informed model of conflict escalation under-performs more mechanical, inductive alternatives. This claim is false. BM&W’s critiques are misguided or inconsequential, and their conclusions hinge on a minor technical question regarding receiver operating characteristic (ROC) curves: should the curves be smoothed, or should empirical curves be used? BM&W assert that empirical curves should be used and all of their conclusions depend on this subjective modeling choice. We extend our original analysis to show that our theoretically-informed model performs as well as or better than more atheoretical alternatives across a range of performance metrics and robustness specifications. As in our original article, we conclude by encouraging conflict forecasters to treat the value added of theory not as an assumption, but rather as a hypothesis to test.
In our article “Forecasting Civil Wars” (Blair and Sambanis 2020, hereafter B&S), we sought to make three contributions to the literature on conflict forecasting. First, we explored whether incorporating theoretical insights into predictive models improves forecasts of civil war. We did this by building a model grounded in “procedural” theories of conflict escalation from the social movements literature, then comparing the predictive performance of this escalation model to the performance of more mechanical, inductive alternatives. Second, we considered whether incorporating “structural” characteristics of countries, such as regime type or per capita income, might improve the predictive performance of our escalation model, which was otherwise based on procedural variables alone. Finally, we used the escalation model to generate “true” prospective forecasts for the first half of 2016, which we preregistered with the Evidence in Governance and Politics (EGAP) network. 1 We returned to these forecasts to evaluate their accuracy with the benefit of hindsight. We found that the theoretically-informed escalation model generally outperformed more atheoretical alternatives (though the margins were often small); that adding structural characteristics to the model did not significantly improve predictive performance, especially over narrow forecasting windows; and that our pre-registered predictions were generally fairly accurate.
Beger, Morgan, and Ward (2021, hereafter BM&W) replicate and modify our analysis and challenge our conclusions. Their article can be read as a defense of machine learning in conflict forecasting, and the crux of their argument is that the theoretically-grounded escalation model under-performs more mechanical, inductive alternatives. We certainly agree that machine learning can improve predictive performance, which is precisely why we applied machine learning methods in our original article. What is puzzling is that, while BM&W’s paper is entirely dedicated to proving the purported superiority of more atheoretical approaches to conflict forecasting, they end up supporting our own conclusion when they “concede that predicting on the basis of a strong theoretical model is preferable to inductive prediction” (p. 19). We are pleased that they agree with our intuition, but their conclusion is not supported by their analysis. It is, however, supported by ours.
BM&W raise six empirical critiques of our study. We consider these in detail and show that five of them are either misguided or largely inconsequential: even if we had embraced all of them before publishing our original article, they would not have changed our substantive conclusions. They are also irrelevant to the question of whether theory is useful for prediction, as they have no bearing whatsoever on our comparison of more and less theoretically-informed models. In any event, we do not accept all of BM&W’s critiques: some are simply false, and others are based on questionable premises or obvious mischaracterizations of our arguments. We do admit to making two minor coding errors, which BM&W identify, but correcting these errors does not change our substantive conclusions.
The only remaining point of difference—and the only one that has any consequences at all for our comparison of the escalation model to more mechanical, inductive alternatives—is also relatively minor: it hinges on whether “empirical” or “smoothed” receiver operating characteristic (ROC) curves should be used to assess predictive performance. We used smoothed ROC curves; they use empirical ones. As we discuss in detail in the appendix, we decided to use smoothing during preliminary work—before we had finalized our models, our data, or even the outcome we were trying to predict—because the empirical curves proved to be so “jagged” and discontinuous that they were difficult to interpret. BM&W nonetheless prefer the empirical curves, and make three claims: (1) that a “low-effort” atheoretical model outperforms the escalation model “in all instances;” (2) that a simple ensemble model outperforms the escalation model “in all instances” as well; and (3) that a hybrid model that incorporates structural variables “strictly dominates” the escalation model alone (p. 19). As we demonstrate below, these conclusions depend entirely on BM&W’s use of empirical ROC curves. Contrary to their assertions, the escalation model performs as well as or better than the alternatives across a range of different performance metrics and robustness specifications.
Indeed, ROC curves (smoothed or empirical) are just one of many metrics one might use to evaluate predictive performance. In what follows we present an expanded analysis that asks what happens if, for the sake of argument, we grant BM&W’s critiques in order to probe the sensitivity of our results using additional performance metrics and robustness specifications. We find that the escalation model almost always performs as well as or better than the alternatives across a variety of widely used metrics, including precision-recall curves, mean and maximum F1 scores, mean and maximum F2 scores, Brier scores, and smoothed ROC curves. Each of these metrics has pros and cons, as we explain below. The escalation model also performs as well as or better than the alternatives on the MCC-F1, a new measure that was recently proposed as an alternative to the ROC curve because it can distinguish good from bad classifiers even with class-imbalanced data across a wide range of classification thresholds. Indeed, of the various metrics we test, the only one according to which the escalation model clearly under-performs the alternatives is the empirical ROC curve.
Finally, we show that the escalation model performs as well as or better than the alternatives when we replicate our prospective forecasting exercise using all of our models (rather than just the escalation model, as in B&S), or when we consider additional permutations of the model parameters and hyperparameters. BM&W’s first two claims thus hinge entirely on a single subjective modeling choice. To assess their third claim, we extend the analysis in B&S to show that the inclusion of structural characteristics has ambiguous effects on the escalation model’s performance. Here again, we find that BM&W’s conclusions are completely dependent on their use of empirical ROC curves, though our extensions of this analysis leave us somewhat more optimistic about the prospects for hybrid models of this sort. Importantly, since the structural variables in our hybrid models were selected on the basis of theory (Goldstone et al. 2010), any resulting improvement in the escalation model’s performance only supports our conclusions about the value of theory for prediction.
Why Conflict Forecasters Should Treat the Value of Theory as a Hypothesis to Test
Before we turn to the analysis, we want to make clear that the main contribution of our original article was to introduce the idea that researchers should formally test the usefulness of theory for prediction, rather than dismiss theory out of hand, as some have done in the past (Ward 2017), or, conversely, assert that theory is necessary without proof (Beger, Morgan, and Ward 2021). Although we believe that focusing on process rather than just structure is important in building predictive models of conflict over short time intervals, we are not wedded to this particular model, and we do not claim to have made any significant theoretical breakthroughs in building it. Nor are we wedded to these particular atheoretical alternatives, or even this particular method for generating forecasts. There are many alternatives that one might test.
BM&W criticize our original article for failing to provide “law-like evidence for general theory” (p. 19). But this is obviously an elusive and inappropriate standard to apply to a single study. Our paper was clearly intended as a small step in a larger effort to formally compare more and less theoretically-motivated models in conflict research. Members of the BM&W team (along with other researchers whose work we cite) have, in previous studies, invoked theoretical arguments and intuitions as the motivation underlying their forecasting models (Beger, Dorff, and Ward 2016; Chiba, Metternich, and Ward 2015; Gleditsch and Ward 2013; Metternich et al. 2013; Weidmann and Ward 2010). But as we noted in both the introduction and conclusion of B&S, whether theory is necessary for prediction should be treated as a hypothesis to test, not as an ad hoc assumption. Ironically, we suspect BM&W would agree with this proposition. We may have provided the first such test, but it certainly is not going to be the last.
Given that BM&W devote much of their paper to attempting to show the purported inferiority of theoretically-grounded models of conflict, it is surprising that they conclude by arguing for the usefulness of theory for prediction after all—a position that is very similar to our own. They attempt to explain this apparent contradiction by drawing a distinction between “weak” and “strong” theory (p. 16); ours, they claim, is a “weak” theory, when what’s needed is a “strong” one. But this distinction is of little practical use, and it strikes us as unreasonable and somewhat disingenuous. Indeed, BM&W judge most empirical social science research to be deficient according to their standard of “strong” theory (p. 16)—a verdict that would surely apply to much of their own previously published research. 2
We do not wish to engage in an unproductive debate over what should count as “good” or “bad” theory. To reiterate, we do not believe—nor did we claim—that we broke new theoretical ground in our original article. We do believe that our use of the term “theory” was consistent with prevailing practice in our field; indeed, BM&W concede as much when they note that “most empirical social science research” follows an approach similar to ours (p. 17). BM&W also attempt to distinguish their argument from ours by pointing out that even our more atheoretical models necessarily incorporate some amount of theory, since they were built using data that must have been gathered with some theoretical purpose in mind (p. 17). But we already made this exact same observation in our original article. 3 On this point, BM&W’s substantive argument turns out to be strikingly similar to our own.
In any event, most reasonable people will recognize that there is, in fact, more theory guiding the escalation model than there is in an alternative that maximizes predictive performance by mechanically selecting among hundreds of variables, many of which have little to no a priori connection to the outcome under study. We reiterate that it is of course possible to construct different empirical specifications that operationalize the theoretical logic underlying the escalation model in different ways. We do not believe we have ended the debate regarding the value of theory for prediction, and we do not aspire to do so. But we do think we have pushed the discussion forward by asking the question clearly and offering an answer, however incomplete that answer may be. It should be obvious to readers of this journal that the broader question of the value of theory for prediction cannot possibly be settled by a single article, just as it is clear to us that the narrow scope and misguided criticisms in BM&W do not help advance this important debate.
Does Theory Improve Predictive Performance?
In B&S we tested whether theory improves predictive performance by using ICEWS event data to compare a theoretically-informed escalation model to more atheoretical alternatives. These alternatives included models based on “Quad” counts 4 and “Goldstein” scales, 5 as well as a “kitchen sink” CAMEO model comprising over 1,000 cooperative and conflictive events between government, opposition, and rebel groups, ranging from “make pessimistic comment” to “consider policy option” to “bring lawsuit against” to “kill by physical assault.” We also generated forecasts based on a simple unweighted average of the predicted probabilities from these latter four models. We compared our five models—escalation, Quad, Goldstein, CAMEO, and average—using a variety of performance metrics, focusing in particular on receiver operating characteristic (ROC) curves.
The main point of contention is our approach to computing the area under the ROC curves (AUC-ROC). We smoothed our ROC curves before estimating the area under them. BM&W criticize this decision, arguing that we should have used “empirical” ROC curves instead. While there is much less consensus on whether or when smoothing should be used than BM&W seem to believe, we view this as a relatively minor technical question, and we relegate a more detailed discussion to the appendix. Fortunately, resolving the debate is unnecessary, as there are many other ways to evaluate predictive performance. Indeed, in the case of heavily class-imbalanced civil war data, the AUC-ROC might not always be the most appropriate performance metric to use (Saito and Rehmsmeier 2015)—an argument that members of BM&W’s own team have made in the past. 6 In B&S we reported several additional metrics, including maximum F1 and Brier scores. We computed these metrics for the base specifications of our five models. Here we replicate and extend our analysis, adding these and other metrics for all specifications of all models, including eight robustness specifications.
Following B&S, in the base specification we implement random forest models with five terminal nodes and 100 observations per tree (sampled without replacement) and 100,000 trees per forest. We train each model on a subset of data beginning January 1, 2001, and ending December 31, 2007, then test the models on the remaining data, forecasting over intervals of either one or six months. We also consider models with 10 (rather than five) terminal nodes, 500 (rather than 100) observations per tree, and 1,000,000 (rather than 100,000) trees per forest. We also consider three alternate start dates for the test set (January 1, 2009, 2010, or 2011) and two alternate coding rules for the dependent variable, described below. In the online supplementary materials we show that our conclusions remain substantively unchanged when we use the default parameters and hyperparameters for each model (using the randomForest package in R), and when we probe robustness across a variety of additional parameter and hyperparameter choices.
We report predictive performance over one-month and six-month forecasting windows in Figures 1 and 2, respectively. The shading indicates which model performs best relative to the others; darker shading indicates better performance. 7 We focus on relative rather than absolute performance, since that is the main point of contention with BM&W. 8 We report nine performance metrics: the smoothed AUC-ROC, the empirical AUC-ROC, the area under the precision-recall curve (AUC-PR), 9 the maximum F1 score, 10 the mean F1 score, the maximum F2 score, 11 the mean F2 score, the Brier score, 12 and the MCC-F1 score. 13 Most of these performance metrics have been used by conflict forecasters in recent research, including by members of BM&W’s team. 14 The MCC-F1 combines the F1 score with another widely used performance metric, the Matthews Correlation Coefficient (Cao, Chicco, and Hoffman 2020).

Out-of-sample performance of one-month models. The top row in each panel reports results for the base specification of each model. We also report results with ten rather than five terminal nodes (second row); 500 rather than 100 observations per tree (third row); 1,000,000 rather than 100,000 trees per forest (fourth row); a test set that begins January 1, 2009, January 1, 2010, or January 1, 2011 (fifth, sixth, and seventh rows, respectively); and alternate codings of the dependent variable (eighth and ninth rows).

Out-of-sample performance of six-month models. The top row in each panel reports results for the base specification of each model. We also report results with ten rather than five terminal nodes (second row); 500 rather than 100 observations per tree (third row); 1,000,000 rather than 100,000 trees per forest (fourth row); a test set that begins January 1, 2009, January 1, 2010, or January 1, 2011 (fifth, sixth, and seventh rows, respectively); and alternate codings of the dependent variable (eighth and ninth rows).
BM&W draw three conclusions from their reanalysis of our data. They claim (1) that the theoretically-grounded escalation model under-performs the more atheoretical CAMEO model “in all instances;” (2) that the escalation model under-performs the average model “in all instances” as well; and (3) that adding structural characteristics “substantially improves” the performance of procedural models—in particular, that a hybrid with PITF model “strictly dominates” the escalation model alone (p. 19). Our results in Figures 1 and 2 show that BM&W’s first and second conclusions are false. (We evaluate their third conclusion in the next section.) While the differences between models are generally small, 15 they are broadly consistent with our results in B&S. In B&S we found that the escalation model generally outperforms the alternatives over one-month windows. Our results in Figure 1 are consistent with this finding. In B&S we also found that the Goldstein model is a “close competitor” to the escalation model over six-month windows (Blair and Sambanis 2020, 1899). Our results in Figure 2 are consistent with this finding as well. Indeed, the only performance metric according to which the CAMEO and average models consistently outperform the escalation model is the empirical AUC-ROC. BM&W’s first and second conclusions depend entirely on this subjective modeling choice.
Of the nine robustness specifications we test, the escalation model performs most inconsistently on “coding of DV 2.” As we noted in our original article, with such a rare outcome, small changes to coding rules could produce large changes in relative predictive performance. This problem is by no means unique to our study: conflict prediction is made inherently more difficult by the fact that civil wars are rare, and that it is not always straightforward to distinguish them from adjacent forms of political violence, or to code precisely when a civil war starts and ends. This observation motivated two robustness checks in B&S. In “coding of DV 1,” we recoded four especially ambiguous civil wars; 16 in “coding of DV 2,” we artificially lengthened all civil wars by coding the month of onset as January and the month of termination as December. The escalation model sometimes outperforms and sometimes under-performs the alternatives in this latter specification, depending on the performance metric we use. Nonetheless, taken together, our results in Figures 1 and 2 are broadly consistent with B&S, and inconsistent with BM&W.
Does Structure Improve on Process?
In B&S we presented a “procedural” model of conflict based on insights from the literature on social movements and contentious politics. Rather than focusing on slow-moving or time-invariant “structural” characteristics of countries to forecast civil war, our procedural model uses predictors that attempt to capture the dynamics of escalation and de-escalation that characterize interactions between opposition groups, rebels, and the state. The escalation model includes proxies for the demands that opposition and rebel groups may make of the state; the accommodations or acts of non-violent repression with which the state may respond to those demands; and the small-scale, low-level violence that may precede civil war. We sought to test whether such a procedural model of escalation can predict the onset of civil war using machine learning techniques that are “well suited to estimate contingent relationships without pre-specifying all potentially relevant interactions” (Blair and Sambanis 2020, 1890).
In focusing on process, however, we certainly did not mean to suggest that there is nothing to learn from models that emphasize structural differences between countries. Previous research by members of the Political Instability Task Force (PITF) has identified a number of country-level correlates of conflict, and has shown that these can be used to predict civil war relatively well, at least at the annual level (Goldstone et al. 2010). We decided to explore whether the addition of structural variables might improve the performance of the escalation model, which otherwise consists of procedural variables alone. Our prior was that the addition of structural variables would improve predictive performance, especially over longer temporal windows.
We tested four approaches. First, we simply added the variables from the PITF model to the escalation model (this model is called with PITF). Second, we re-weighted the predicted probabilities from the escalation model using forecasts gleaned from the PITF’s annual “Watch List” (this is the weighted model). Third, we used the PITF’s Watch List to run the escalation model separately for countries at high and low risk of conflict, then we recombined the high-risk and low-risk models to generate a single set of predictions. Our intuition was that the process of conflict escalation (and thus the predictors of civil war) might be different in high-risk countries than in low-risk ones. (We called this the PITF split population model.) Finally, we ran a model composed of PITF variables alone (the PITF model).
BM&W criticize several aspects of this analysis. These critiques are all relatively narrow, and none has anything to do with the question of whether or not theory is useful for forecasting. We respond to them briefly here, and leave a more detailed discussion to sections S.1 and S.2 of the online supplementary materials. BM&W make three points: (1) for the weighted model, we accidentally weighted the test set predicted probabilities from the escalation model by the training set forecasts from the PITF Watch List, due to a coding error; (2) the PITF split population model is unorthodox, and is really just a variation on the escalation model with a larger forest; and (3) our comparison of the procedural model to the structural and hybrid alternatives involves test sets of different sizes. (The N for the escalation model is slightly larger due to missingness in the PITF dataset.) The first critique relates to an easily corrected coding error; as we will see, fixing it significantly improves the AUC-ROC of the weighted model, but does not change our original conclusions in any substantive way. The second critique is contradicted by BM&W’s own results, as we discuss in section S.2 of the online supplementary materials. But the PITF split population model played only a tangential role in our original article, and dropping it does not change our conclusions in a substantive way either.
The third critique is a bit more complex. As BM&W note, the disadvantage of our approach is that differences in relative predictive performance may be artifacts of differences in the number of observations in each test set—though, as we show in section S.1 of the online supplementary materials, this is not actually a problem in our case, since relative performance remains roughly the same one way or the other. In other words, BM&W’s preference for standardizing the size of the test sets does not change any of our substantive conclusions. A disadvantage of BM&W’s approach is that it discards potentially useful data—a non-trivial issue when the outcome we seek to predict is so rare—and penalizes one model for missing data problems in the others. In any event, it is not obvious whether a larger test set helps or hurts the performance of any particular model. Given the rarity of the predictand, the performance of a model tested with more data could easily be worse than the performance of a model tested with less.
Nonetheless, for the sake of argument, in Figures 3 and 4 we grant all three of BM&W’s critiques and extend the analysis in B&S to test the robustness of our results across the same nine performance metrics and the same nine robustness specifications as above. Following BM&W, we (1) correct the minor coding error in the weighted model, (2) drop the PITF split population model, and (3) standardize the size of the test sets used to compare the procedural model to the structural and hybrid alternatives. We make two observations. First, even if we had addressed all three of BM&W’s critiques before publishing our original article, doing so would not have changed any of our substantive conclusions, unless we had also switched from smoothed to empirical ROC curves. 17 Here, as elsewhere in BM&W, the crux of their argument turns out to hinge entirely on the minor technical question of smoothing.

Out-of-sample performance of one-month models with PITF and standardized test sets. The top row in each panel reports results for the base specification of each model. We also report results with ten rather than five terminal nodes (second row); 500 rather than 100 observations per tree (third row); 1,000,000 rather than 100,000 trees per forest (fourth row); a test set that begins January 1, 2009, January 1, 2010, or January 1, 2011 (fifth, sixth, and seventh rows, respectively); and alternate codings of the dependent variable (eighth and ninth rows).

Out-of-sample performance of 6-month models with PITF. The top row in each panel reports results for the base specification of each model. We also report results with 10 rather than five terminal nodes (second row); 500 rather than 100 observations per tree (third row); 1,000,000 rather than 100,000 trees per forest (fourth row); a test set that begins January 1, 2009, January 1, 2010, or January 1, 2011 (fifth, sixth, and seventh rows, respectively); and alternate codings of the dependent variable (eighth and ninth rows).
Second, incorporating structural characteristics has ambiguous effects on the performance of our procedural model. Contrary to BM&W’s claim, the with PITF model does not “strictly dominate” the escalation model (p. 13); this conclusion again depends entirely on the use of empirical ROC curves. Comparing across the remaining metrics, we find that structural variables sometimes improve and sometimes diminish the performance of a model based on procedural variables alone. This is consistent with our results in B&S (p. 1904), and inconsistent with BM&W.
That said, the additional performance metrics and robustness specifications in Figures 3 and 4 do leave us somewhat more optimistic about the prospects for hybrid models of this sort. While it is not obvious which approach is best, there appears to be some promise in combining process with structure. This is intuitive, since structural variables can help predict where civil wars are likely to occur cross-nationally, even if they are less helpful in predicting when civil wars are likely to occur sub-annually. Again, our prior when we embarked on this exercise was that adding PITF predictors would improve the performance of the escalation model. To the extent that our extensions of B&S confirm our prior, this only strengthens our finding that theory can aid in prediction, since the variables in the PITF model were explicitly “drawn from the theoretical literature” (Goldstone et al. 2010, 194).
How Well Do Our “True” Forecasts Perform?
In B&S we generated “true” prospective forecasts for the first half of 2016 using the escalation model and data through the second half of 2015. We preregistered a list of the 30 countries with the highest risk of civil war, then returned to evaluate the accuracy of our predictions with the benefit of hindsight. Since there is inevitably some ambiguity in the way civil wars are coded, we considered two scenaria: (1) a “persistence” scenario in which we assumed that all countries that were in conflict at the end of 2015 continued to be in conflict at the beginning of 2016, and all countries that were at peace continued to be at peace; and (2) a “change” scenario in which we coded the civil war in Colombia as ending in the first half of 2016, and two new civil wars as starting in Turkey and Burundi. We classified our predictions as successes (“true positives”) any time a country with a new or ongoing civil war in the first half of 2016 appeared on our top thirty list. The procedure we followed was clearly explained in our article, and should not have created confusion.
BM&W nonetheless raise two objections to this exercise. First, they claim that we generated forecasts for the second half of 2016, rather than the first half, and tested them on data from the first half (p. 21). This claim is false. By simply reviewing the replication materials for our original article, 18 it is clear that our prospective forecasts were for the first half of 2016. Second, BM&W argue that we should have classified our predictions as true positives only when countries with new civil wars appeared on our top thirty list. They claim that we should not count it as a “win” for our model if it assigns a high probability of conflict in the first half of 2016 to countries with ongoing conflicts in the second half of 2015. But this argument is purely subjective, and there are good reasons to care whether our model assigns a high probability of conflict to countries with ongoing conflicts at the end of 2015. We already discussed some of these reasons in B&S. We reiterate and elaborate on them in the appendix.
Nonetheless, for the sake of argument, here we again grant all of BM&W’s critiques and extend our analysis in B&S to explore the robustness of our results. Whereas in B&S we conducted this prospective forecasting exercise using the escalation model alone, we now generate forecasts using all five models in order to compare the more theoretical to the less theoretical ones. In addition, in the original article we created confusion matrices using the threshold implied by the top thirty list in our PAP. As a robustness check, we see what happens when we set the threshold higher (top ten or top twenty) or lower (top forty). 19 Per BM&W’s preference, we code a true positive only when a model correctly predicts a new onset of civil war.
Following B&S, we begin by assuming that all countries that were at peace in the second half of 2015 were still at peace in the first half of 2016, and that all countries that were at war in the second half of 2015 were still at war in the first half of 2016. This is the “assuming persistence” scenario in B&S. Unfortunately, when we adopt BM&W’s preferred approach, this scenario becomes mostly uninformative: since we assume persistence, there are by definition no new onsets, and thus no true positives or false negatives. The only informative performance metric in this case is the number of false positives. Other potentially useful performance metrics either do not exist (e.g. recall), or are mechanically 0 (e.g. precision). We relegate presentation of these results to section S.3 of the online supplementary materials. (To the extent that they are informative at all, these results suggest that the escalation model performs as well as or better than the alternatives.)
Again following B&S, we then code new civil wars as starting in Burundi and Turkey and the ongoing civil war in Colombia as ending in the first half of 2016. This is the “assuming change” scenario. Again, per BM&W’s preference, we code a true positive only when a model correctly predicts a new onset. This exercise is more informative, and allows us to compute a broader range of performance metrics. In Figure 5 we report accuracy, precision, recall, specificity, false positive rate (FPR), false negative rate (FNR), F1, F2, and the Matthews Correlation Coefficient (MCC). We find that the escalation model outperforms all of the alternatives when we predict onsets in relatively few countries (ten or twenty), outperforms the Goldstein and CAMEO models when we predict onsets in more countries (thirty), and under-performs the CAMEO model when we predict onsets in more than thirty countries—though the difference is quite small. 20 Given that our forecasts become less useful as the number of countries in which we predict an onset grows, we view these results as further evidence of the escalation model’s value.

Performance metrics for out-of-sample test, onset assuming change. For this figure we assume that the civil war in Colombia ended by the first half of 2016, and that new civil wars began in Turkey and Burundi. The top left panel reports results when we predict civil war onsets in the top ten highest ranked countries. The top right, bottom left, and bottom right panels report results when we predict civil war onsets in the top twenty, thirty, and forty highest ranked countries, respectively.
What Is Theory?
Although BM&W purport to show that a theoretically-informed model of conflict escalation under-performs more inductive, mechanical alternatives, they nonetheless end up concluding that theory is valuable for prediction after all. For this substantive conclusion to be internally consistent with their empirical analysis, they must necessarily argue that our escalation model is not in fact informed by theory. This is precisely the argument they attempt to make in the final section of their paper, where they distinguish “strong” from “weak” theory and criticize our model as an example of the latter. We are as much confused by this posturing as we are by this team’s previous positions on the usefulness of theory for prediction. Members of BM&W’s team have argued both against theory and in favor of it, and in some cases have professed not to know what theory even is (Beger, Dorff, and Ward 2016; Ward et al. 2013; Ward 2017). In light of the ambiguous and contradictory positions this team has taken, it is distressing to note the extent to which they misrepresent our own.
First, BM&W mischaracterize our study as treating “theoretical modeling and machine learning as competing endeavors” (p. 18). This is plainly not true; indeed, the purpose of our study was to use machine learning methods to evaluate the predictive performance of a theoretically-informed model. In this sense, we offered an example of exactly the approach that BM&W recommend: we considered “how [theory and machine learning] can reinforce each other” (Beger, Morgan, and Ward 2021, 18). Because there is inevitably uncertainty about functional form and other features of model specification, machine learning methods offer ways to estimate models flexibly, and to explore complex and interactive relationships. We repeatedly emphasized these virtues of machine learning in our discussion of the methods underlying our models (Blair and Sambanis 2020, 1890, 1895). BM&W attack a straw-man version of an argument that we explicitly rejected.
BM&W also ignore the many caveats that were already present in our description of the data and models used in B&S. We noted that events in the ICEWS dataset “correspond closely but not perfectly to the concepts we wish to measure,” and that “one could experiment with different aggregations than the four we propose” (Blair and Sambanis 2020, 1893). We suggested that exhaustively comparing these permutations was potentially worthwhile, but was—and is—beyond the scope of our analysis. We also noted that, as a result of the structure of ICEWS, we could not link sequences of actions and reactions between actor dyads, and thus could not capture the “tit-for-tat dynamics that typically characterize conflict escalation” (p. 1908). In other words, we were quite transparent about the limitations of our approach. Here again, BM&W attack a straw-man version of an argument that we never made.
One could ask, as BM&W seem to do, whether there is any point in theorizing unless you can assemble the empirical model that captures everything important about the underlying theoretical dynamics you are studying. Our answer is “yes,” while theirs seems to be “no.” 21 BM&W’s position on this question strikes us as untenable. The fact that there can be other plausible models of conflict escalation does not necessarily invalidate this particular model. Our approach in B&S is fully consistent with prevailing practice in our discipline and with publications in top field journals, including the Journal of Conflict Resolution. It is also fully consistent with the approach that members of BM&W’s team have taken in their own prior research. Beger, Dorff, and Ward (2016, 102), for example, use theory as an “initial starting point” for building their conflict forecasting models; our models use theory in much the same way. Of course, there are many other ways to combine theory with machine learning and prediction. But most of them reflect an understanding of theory that resonates with our own.
Conclusion
We have shown that the thrust of BM&W’s critique is miscast on relatively minor technicalities and misses the bigger picture. We are of course aware that our analysis does not resolve the debate about the usefulness of theory for prediction. We do not offer a new theory of conflict escalation, and do not test all possible empirical permutations of all possible theories of conflict escalation. Indeed, we do not even test all possible empirical permutations of the one theory of conflict escalation that we propose. We made no claims to the contrary in B&S, and we recognized the limits and scope of our contributions to this debate. But we do believe we identified an important question, and that we pushed the literature in a productive direction that future scholars should continue to pursue. In the abstract of our original article, we wrote that we found “a more direct connection between theory and forecasting than is sometimes assumed,” but urged conflict forecasters to “treat the value-added of theory for prediction not as an assumption but rather as a hypothesis to test” (Blair and Sambanis 2020, 1885). In this paper, in the context of responding to criticism by BM&W, we reiterate this conclusion and encourage researchers to take up the challenge of empirically evaluating the usefulness of theory for prediction in the future.
Supplemental Material
Supplemental Material, sj-pdf-1-jcr-10.1177_00220027211026748 - Is Theory Useful for Conflict Prediction? A Response to Beger, Morgan, and Ward
Supplemental Material, sj-pdf-1-jcr-10.1177_00220027211026748 for Is Theory Useful for Conflict Prediction? A Response to Beger, Morgan, and Ward by Robert A. Blair and Nicholas Sambanis in Journal of Conflict Resolution
Supplemental Material
Supplemental Material, sj-zip-1-jcr-10.1177_00220027211026748 - Is Theory Useful for Conflict Prediction? A Response to Beger, Morgan, and Ward
Supplemental Material, sj-zip-1-jcr-10.1177_00220027211026748 for Is Theory Useful for Conflict Prediction? A Response to Beger, Morgan, and Ward by Robert A. Blair and Nicholas Sambanis in Journal of Conflict Resolution
Footnotes
Appendix
Acknowledgment
For helpful comments we thank Chad Hazlett, Nicholas Miller, and David Siroky. Marie Schenk provided excellent research assistance.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
Notes
Supplemental Material
Supplemental material for this article is available online.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
