Abstract
This paper revisits the latest statistical evidence for the nuclear emboldenment thesis—nuclear-armed states are more likely to initiate military aggression than non-nuclear states—from (Bell and Miller 2015). If correct, their findings have important theoretical and policy implications regarding the effect of nuclear proliferation on international conflict. This paper shows, however, that Bell and Miller’s findings heavily rely on two important components of their statistical analysis: (1) using all state dyad observations, and (2) employing pooled regression models to analyze time-series-cross-sectional (TSCS) data. I argue that those components are based on questionable assumptions on heterogeneity in their dataset. Based on alternative strategies dealing with heterogeneity in dyadic data, my reanalysis shows that the emboldening effect of nuclear weapons is not as robust as originally claimed. Instead, I find the robust deterrent effect of nuclear weapons: nuclear-armed states are less likely to be targeted in military disputes. These findings highlight the need for careful application of quantitative methods to produce a more robust understanding of nuclear issues.
Introduction
Do nuclear weapons encourage states to act more aggressively? Answers to this question are of great importance for scholars and policymakers. Knowing whether nuclear-armed states are more prone to use force and initiate interstate disputes is central to our understanding how nuclear weapons alter states’ foreign policy behavior. If nuclear-armed states can successfully convince their opponents that resistance would result in nuclear retaliation, they can significantly reduce the costs of using force to achieve political aims and therefore have greater incentives than non-nuclear states to solve conflicts of interest by force. U.S. policymakers have also echoed this logic: for instance, the past CIA Director Mike Pompeo warned of the dangerous consequences of Pyongyang’s nuclear acquisition, arguing that North Korea would “use these tool sets beyond self-preservation” (Z. Cohen 2018). In fact, the possibility that nuclear-armed states can successfully either threaten to use or employ military force to achieve their foreign policy aims behind the “nuclear shield” has been a core justification for U.S. counterproliferation policies: Given the danger posed by the aggressive actions of nuclear-armed states, Washington should not allow a new member to join the nuclear club (Gavin 2015, 23).
However, empirical evidence for this “nuclear emboldenment” thesis—the argument that nuclear-armed states are more likely to initiate military aggression than non-nuclear states—is mixed. On the one hand, several studies show qualitative evidence that nuclear-armed states demonstrate a greater tendency to engage in military aggression (Bell 2015, 2019, 2021; M. Cohen 2018; Kapur 2005, 2007). Other studies challenge this view, based on quantitative analyses, that nuclear-armed states do not initiate interstate conflict more frequently than non-nuclear states (Beardsley and Asal 2009; Gartzke and Jo 2009).
In their analysis of the relationship between nuclear weapons and interstate conflict, Mark S. Bell and Nicholas L. Miller (2015) find that nuclear-armed states are more likely than non-nuclear states to initiate militarized disputes against non-nuclear states, but under limited conditions. Specifically, they find that nuclear-armed states’ dispute initiation is largely driven by their effort to expand foreign policy interests against opponents with whom they previously did not have any history of conflict (Bell and Miller 2015, 84-85).
As one of the latest studies on nuclear weapons and military conflict, Bell and Miller’s findings carry significant weight. Bell and Miller use a better empirical strategy to consider the endogenous relationship between nuclear weapons and international conflict by accounting for drivers of nuclear proliferation (Gartzke and Kroenig 2016, 402). Therefore, their analysis provides more reliable quantitative evidence for the emboldening effect of nuclear weapons. They also present new evidence that nuclear emboldenment is mainly driven by nuclear-armed states’ expansion of foreign policy aims. This “expanded interests” thesis contrasts with the conventional wisdom that conventionally weak nuclear-armed states use nuclear weapons to compensate for their conventional military disadvantages: instead, both conventionally strong and weak states can be targets of nuclear-armed states’ conflict-triggering behavior. From a policy perspective, this implication could also help policymakers form informed expectations about conditions under which new nuclear-armed states demonstrate aggressive behavior against which targets and thus understand the nature of challenges posed by nuclear emboldenment. As such, assessing the validity of Bell and Miller’s empirical conclusions have important implications for scholarly research and policy debates on nuclear proliferation.
This paper argues that Bell and Miller’s statistical evidence for the nuclear emboldenment hypothesis heavily relies upon two crucial components of their analysis: using all state dyad observations and pooled regression models. I argue that these components are based on questionable assumptions about heterogeneity in dyadic data. I also demonstrate how each component affects Bell and Miller’s central results. First, Bell and Miller use all state dyad observations, which is the most inclusive sample selection criteria. This practice assumes that conditional on observed covariates, all pairs of states have a comparable chance of entering a conflict. I argue that this is a problematic assumption and demonstrate that considering the possibility that some dyads (e.g. irrelevant dyads) have practically zero chance of conflict occurrence results in notable changes in Bell and Miller’s core results. Second, in their analysis, Bell and Miller do not allow time-invariant unobserved dyadic differences that might be correlated to their key independent and dependent variables. As demonstrated by research on times-series-cross-sectional (TSCS) data, pooled regression models may ignore the significant effect of unmodeled dyad-specific heterogeneity, which in turn threatens our statistical inference. I show that after addressing these concerns about unit-level heterogeneity properly, the empirical link between nuclear weapons possession and an increased chance of dispute initiation is significantly weakened. Instead, I find a different and robust association between the bomb and international conflict: nuclear-armed states are less likely to be targeted in military disputes than non-nuclear states. I conclude that the latest statistical support for the nuclear emboldenment thesis is not as robust and generalizable as originally reported.
The rest of this paper proceeds as follows. First, I review scholarly arguments and existing findings on the relationship between nuclear weapons and conflict initiation. Second, I review scholarly attempts to address heterogeneity in dyadic data and the assumptions Bell and Miller implicitly make when designing their statistical analysis. Third, I revisit Bell and Miller’s key results by demonstrating how using strategies involving politically relevant dyads affects their empirical conclusion. Fourth, I show how using models that capture unobserved dyad-specific heterogeneity changes Bell and Miller’s statistical results. Fifth, I reexamine Bell and Miller’s evidence for the expanded interests thesis that nuclear-armed states are more likely to attack states with whom they have no previous disputes. I conclude with a brief discussion of the implications of my findings.
Do Nuclear Weapons Embolden Their Possessors? Theory and Evidence
The question of whether nuclear weapons acquisition encourages states to engage in military aggression more frequently has received continuous attention. One line of scholarship argues that nuclear weapons could lead their possessors to use force more frequently to achieve their aims. Among a range of mechanisms through which nuclear-armed states are emboldened, raising the costs of resistance by increasing the risk of nuclear escalation has been regarded as a crucial one. 1 For example, Bell (2019, 15) argues that nuclear weapons can increase the costs of resisting nuclear-armed states’ attempts to expand their foreign policy interests by raising the costs of escalation. Kapur (2005, 2007) makes a similar argument in the India-Pakistan dyad context that Pakistan employs the risk of nuclear escalation as a tool for deterring Indian conventional retaliation and drawing international mediation to achieve its territorial ambitions by force. Similarly, Kroenig (2014, 134) argues that a nuclear-armed Iran would be emboldened to launch low-level conflict more frequently because it can raise the risk of nuclear war and use it as a “cover” to implement its geopolitical aims.
Not all scholars, however, agree with the proposition that nuclear acquisition is a source of increased military aggression. While not directly engaging with the debate on nuclear emboldenment, several scholarly works argue that threats of nuclear retaliation are not always credible because using nuclear weapons entails significant strategic, political, and military costs. As a result, nuclear-armed states do understand that the benefits of using nuclear weapons do not always exceed its costs, and threats to use nuclear weapons are not always credible enough to change the calculations of adversaries. Therefore, nuclear-armed states do not necessarily behave more aggressively. For instance, Tannenwald (1999, 2007) argues that there exists a significant normative prohibition against the use of nuclear weapons, which suggests that threats of nuclear escalation are not always credible because fulfilling them could result in significant political backlash. Similarly, Todd S. Sechser and Matthew Fuhrmann (2017, 48-50) argue that using nuclear weapons could trigger diplomatic isolation, ruptures in alliance relationships, and further nuclear proliferation. Avey (2019, 18-19) also notes that not only does the use of nuclear weapons create significant collateral damage, which could complicate conventional military operations, but also it is likely to expand the level of violence in conflict, which may not be a desirable outcome even for nuclear-armed states. These arguments suggest that unless the interests at stake are high and alternative military options are not available so that the benefits of nuclear strikes exceed its substantial costs, threats to use nuclear weapons would lack credibility, and the danger of nuclear escalation would be seen as important only in a small number of high-stake military conflicts. In other words, nuclear-armed states do not necessarily engage in military aggression more frequently because they understand that the costs of implementing threats to use nuclear weapons are significant and the risk of nuclear escalation is not always exploitable.
Given these divergent claims, what evidence does existing scholarship find? To begin with, a significant portion of the vast literature on nuclear weapons examines how nuclear weapons influence states’ conflict behavior using qualitative case study methods. For example, Bell (2019, 10-13) argues that nuclear-armed states facing grave security threats, such as territorial threats or interstate wars, are especially more likely to engage in aggressive behavior. Cohen (2018) posits that leaders of nuclear-armed states are more likely to engage in coercive behavior, at least in the short term. Kapur (2005, 2007) also claims that Pakistan’s nuclear weapons, combined with its conventional military inferiority and territorial ambitions, make Islamabad more prone to engage in militarized attempts to draw territorial concessions from India.
On the other hand, existing quantitative studies on nuclear weapons found no supportive evidence for the nuclear emboldenment thesis. For example, Beardsley and Asal (2009, 245-47) show that nuclear weapons have little effect on the probability of crisis occurrence. Likewise, Gartzke and Jo (2009, 218-21) argue that while nuclear acquisition appears to increase the probability of the initiation of military conflict, this association no longer holds after addressing the endogenous relationship between nuclear weapons possession and conflict.
In their comprehensive analyses on the relationship between nuclear weapons and conflict occurrence, Bell and Miller provide an updated and nuanced empirical assessment. Using Gartzke and Jo’s (2009) dataset, they provide two results regarding the nuclear emboldenment thesis. First, nuclear-armed states are more likely than non-nuclear states to initiate military disputes against non-nuclear states. Second, this effect, however, is strongly conditional: nuclear-armed states are more likely to initiate aggression only against non-nuclear states with whom they have no history of conflict (Bell and Miller 2015, 84-86).
As the widely cited quantitative evidence revealing a nuanced relationship between nuclear weapons and conflict, Bell and Miller’s findings should be appraised as making significant progress in our empirical knowledge of nuclear weapons. Furthermore, their study uses a more suitable technique to address the concern that nuclear weapons possession is potentially endogenous to international conflict as military conflict is believed to be a powerful driver of a decision to acquire nuclear weapons (Gartzke and Kroenig 2016, 402). 2 Therefore, our confidence in the validity of Bell and Miller’s results is stronger than in previous studies. Lastly, Bell and Miller offer important supplementary evidence to existing qualitative research’s support for the nuclear emboldenment hypothesis. Nuclear-driven military aggression is not just applicable to a few nuclear-armed states: it is a generalizable pattern that holds across different dyads and temporal periods.
Heterogeneity in Dyadic Data
I argue that Bell and Miller’s empirical results are open to question because they are critically dependent on two decisions: (1) using all state dyad observations and (2) pooled logit models. I subsequently demonstrate why both decisions are potentially problematic given unobserved heterogeneity in dyadic data and how each decision influences Bell and Miller’s key results. My analysis indicates that the emboldening effect of nuclear weapons is even more tentative and weaker than Bell and Miller claim: there is no robust statistical support for the nuclear emboldenment thesis, even for its qualified version.
Heterogeneity in Dyadic Data and Strategies for Conflict Researchers
I begin by discussing why dealing with heterogeneity is an important issue for analyses of TSCS data. Debates on the effect of heterogeneity on statistical inferences are not new for quantitative approaches to studying international politics, and heterogeneity across dyads is a critical issue for dyadic analyses of international conflict. It is now well known that ignoring heterogeneity in dyad data creates a serious challenge to our statistical inferences because it violates an important assumption of exchangeability: observations become comparable after taking into account our explanatory variables (King 2001, 498). This assumption is violated if one can make a compelling argument that there are significant unit-level differences between different dyads that are not captured by our observed explanatory variables. The consequences of the violation are severe, often resulting in substantial biases in our statistical inferences (Green, Kim, and Yoon 2001; King 2001). Therefore, researchers need to make proper decisions regarding how to address unit-level heterogeneity in their dyad data.
One common strategy of dealing with heterogeneity in dyadic data is to differentiate a group of dyads where states have a reasonable chance of experiencing conflict from another group of dyads where there is practically no chance of conflict by using certain criteria based on observable variables (Beck 2001, 290; Box-Steffensmeier, Reiter, and Zorn 2003, 280; Neumayer and Plümper 2020). The underlying idea is that some dyads are extremely different from other dyads regarding their baseline chance of experiencing conflict in that the former group’s baseline rate is practically zero, while the latter group has a reasonable chance of conflict occurrence. Ignoring this unobserved heterogeneity in conflict propensity could be a significant threat to our inferences. Identifying dyads where there is no realistic chance of conflict is also important because our sample needs to include only proper “negative” cases—cases where the outcome of interests (e.g. conflict initiation) does not occur even if it could (Clark and Nordstrom 2003; Mahoney and Goertz 2004, 656-57). The argument that some cases are “proper” negative cases means that other cases are “inappropriate” negative cases, in which the outcome of interest does not occur because it cannot—two countries in a pair might have no realistic chance of conflict.
To delineate a proper scope of dyads having a reasonable chance of conflict occurrence, scholars often employ what is called politically relevant dyads. Politically relevant dyads consist of state dyads in which at least one state is a major power, or two states are geographically contiguous. 3 Previous studies argue that geography and major power status are powerful determinants of the possibility of interstate military conflict. 4 By using “relevant” dyads only, a sample can include only dyads where a reasonable chance of conflict exists (Maoz and Russett 1993; Oneal and Russett 1997; Weede 1976). The risk of including “irrelevant” cases in a sample, on the other hand, is to artificially increase the number of cases and result in erroneous statistical inference (Clark and Nordstrom 2003; Mahoney and Goertz 2004, 656-57).
While this “domain restriction” strategy has been widely used by studies on conflict initiation and some past studies on nuclear weapons also employed it (Fuhrmann and Sechser 2014; Narang and Mehta 2019), it is not free from criticism. Scholars argue that (1) conflict does occur even within non-relevant dyads and (2) the domain restriction strategy may induce selection bias (Braumoeller and Carson 2011; Lemke and Reed 2001; Mahoney and Goertz 2004, 662-63). Recent studies also propose several techniques to explicitly model differences in the baseline chance of conflict occurrence between politically relevant and non-relevant dyads, rather than dropping non-relevant dyads (Braumeoller and Carson 2011; Xiang 2010).
Another strategy for addressing heterogeneity in dyadic data is to use regression models that directly consider unobserved dyad-specific heterogeneity, such as fixed effects models. Rather than assuming that two groups of dyads have two different opportunities for conflict, this approach directly models unobserved heterogeneity across dyads. Fixed effects models are widely used but appear to be more contested than the politically relevant dyads approach. Some experts suggest that fixed effects logistic regression models offer a simple but powerful solution (Green, Kim, and Yoon 2001), but others vehemently disagree, claiming that if the event of interest does not frequently occur, as in the case of military conflict, this solution reduces the number of available observations in a sample dramatically, rendering any inference based on such a reduced sample dubious (Beck and Katz 2001). 5 In addition, recent studies propose an alternative model that can account for unit-level heterogeneity without discarding observations that show no variation in the dependent variable (e.g. Bell and Jones 2015).
The use (or non-use) of these strategies indicates different assumptions researchers make (often implicitly) about comparability after taking into account observed explanatory variables. For instance, the use of strategies based on politically relevant dyads suggests that a researcher assumes that after identifying and addressing a group of dyads (e.g. non-relevant dyads) where the expected opportunity for conflict is practically zero, the remaining observations (e.g. politically relevant dyads) are sufficiently homogenous. The use of models that capture dyad-specific heterogeneity but without considering politically relevant dyads indicates a different assumption that once unobserved heterogeneity is captured by the models, all dyads become sufficiently exchangeable. Thirdly, if a researcher uses the models capturing dyad-level heterogeneity and politically relevant dyads (e.g. fixed effects models with only politically relevant dyads), she assumes that only after identifying dyads that have a realistic chance of conflict and accounting for dyad-specific heterogeneity, observations become comparable. Lastly, the non-use of all of those strategies indicates yet another assumption, which is that all dyad observations are comparable.
When building on Gartzke and Jo’s (2009) analysis, Bell and Miller adopt the last approach: they use all state dyad observations and do not use models that consider dyad-specific heterogeneity. 6 This approach certainly has its own merits: for instance, it does not discard any observations. However, such a decision comes with the risk that even after controlling for observed covariates, there might be a substantial amount of unobserved dyad-specific heterogeneity that could significantly threaten the validity of their inference.
Using the Politically Relevant Dyads Approach Weakens the Quantitative Evidence for the Nuclear Emboldenment Thesis
I first investigate whether using empirical strategies based on politically relevant dyads could significantly influence Bell and Miller’s core findings. 7 If the findings are robust even after adopting those strategies, we can confirm the validity of Bell and Miller’s original findings. In fact, some studies on nuclear weapons and conflict adopt this strategy (e.g. Fuhrmann and Sechser 2014, 925).
I adopt several variants of the politically relevant dyads approach. For instance, I use the traditional politically relevant dyads approach and replicate Bell and Miller’s analysis without non-relevant dyad observations. I also use empirical strategies that explicitly model the effect of heterogeneity in terms of the baseline chance of conflict by using politically relevant dyads. These strategies depart from traditional approaches in that rather than assuming that conflict is possible only within politically relevant dyads, they directly model heterogeneity between two groups of dyads in terms of conflict propensity. Therefore, they are less vulnerable to selection bias and do not drop any observations. Here I intend to demonstrate whether different variants of the politically relevant dyads approach could produce similar results and whether those results are different from Bell and Miller’s original results.
Before turning to regression analyses, I first investigate whether politically relevant dyads are different from non-relevant dyads in terms of the risk of conflict initiation. It can demonstrate whether there are indeed substantial differences between politically relevant dyads and non-relevant dyads in terms of conflict propensity. If we can find heterogeneity between two types of dyads, then further analysis of how such heterogeneity influences our inference is warranted, regardless of specific methods of dealing with heterogeneous conflict propensity.
In Bell and Miller’s data consisting of all dyad observations, the total number of militarized interstate dispute (MID) initiations, Bell and Miller’s dependent variable, is 1,545 and the probability of MID initiation is 0.14%. Once we divide this sample into politically relevant dyads and non-relevant dyads, however, there is significant heterogeneity between the two sets of dyads in terms of conflict propensity. While the risk of MID initiation is 1.28% in the politically relevant dyads sample, it is only 0.03% in the non-relevant dyads sample. It demonstrates that politically relevant dyads and non-relevant dyads are indeed significantly different from each other in terms of conflict propensity and pooling these two types of observations may result in ignoring substantial heterogeneity between those observations.
A Cross Tabulation Analysis of Challenger’s Nuclear Status against MID Initiation, with Different set of Dyads. 8
A Cross Tabulation Analysis of Target’s Nuclear Status against MID Initiation, with Different set of Dyads.
Now I turn to multivariate regression analysis to see whether dropping non-relevant dyads creates noticeable changes in Bell and Miller’s original results. I first replicate Bell and Miller’s result that nuclear states are more likely than non-nuclear states to initiate conflict against non-nuclear states. 9 Their independent variables are two dichotomous variables indicating the nuclear weapons state status of both a challenger and a target in a dyad. The dependent variable is MID initiation. I then replicate Bell and Miller’s original analysis but only using politically relevant dyad observations. Dropping non-relevant cases reduces the number of observations in the sample from 1,079,328 to 99,638. There are 1,272 cases of MID initiations (82.33%) within politically relevant dyads, and 273 cases of MID initiation (17.67%) that occurred within non-relevant dyads.
Alternatively, I use an expanded operationalization of politically relevant dyads. In his thorough examination of different operationalizations of political relevance, D. Scott Bennett (2006, 260) proposes an expanded, but still reasonably restrictive operationalization of politically relevant dyads, which is based on the major power status and current or historical geographic contiguity through either direct territorial contact or colonial possessions. He argues that this could be a balanced operationalization of politically relevant dyads which can achieve a more reasonable trade-off between including more cases of conflict initiation and excluding irrelevant dyads. With this new operationalization, now 91.6% (1,415) of all MID initiation occurs within politically relevant dyads. By doing so, this approach substantially reduces the number of MID initiations that are not included in the sample (from 273 to 130) and therefore mitigates the missing data concern.
Figure 1 visualizes the estimated coefficients of our variables of interest, the challenger and the target’s nuclear weapons possession.
10
If a variable’s confidence interval bar crosses the vertical dashed line, which indicates zero, then this means that the confidence interval contains zero and the variable is not significantly associated with the dependent variable at the 5% level. The Effect of Nuclear Weapons Possession on MID Initiation (Politically Relevant Dyads). Point estimates and 95% confidence intervals. All other covariates are not included in the figure. A coefficient’s confidence interval containing zero means that it is not statistically significant.
I successfully replicate Bell and Miller’s original finding, indicating that nuclear-armed states are more likely than non-nuclear states to initiate MIDs against non-nuclear states (model 1). However, this association between nuclear weapons and MID initiation changes dramatically in models using politically relevant dyads. In models 2 and 3, the challenger’s nuclear status is no longer significantly associated with MID initiation, while the target’s nuclear status is significantly and negatively related to MID initiation. Moreover, the coefficients representing the challenger’s nuclear status reported in models 2 and 3 are statistically different at the 5% level from the corresponding coefficient in model 1. 11 In other words, the robustness of Bell and Miller’s findings is highly dependent on the inclusion of non-relevant dyads. 12
As noted earlier, dropping irrelevant dyads is not free from criticism. For example, it might induce selection bias because determinants of relevant dyads might be correlated with other covariates (Lemke and Reed 2001, 136). In addition, conflict does occur within non-relevant dyads (Braumoeller and Carson 2011, 293; Mahoney and Goertz 2004, 662-63). To address these concerns, I conduct additional analyses. First, I simply control for whether or not a dyad is politically relevant. The results remain identical: the challenger’s nuclear status is not a statistically significant predictor of MID initiation, while nuclear-armed targets are less likely to face MID-triggering action by non-nuclear states.
Second, I estimate a multiplicative interaction model in which the nuclear weapons variables interact with the political relevance variable. This is based on the idea that if we understand political relevance as a variable indicating either an opportunity for conflict or a willingness to engage in conflict, then our independent variables—whether a state is a nuclear-armed state—may interact with political relevance. For example, a nuclear-armed state is emboldened to initiate conflict against another state only if both states are in a politically relevant dyad. This approach employs all dyads and still accounts for the potential impact of political relevance, thereby allowing us to avoid a “sin of omission” (Braumoeller and Carson 2011, 293).
Interpreting a multiplicative interaction model benefits from calculating the substantive effect of variables of interest (Brambor, Clark, and Golder 2006). Therefore, I calculate the substantive effect of our nuclear variables. 13 When both states are in a non-relevant dyad, the nuclear status of the challenger increases the probability of MID initiation by 0.056 percentage point (pp) and the effect is statistically significant. If both states are in a politically relevant dyad, however, the effect of challenger’s nuclear status on the risk of MID initiation is statistically indistinguishable from zero. 14 On the other hand, when a dyad is not politically relevant, the effect of the target’s nuclear weapons possession on the likelihood of MID initiation is statistically indistinguishable from zero. In a politically relevant dyad, however, the target’s nuclear status significantly reduces the probability of being targeted in a MID by 0.051 pp. This pattern is consistent with the above results: within politically relevant dyads, nuclear-armed states are no more likely to initiate military conflict against non-nuclear states, but they are less likely to be targeted in disputes by non-nuclear states. 15
Next, I adopt the methods that explicitly model the effect of political relevance. Rather than assuming that conflict is possible only within politically relevant dyads, these methods directly take into account heterogeneity in the baseline conflict propensity between politically relevant dyads and non-relevant dyads. First, I estimate a split population binary choice model which treats political relevance as a latent variable, sorts out the dyads which lack the chance of conflict, and models the probability of conflict initiation simultaneously (Xiang 2010). In the MID initiation equation, I include all covariates used by Bell and Miller. The relevance equation includes contiguity, geographical distance, the presence of an alliance relationship within a dyad, and whether either one of the two states is a major power.
Second, I consider the role of the latent probability of conflict, captured by political relevance, as a ceiling effect that influences the effect of nuclear weapons on the risk of military conflict (Braumoeller and Carson 2011). Following Braumoeller and Carson (2011), I assume that the effect of nuclear weapons is significantly attenuated by democracies and (the lack of) political relevance and estimate a Boolean logit model. This model allows us to directly estimate how variation in the risk of military conflict between two groups of dyads influences the effect of nuclear weapons on conflict initiation.
Figure 2 shows the results. In model 4, the challenger’s nuclear possession is not significantly associated with the risk of MID initiation, while the target’s nuclear weapons significantly reduce the probability of MID initiation. In model 5, nuclear-armed states are now less likely to initiate MIDs than non-nuclear states; and they are less likely to be targeted in MIDs by non-nuclear states. These results show that even if we explicitly model the effect of political relevance on the relationship between nuclear weapons and conflict, nuclear weapons do not encourage states to act more aggressively than they would otherwise. Instead, my results confirm the new pattern that nuclear weapons act as a powerful deterrent. The Effect of Nuclear Weapons Possession on MID Initiation (Politically Relevant Dyads: More Analyses). Point estimates and 95% confidence intervals. All other covariates are not included in the figure. A coefficient’s confidence interval containing zero means that it is not statistically significant.
The above analyses show that once we take seriously the possibility that non-relevant dyads, in general, do not have a realistic chance of conflict, the supportive evidence for the nuclear emboldenment thesis no longer holds. In all my analyses, I find little evidence for the argument that nuclear-armed states are more likely than non-nuclear states to initiate militarized disputes against non-nuclear states. Instead, I consistently find that nuclear weapons significantly reduce the probability of being challenged in MIDs. At least, therefore, it is reasonable to say that the statistical support for the emboldening effect of nuclear weapons is not as robust as originally believed.
Using Models Capturing Dyad-Specific Heterogeneity Also Weakens the Statistical Evidence for the Nuclear Emboldenment Thesis
Bell and Miller’s second consequential decision is using pooled logit models. Bell and Miller use TSCS data and estimate ordinary logit models with a correction for potential bias from rare events. By doing so, they assume that there are no meaningful unobserved differences between different dyads. This is in turn equal to assuming that each dyad has the same baseline of conflict initiation, conditional on observed covariates. If, however, there are unmodeled dyad-specific differences that are associated with the initiation of militarized conflict and they are correlated with covariates in the models, any inference drawn from ordinary models may be subject to serious bias. 16 In other words, pooling TSCS observations without considering unit-specific heterogeneity may result in a biased inference (Green, Kim, and Yoon 2001).
As a solution, scholars have long noted the utility of fixed effects models, but there has been no consensus regarding their costs and benefits. The advocates of fixed effects models argue that they offer a simple and methodologically sound solution (Green, Kim, and Yoon 2001), while others disagree, claiming that the application of fixed effects models to the study of military conflict has significant costs, such as a dramatic loss of observations (Beck and Katz 2001).
Another solution that has been recently proposed is a within-between random effect (WB-RE) model (Bell and Jones 2015). Without discarding any observations that do not experience conflict, this model captures heterogeneity across dyads in a random effects model framework and estimates the between-unit effect and within-unit effect simultaneously by disaggregating the independent variables into their unit-specific means and the deviation from the unit-specific means (Bell, Fairbrother, and Jones 2019, 1055; Bell and Jones 2015). 17
Does unobserved unit-level heterogeneity that is correlated to the explanatory variables exist in Bell and Miller’s dataset? To answer the question, I conduct a test proposed by Mundlak (1978), which examines whether there is an association between time-invariant unit-level heterogeneity and the observed independent variables.
18
The results imply that there is a significant correlation between the unit-specific time-invariant characteristics and the independent variables (
Accounting for Time-Invariant Unobserved Dyadic Heterogeneity
Here I examine whether accounting for time-invariant unobserved dyadic heterogeneity alters Bell and Miller’s key results. The goal is to examine the robustness of Bell and Miller’s central findings from pooled binary response models. If their findings remain substantively identical, then our confidence in the validity of Bell and Miller’s results becomes higher. If accounting for unmodeled dyadic heterogeneity alters their key results, however, some healthy skepticism may be warranted because it signals that the validity of the findings is highly dependent on the assumption about the costs and benefits of ignoring unmodeled unit-level heterogeneity.
Estimating fixed effects logit models is one way of accounting for time-invariant dyadic heterogeneity. Given scholarly disagreement over the costs and benefits of fixed effects logit models, however, one might argue that any meaningful difference between the results from fixed effects logit models and those from ordinary logit models might not be enough for questioning the robustness of Bell and Miller’s findings. That difference, the argument goes, might occur as a result of dropping observations that never experience conflict. In fact, when estimating a fixed effects logit model, roughly 97.6% of the observations used by the pooled model (1,027,739) is dropped from Bell and Miller’s dataset as a vast majority of dyads never experience militarized disputes.
As an alternative approach, therefore, I estimate a WB-RE model (Bell and Jones 2015). In the model, the within-unit effect of nuclear weapons measures the effect of change in a state’s nuclear status in a given dyad-year (e.g. the state acquires nuclear weapons in that dyad-year) on its probability of conflict initiation, while my analyses between-unit effect measures the effect of different states’ average nuclear year (the average years of a state’s nuclear possession) in different dyads on the probability of conflict initiation. I focus on the within-unit effect of the challenger’s and target’s nuclear weapons possession, as the questions often asked by scholars or policymakers take the form of “would North Korea behave more aggressively if it acquires nuclear weapons?”—which potentially implies the within-unit effect of nuclear weapons. It should be noted, however, that I am not arguing that the between-unit effect of nuclear weapons is theoretically irrelevant. Theories of nuclear weapons rarely make a distinction between two effects, and answers to both the question of “whether nuclear weapons make a state more aggressive?” and “whether nuclear-armed states are more aggressive than non-nuclear states?” have important implications.
Figure 3 displays that the results from the WB-RE model are different from Bell and Miller’s results. Model 6 indicates that the within-unit effect of the challenger’s nuclear status is not statistically significant, meaning that a state’s acquisition of nuclear weapons is not significantly related to the likelihood of MID initiation.
19
On the other hand, the within-unit effect of the target’s nuclear possession is both statistically significant and negative, implying that being a nuclear-armed state significantly reduces the risk of being targeted in MIDs. The Effect of Nuclear Weapons Possession on MID Initiation (Accounting for Dyad-Specific Heterogeneity). Point estimates and 95% confidence intervals. All other covariates are not included in the figure. A coefficient’s confidence interval containing zero means that it is not statistically significant.
While not reported in Figure 3, I also find that the between-unit effect of the challenger’s nuclear status on MID initiation is both positive and statistically significant, which is consistent with Bell and Miller’s original findings. On the other hand, the between-unit coefficient of the target’s nuclear status variable is statistically indistinguishable from zero. This suggests that there is a possibility that nuclear-armed states, compared to non-nuclear states, are more likely to initiate MIDs against non-nuclear states, but nuclear weapons do not make their possessors act more aggressively than they would without nuclear weapons. If found consistently in other models, this distinction between the between-unit effect and the within-unit effect of nuclear weapons might deserve further theoretical and empirical scrutiny. 20
To summarize, my analyses demonstrate that Bell and Miller’s statistical support for the nuclear emboldenment hypothesis is not as strong as originally argued. The acquisition of nuclear weapons does not make states more likely to initiate military conflict against non-nuclear states, while it makes states less likely to be targets of military conflict. This is consistent with the pattern that has been repeatedly observed in the models discussed above. It indicates that accounting for unmodeled dyad-specific heterogeneity could significantly influence our inference of how nuclear weapons change conflict behavior.
The Expanded Interests Thesis Reconsidered
The previous sections show that the empirical association between nuclear weapons and an increased probability of conflict initiation is significantly weakened once we adopt empirical strategies dealing with heterogeneity in dyadic data. In this section, I also review the robustness of Bell and Miller’s another central finding: the emboldening effect of nuclear weapons is mainly driven by nuclear-armed states’ expansion of foreign policy aims and their pursuit of those aims with military force against new adversaries (Bell and Miller 2015, 84-86).
To examine whether employing strategies using politically relevant dyads affect the original finding, I rerun Bell and Miller’s model in which they interact the dispute history variable with the challenger’s nuclear status variable by applying the methods I used in the preceded analyses. To ease comparison between different results, I calculate the substantive effect of the challenger’s nuclear possession on the dependent variable when the dispute history variable’s value is zero, which indicates the absence of any previous conflict involvement between the challenger and the target. If Bell and Miller’s results are robust, then the challenger’s nuclear status should have a positive and significant effect on MID initiation if there is no previous conflict between the challenger and the target.
The Effect of the Challenger’s Nuclear Status on MID Initiation against the Target with Whom There Is No Previous MID (Accounting for Political Relevance).
The Effect of the Challenger’s Nuclear Status on MID Initiation against the Target with Whom There Is No Previous MID (Accounting for Dyad-Specific Heterogeneity).
Taken together, these results demonstrate that the validity of the evidence for the expanded interest thesis is more tenuous than previously claimed. Once we account for heterogeneity in dyadic data, the evidence that nuclear-armed states pursue broader foreign policy objectives with military force against new non-nuclear opponents is not strongly robust.
Conclusion
This paper argues that Bell and Miller’s conclusion that nuclear-armed states are more likely than non-nuclear states to initiate conflicts against non-nuclear opponents highly depends on their assumptions about heterogeneity in dyadic data. By using all state dyad observations and ordinary logit models with rare events bias correction, they implicitly assume that all dyads have a comparable and practically non-zero baseline rate of conflict. I argue that this assumption is potentially problematic and has crucial impacts on Bell and Miller’s findings. My reanalysis and extension of Bell and Miller’s two core results show that when relaxing this assumption and using alternative model specifications, the evidence that nuclear weapons encourage states to expand their foreign policy goals and initiate militarized disputes is significantly weakened.
Strikingly, the empirical association that is most consistently observed from my analysis is the deterrent, not emboldening, effect of nuclear weapons: nuclear-armed states face a reduced probability of being targeted in MIDs. Additional analyses using models that capture dyad-level heterogeneity based on only politically relevant dyad observations lend further support. Figure 4 shows the results from WB-RE logit models and Bell and Miller’s original model. The WB-RE models exclusively use either politically relevant dyads or an expanded set of politically relevant dyads (Bennett 2006). In all models, the within-effect of the target’s nuclear possession on the probability of MID initiation is still negative and statistically significant.
23
These findings show that even under a more restrictive assumption about the comparability of dyad observations, the negative effect of the target’s nuclear possession on MID initiation remains robust.
24
The Effect of Nuclear Weapons Possession on MID Initiation (Accounting for Dyad-Specific Heterogeneity and Using Politically Relevant Dyads). Point estimates and 95% confidence intervals. All other covariates are not included in the figure. A coefficient’s confidence interval containing zero means that it is not statistically significant.
This paper contributes to debates on nuclear emboldenment by demonstrating that there is little statistical support for the argument that nuclear weapons make their possessors more aggressive; there is also no robust support for a qualified argument that nuclear-armed states pursue their expanded interests by finding new opponents, without pushing harder existing rivals. At least, therefore, it is reasonable to conclude that there is no strong link between nuclear weapons and an increased risk of conflict initiation, even in a conditional form.
If nuclear weapons do not facilitate military aggression, then why are nuclear-armed states less likely to be targets in MIDs? Gartzke and Jo (2009, 215) once note that the observed relationship between nuclear weapons possession and conflict is likely to be weakened by opposing effects, such as the status-quo enhancing effect and the destabilizing nature of nuclear weapons. One could argue that given the evidence presented in this paper, the deterrent power of nuclear weapons may prevail under a wider range of circumstances than the emboldening effect of nuclear weapons. For instance, the costs of nuclear escalation may loom larger for an initiator of military conflict than a target. While a side using force first may not necessarily be the side responsible for the “root” causes of the conflict, the fact that it first initiates visible military actions could negatively affect its international reputation, which potentially amplifies the political and diplomatic backlash against nuclear escalation. If this is the case, then threats of nuclear escalation may not give an effective cover for states wanting to use force in case of bargaining failure. On the other hand, nuclear-armed targets may be less constrained by the negative consequences of nuclear use when their adversaries initiate the use of military force: they may be able to enjoy favorable international opinion and strong domestic support, which could reduce the costs of breaching normative prohibitions against nuclear use (Sechser and Fuhrman 2017, 59). Therefore, threats of nuclear escalation by a target in a military dispute could be perceived as more credible, even if other costs of nuclear escalation would still exist (e.g. collateral damage, the complication of conventional military operations). Such an expectation, in turn, makes opponents of nuclear-armed states reluctant to use force against them and more willing to settle conflicts of interest peacefully, given the higher chance of nuclear escalation in case of military attacks and its catastrophic consequences.
It is important to note that my findings do not necessarily preclude the possibility that nuclear weapons embolden their possessors in different ways. While my analyses strongly challenge the notion that nuclear-armed states show an increased propensity for military aggression, policy practitioners’ concerns over the emboldening power of nuclear weapons are rarely limited to more occurrences of military conflict. Other forms of assertive behavior are also of great importance, such as intra-conflict escalation, aggressive rhetorical threats, and increased military and political support for allies. 25 I suggest two possible ways in which future research could approach the issue of nuclear emboldenment. First, nuclear weapons could still embolden states, but its effect is strongly conditional on other facilitating conditions that are not captured by the datasets used in past studies, such as serious territorial threat (Bell 2019, 2021) or revisionist territorial ambition (Kapur 2007). Second, nuclear weapons may encourage states to engage in a range of behavior other than initiating military conflict, and a simple indicator of the occurrence of militarized conflict may not be suitable for capturing such behavior. Thus, new data on state preferences, external security environment, interstate coercion, and conflict behavior would be helpful for future quantitative research to explore how nuclear weapons change state behavior in different ways and add novel empirical evidence for debates on nuclear emboldenment.
My findings that nuclear weapons deter the initiation of international conflict lend partial support to nuclear optimists’ belief that nuclear weapons are a driver of peace because it dramatically reduces the chance of military conflict (e.g. Waltz 1981). The support is partial, however, because it does not answer the questions of whether the process of the spread of nuclear weapons is conflict-prone (e.g. Sobek, Foster, and Robison 2012) and the effect of nuclear proliferation on the risk of accidental nuclear use (Sagan 1994). Nor does it dispute that the consequences of war fought by nuclear weapons are much more devastating, even if its probability becomes remote (Kydd 2019). Furthermore, the dependent variable used in my analyses does not differentiate various forms of international conflicts with different levels of violence. As one of the important questions in the optimist-pessimist debate is about the impact of nuclear weapons on interstate war, the most violent form of conflict, my findings do not offer a direct answer to the question of the relationship between nuclear weapons and interstate war. Nor does it assess whether the predictions of “the stability-instability paradox” thesis is correct, an issue that has been subject to extensive empirical scrutiny (e.g. Bell and Miller 2015; Kapur 2007; Rauchhaus 2009). These questions remain important avenues for future investigation.
While this paper demonstrates that scholars should carefully consider the impacts of each element of their empirical research design to reach robust conclusions, it does not suggest the field should abandon quantitative research methods for studying the effect of nuclear weapons on world politics. On the one hand, my argument that assumptions about heterogeneity in dyadic data hugely affect our empirical conclusions echoes an emerging body of studies challenging the utility of quantitative approaches to understanding nuclear security issues. This literature argues that past quantitative studies often suffer from several problems, including poorly coded variables and the lack of sufficient robustness tests (Bell 2016; Montgomery and Sagan 2009), the failure to pay attention to heterogeneity across cases and the small-N problem (Gavin 2014), and the lack of careful consideration of the underlying assumptions of statistical models (Winter and Lenine 2020). On the other hand, it is the transparency of existing studies, one of the notable strengths of quantitative studies (Gartzke 2014; Fuhrmann, Kroenig, and Sechser 2014), that contributed to my reanalysis of Bell and Miller’s (2015) results by promoting reliable replication. My replication also benefits from existing quantitative scholars’ attempts to provide solutions to inferential challenges to quantitative studies of interstate conflict. As quantitative research methods still possess unique strengths, including greater generalizability, transparency, and ability to model probabilistic processes (Gartzke 2014; Fuhrmann, Kroenig, and Sechser 2014; Narang 2014), they could still produce meaningful progress in our knowledge of nuclear weapons and international security with thoughtful consideration of the core elements of research design.
Supplemental Material
Supplemental Material - Does the Bomb Really Embolden? Revisiting the Statistical Evidence for the Nuclear Emboldenment Thesis
Supplemental Material for Does the Bomb Really Embolden? Revisiting the Statistical Evidence for the Nuclear Emboldenment Thesis by Kyungwon Suh in Journal of Conflict Resolution
Supplemental Material
Supplemental Material - Does the Bomb Really Embolden? Revisiting the Statistical Evidence for the Nuclear Emboldenment Thesis
Supplemental Material for Does the Bomb Really Embolden? Revisiting the Statistical Evidence for the Nuclear Emboldenment Thesis by Kyungwon Suh in Journal of Conflict Resolution
Footnotes
Acknowledgements
I would like to thank Daniel McDowell, Ryan D. Griffiths, Simon Weschle, Mark S. Bell, and Jason Jung Jae Kwon for their helpful suggestions on earlier drafts of this paper. I also appreciate valuable feedback from participants at the Brown Bag Research Workshop in the Department of Political Science at Syracuse University, as well as three anonymous reviewers. All errors are mine alone.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
Supplemental Material
Supplemental material for this article is available online.
Notes
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
