Abstract
The Department of Defense (DOD) has documented a strategic gap how exploratory analyses are accomplished to support capability development which was decomposed into areas of needed focused research. We began with an exploration of current methods for integrating different models to meet the concerns of Congress with regard to quality, accuracy, and dependability, noting that they have become too computationally prohibitive for exploring large trade spaces. In addition, current model abstraction methods have difficulty accounting for the increasing dimensionality associated with increasingly complex simulations. These observations led to the formulation of the Reduced Order Non-INtrusive (RONIN) modeling methodology, which generates predictive reduced order surrogate models, which capture more information regarding behaviors as compared to traditional methods. The RONIN modeling methodology works to create surrogate models which emulate stochastic full-order models (FOMs) by leveraging order reduction approaches, stochastic modeling methods, and regression techniques. To demonstrate the RONIN modeling methodology, a notional United States Air Force use case was defined, and a DOD standard simulation framework was used to create relevant simulation scenario which output a set of response distributions. Ultimately, the RONIN modeling method was used to create a predictive surrogate model which was able to reconstruct output distributions which are statically consistent with the original FOM on average over 99% of the time while reducing the time needed to generate a distribution of outputs from minutes down to less than a second.
1. Introduction
As the Department of Defense (DOD) works to prepare for the threats called out in strategic guidance, new military paradigms are being developed and analyzed through the Defense Planning process. The DOD defines the term Defense Planning as “the employment of analytical, planning, and programming efforts to determine what sort of armed forces are appropriate for a state.” 1 This process consists of translating strategic level guidance provided by the National Defense Strategy (NDS), National Military Strategy (NMS), and/or National Security Strategy (NSS) into spending priorities, a portfolio of provided capabilities, and the needed force structure. The goal of this process is to generate a comprehensive force structure that can meet future strategic and operational challenges while implementing feasible and affordable capabilities. In implementation, this process leverages the use of specific scenarios that are defined by senior leaders based on global trends and intelligence reports. These scenarios are then used to conduct a variety of analyses with the goal of being able to compare alternative decisions based on different forms of modeling and simulation. As noted in a RAND report in 2019, 1 there are some challenges with this scenario-based analysis approach which includes the following:
Appropriate scenario implementation;
Scenario selection process can be subjective and arbitrary;
Scenario assessment has become so intensive and rigid that it constrains responsiveness;
Concepts arbitrarily over-determine requirements derived from a limited set of scenarios or force structure sizing constructs.
The overarching conclusion from the report showed that the fundamental structure of the process is sound and can be effective in supporting senior decision makers. However, the current implementation could be improved to increase the responsiveness of the process to better address the dynamic nature of the future threat environment. In addition, there is a need to find ways to explore more non-traditional scenarios and concepts not only to foster innovation but also to make future force constructs more robust to unexpected situations. 1
Another report in 2019 2 was issued by the US Government Accountability Office (GAO) recommending a revised analytical approach to the current defense planning process. This process was originally established in 2002, called Support for Strategic Analysis (SSA). 3 In this report, the GAO documented that the current analytical processes were not properly supporting senior decision makers in determination and evaluation of needed force structures. Major findings included the following:
Analysis products are inflexible and cumbersome;
Analyses lack significant deviation from currently programmed service structures;
Key assumptions are not tested;
There is a lack of joint analytical capabilities to assess different force architectures.
This research addresses shortcomings with strategic analyses through the formulation and demonstration of an advanced modeling methodology, specifically to avoid computationally prohibitive aspects of military operations analyses. To further decompose and refine this objective, the following sections introduce background on various modeling and simulation techniques, along with specific concerns related to military analyses. Following that, the Reduced Order Non-INtrusive (RONIN) modeling methodology is defined and applied to a demonstration use case. Finally, a discussion of results from the demonstration is presented, along with potential future efforts related to this research.
2. Background
In the following section, different aspects related to how the military leverages modeling and simulation (M&S) are discussed, along with different methodologies for addressing different concerns pertaining to providing relevant information for decision makers. In addition, the concepts of surrogate modeling and reduced order modeling will be introduced as specific methods for abstracting model behavior, specifically focused on capturing as much variability as possible in a computationally efficiently manner.
2.1. Military operations modeling and simulation
While the current implementation of SSA might not be delivering the information needed to senior leaders, this does not mean that the DOD lacks the means to conduct the required analyses. In fact, the DOD has many different analytical capabilities that have been leveraged in the past. 4 RAND has published work that assesses the different strengths and weaknesses of varying types of analyses grading each analysis on a variety of matters from representation of different phenomena to strategic-level functionality. Figure 1 shows this comparison of different analytical instruments with regard to their resolution, functionality, and ability model different phenomena.

Assessment of defense planning analyses. 4 Ratings are 1 (very poor) to 5 (very good), with red, orange, yellow, light green, and green corresponding to 1, 2, 3, 4, and 5, respectively. Scores depend on assumptions.
In general, analyses that are used to provide strategic decision support and integrate well with other strategic-level analyses are less suited to representing human and physical phenomena. This is due to the scope and scale of these types of analyses. In addition, the computation, temporal, and data expenses to conduct these analyses with high-fidelity M&S becomes too prohibitive to be practical. Likewise, the analyses that are excellent at capturing the behavior of different physical and human phenomena are not good for providing strategic insights because the scope is too narrow to provide senior leaders the required information to effectively make informed decisions. 4
Tolk 5 describes how to model key elements, such as the environment, movement, command, and control, among others, to create effective combat simulations. These elements contribute in different ways to understanding how decisions made about Concepts of Operation (CONOPs), tactics, and capability employment affect the outcomes of battles or campaigns. While incorporation of all these elements is needed, it is not always feasible or practical to model an element to the highest resolution possible. Tolk 5 describes how focusing on high-fidelity modeling can lead to prohibitive computationally expensive simulations, along with losing the ability to understand the bigger picture, “not being able to see the forest through the trees” as it were. To address these concerns, the concept of model aggregation or abstraction was introduced, which is focused on capturing the key effects, interdependencies, and behaviors of the high-fidelity model in some lower dimensional form. 6 This can be done for both deterministic and stochastic processes, where for stochastic processes, the underlying individual probability density functions should be derived and emulated as effectively as possible. 7
From the works of Tolk, Davis, and others, along with advances in the field of military operations M&S through the years, the DOD has continually sought to leverage these lessons and advances to provide analytical support for strategic decisions where different types of analyses in a multi-level approach are brought together to provide “at-time decision aiding” that includes a framework, story, optional focus points for explanations, trades, and what-if scenarios. Davis and Bigelow 8 introduced the concept of Multi-Resolution Modeling, or Multiresolution Modeling (MRM), that is defined as “building a single model, an integrated family of models, or a combination, with different levels of resolution for the same phenomena.” 9 The idea of being able to analyze a specific phenomenon through different levels of resolution or abstraction offers some advantages when compared to a single resolution approach. The MRM approach includes more explanatory power, traceability, and more efficient use of computational resources, along with a means to account for different emergent behaviors at different levels. 10
The terms resolution and fidelity are present throughout literature related to various modeling and simulation concepts. At times, they are used interchangeably, while the terms are similar, they do differ; for this research, working definitions are pulled from Cox. 11 In simple terms, fidelity is the accuracy of a model or simulation as compared to the real world, whereas resolution is the “level of detail” that is represented in the model or simulation. A high-resolution model may represent how the engine and components combine to generate system performance such as accelerations, whereas a low-resolution model may only indicate the ability of a car to move passengers various distances over time. The MRM approach conceptualized through the different levels of abstraction in DOD M&S efforts is through the “common military M&S pyramid.” 12 This hierarchy provides a method for not only helping define the different levels of models but also their relationship to other models at different levels within the hierarchy to better understand causal relationships. The bottom or foundation of the pyramid is where the highest resolution models are used. As the level of abstraction increases going up the pyramid, the resolution decreases while attempting to maintain the highest fidelity possible. 13 As highlighted in Figure 2, specific methods for accomplishing model abstraction from higher resolution foundation models to lower resolution, hopefully high fidelity, strategic models are introduced as means for incorporating different elements of useful information or behavior across the spectrum modeling levels. These examples include Metamodels, also referred to as surrogate models, reduced order models (ROMs), and RONIN models, which are discussed further along in this paper.

Multi-level approach to M&S.
In this multi-level approach, different models and simulations are leveraged for their various strengths. Ultimately, though, models or simulations of potential DOD capabilities should aspire to be “causal in nature as to allow us to understand degree of mission success as a function of input variables expressed parametrically.” 4 With the DOD moving to reach this aspirational goal, some considerations have been highlighted by different literary sources as key to creating models and simulations that effectively provide support to decision makers.
First, the battlefield is increasingly being characterized as a complex System-of-Systems (SoS) where non-linear interdependencies among elements in the battlespace make it difficult to quantitatively assess the value of causal links between capabilities and mission effects. 15 These non-linear effects stem from concepts like “force multipliers” or critical support functions, which might not direct physically impact an adversary in the battlespace but enable other elements to perform their tasks more effectively. Second, when developing parametric models of potential capabilities, traditional regression approaches have difficulty representing the underlying physical phenomena (i.e., linear methods applied to non-linear problem) and variables used might not be the appropriate ones to focus on. 4 This can be particularly difficult when attempting to model complex system-of-systems environments, which have a plethora of different phenomena and many physical interactions and responses. This becomes prohibitive when attempting to model every interaction in large-scale models, which are shown in Figure 2 as campaign-level model or even some mission-level models, simply from the sheer scale of interactions. Finally, stochastic effects should be incorporated into models because settling on simple expected values does not allow for proper representation of multi-modal or skewed distributions, nor do they provide critical information about the uncertainty of the estimated outcomes. 16 These probabilistic interactions come from the uncertainties discussed previously, where probabilities or distributions of outcomes are used to account for behavior that is not explicitly modeled. Ultimately, the growing complexities of any SoS problem like those of future battlefields require complex M&S approaches which can provide useful information to decision makers at the speed of relevance.
2.2. Toward surrogate modeling
As the complexity of varying models and simulations used by the DOD increases, so has the computational expense, such that the time and data requirements are becoming major limiting factors. This increasing expense, coupled with the push to leverage multi-resolution techniques, has led to the need for methods that can emulate the behavior of the high-fidelity models while reducing the computation burden. Surrogate modeling methods are a popular approach to meet this need. Some of the earliest work to leverage surrogate modeling methods dates to the 1970s where Mihran published a procedure for creating a “simular response surface” to efficiently utilize stochastic engineering models for design and analysis. 17
One of the noted advantages of surrogate modeling methods is the inexpensive computational cost associated with using them for evaluation, which has led to their widespread use in analysis and design problems. In general, surrogate models can be described as mathematical constructs that provide approximations of response behavior at a reduced computational burden compared to the original model.
18
A critical aspect of this form of model abstraction is its ability to minimize the loss of accuracy due to its construction process. In mathematical terms, an expensive model of a physical phenomena could be denoted as
Evaluating a surrogate model is relatively inexpensive; however, there is the potential for a higher computational cost in the surrogate model’s creation. To disambiguate the benefits and drawbacks to surrogate modeling, literature uses the terms online and offline when referring to its associated computational expenses. Where online cost refers to the inexpensive evaluation of the surrogate model compared to the original model, and offline cost refers to the up-front computational expense to create the surrogate model, which at times can be quite large. Even though there is the potential of a large offline cost to create a surrogate model, this cost can be justified by viewing surrogate modeling approaches as means of producing artifacts, which once created can be used in a wide variety of different applications. This characteristic of surrogate modeling is considered one of its most notable benefits. 19 Surrogate modeling methods represent the current state of the art for representing response behavior from computationally expensive models or simulations, often found at the higher resolution levels of the M&S pyramid.
Surrogate modeling is broken down into three classifications by Eldred and Dunlavy, 20 the data fit type, the model hierarchy type, and the ROM type. This research focused on the reduced order model type, which can be generated through a variety of methods, and an effective set of methods are based on the concept of mathematical projection. A projection-based reduced order surrogate model of a high-resolution model is generated using a mathematical projection, as in representing higher dimensional data in a lower dimensional form, of the original high-dimensional field down to a reduced number of generalized points. One example of how advanced ROM techniques are being leveraged is in the generation of predictive models to improve design methods for novel concepts in aerospace engineering where traditional empirical-based methods are not applicable. 21
The generation of non-intrusive ROMs, where the governing equations for a high-fidelity model are treated a black-box, is defined in this research by four fundamental steps:
Data Generation;
Dimensionality reduction/Projection;
Interpolation/Regression;
Back Mapping.
These steps can be organized to show how the input space, output fields, training data, and reduced order representation are all related, as seen in Figure 3.

Method for generation of reduced order model (ROM).
To begin the creation of a ROM, the data generation step, which is universal to all surrogate modeling techniques, is done using the original model with the goal of effectively and efficiently sampling a given input space, such that as much of the original model variability is captured within a given set of bounds using the least number of samples. The data generation step can also include some aggregation or abstraction of the behavior observed in the original full-order model (FOM), the goal of the step is to gather the useful data for the specific application that the ROM will be used for. Following that, a dimensionality reduction process, most often done through some sort of projection, is completed on the data that was generated by the original model. Next, a regression or interpolation method is applied to map the original input vector to the new reduced order representation of the original data. Finally, once a regression model is adequately trained, an estimate of the reduced order space can be generated and used to recover a full-order representation through a back mapping algorithm as defined by the original dimensionality reduction approach; this is done to recover the FOM variability in a computationally efficient manner. 21
After reviewing which advanced modeling methods might be suitable for reaching the aspirational goals highlighted in literature, a ROM-based approach was selected because of its ability to handle high-dimensional problems while still allowing for traceability. In addition to identifying various linear and non-linear ROM methods, several stochastic modeling methods, such as Bayesian Gaussian Mixture models, Metalog distributions, and Copula Joint distributions, were identified and evaluated. Along with an assessment of how well different encoding methods performed when transforming categorical variables into numerical values for effective integration with the various ROM and stochastic methods. After evaluating the strengths and weaknesses of these various methods, the RONIN modeling methodology was formulated, and an example demonstration was developed to illustrate how the methodology could be implemented using a standard DOD M&S framework.
3. RONIN methodology
The Reduced Order Non-INtrusive (RONIN) modeling methodology, shown in Figure 4, takes an adaptation of a standard approach to developing predictive surrogate models 22 and augments the step of creating a regression model with the needed steps to leverage ROM methods to create output fields of distributions rather than deterministic scalar values.

RONIN modeling methodology.
3.1. Step 1: data generation
The first step in the RONIN process, which is common to many techniques, is the generation of data from a full-order agent-based or discrete-event model that the predictive surrogate model is trying to emulate. The FOM is run multiple times using a set of input factors or values to generate a set of response distributions. The input and response distribution data sets are then separated into training and validation data sets for model training and performance assessment.
3.2. Step 2: data encoding
Following the generation of full-order model data, the encoding step works to transform the generated data into a format which can be leveraged by various order reduction techniques. This transformation is completed by first defining a grid of points, which is properly suited for the specific information to be predicted, similar to how a mesh is defined for Finite Element Analysis or Computational Fluid Dynamics. Then, the full-order data is represented as the distribution of responses for each grid point. To complement this grid approach, a variety of different encoding methods could be applied, and for this research, an encoding method which leveraged concepts from One-Hot 23 and Code Counting 24 approaches was used to create a more detailed higher dimensional representation of the information of interest.
3.3. Step 3: order reduction
Once the data encoding is complete, the order reduction process works to create a latent space representation of the response distributions. This can be accomplished through a variety of linear and non-linear approaches, all of which work to capture as much variability as possible from the full-order data in a lower dimensional form. In this demonstration, the non-linear approach known as ISOMAP was leveraged to accomplish the dimensionality reduction through the following high-level steps outlined by Izenman: 25
Construct a graph with the data points as vertices using neighborhood information around each point.
Transform graph into a form that is suitable for the next step using a procedure unique to the algorithm.
Perform a spectral embedding by solving an eigenproblem.
This process yields a lower dimensional representation of the full-order data called an embedding, along with the ability to use the inverse of the transformation function to back map from the lower dimensional space to the full-order space.
3.4. Step 4: fit stochastic model
This step uses the reduced order distributions and stochastic modeling methods to find a set of deterministic values, the hyperparameters of the stochastic model, that will be used in the following regression modeling step. The latent space data is used to fit a parametric stochastic model such that the hyperparameters from the trained stochastic model for each input setting are vectorized. In turn, the individual vectorized hyperparameters are used to create a matrix which is made up of all the hyperparameters from each input factor configuration. For this demonstration, a specific form of a Gaussian Mixture Model (GMM) was utilized called a Bayesian GMM (BGMM). This approach follows the standard GMM approach, which works to capture multi-modal or skewed behavior using multiple Gaussian distributions that are combined. However, the BGMM works to minimize the number of components to characterize the multi-modal behavior through a Dirichlet process prior distribution, which results in a more efficient model that is not as susceptible to over-fitting, singularities, or outliers. 26
A practical formulation to utilize the Dirichlet process distribution using a variational Bayes approach was published by Attias, 27 which started from the model of form:
Such that,
where
Such that,
As Attias noted, the parameter posteriors will emerge in suitable, non-trivial, and generally non-Normal functional forms along with having no issues with divergence, all without utilizing any specific prior assumptions. In addition, the approximate marginal likelihood,
3.5. Step 5: fit regression model
Now that the matrix of hyperparameters has been constructed, the embedded regression step works to train a predictive regression model that uses the original model inputs to map to the stochastic model hyperparameters, which is a deterministic set of values. Once the training of the regression model is complete, it can now be used to estimate hyperparameters for values within the bounds of the original training design of experiments (DOE).
For this research, the regression modeling approach used was a form of Gaussian process regression called the Matérn kernel, which is a generalization of the radial basis function (RBF) kernel. This approach was selected based on its performance related estimating latent space representations of shocks in supersonic flow, which, mathematically speaking, manifest as abrupt discontinuities. 21 In general, Gaussian process regression works to capture the relations between inputs and responses by utilizing a theoretically infinite number of parameters, such that the original data determines the model complexity through Bayesian inference. 28
To accomplish this, the regression process focuses on use of distributions over functions, as in, instead of solving for specific weights,
Such that
where
3.6. Step 6: model assessment
To complete the model assessment, the reconstruction of full-order data must be accomplished. Beginning with the estimation of the hyperparameters, the reconstruction phase works to recover an estimation of the original response distributions. This is accomplished by first using the estimated hyperparameters incorporated into the stochastic model to generate an estimation of the latent space distributions. This is then followed by using said latent space distributions along with the ROM back mapping to generate the estimation of full-order distributions.
Finally, the model performance assessment phase works to quantify how well the predictive RONIN model performed in recreating the original response distributions. This is accomplished through the use of metrics which focus on comparing the estimated probably density functions (PDFs) and cumulative distribution functions (CDFs) and the true PDFs/CDFs with the goal of showing that a majority of the estimated response samples came from the same underlying distribution as the truth samples. For this research, the estimated and original distributions were compared using a 2-sample Kolmogorov–Smirov (K-S) test to compare the CDFs and the Jensen–Shannon distance (JSD) to compare the PDFs. The K-S test uses a null hypothesis that argues that the two samples have the same underlying distribution. If the K-S statistic falls below a determined critical value, the test will fail to reject the null hypothesis, which indicates that there is a lack of evidence to assume that the two data sets come from different distributions.
32
In this assessment, failure to reject the null hypothesis is considered a desirable outcome because it communicates, by this test metric, that predicted data and the truth data appear to come from the same underlying distribution. The equation for the K-S test statistic is shown below, such that,
Alternatively, the JSD is a measure of how closely two PDFs match one another by symmetrizing the Kullback-Leibler (K-L) divergence
33
criteria and bounding it between 0 and 1, such that the closer the value is to 0, the smaller the difference in the PDFs.
34
In this assessment, JSD values close to 0 are considered desirable because it indicates the predicted PDF closely matches the truth PDF. The equations below show the calculation of the JSD (
If this assessment indicates unacceptable performance, a closer look at the methods selected is necessary. Due to the critical placement of the embedded regression within the methodology as a whole, assessment of the regression method selection should be accomplished first, because any errors made in the estimation of the stochastic hyperparameters will propagate the furthest and increase the uncertainty the most. Once it is confirmed that the regression selection is performing adequately, the stochastic model selection should be assessed to capture how well the selected approach is estimating the latent space distributions. Following confirmation of the stochastic model performance, the dimensionality reduction method should be assessed to minimize the error induced from the projection and reconstruction phases. Finally, if all other selections are performing adequately, then grid definition and encoding should be assessed to make sure no extraneous error is being introduced in that step.
3.7. Step 7: model integration
Once satisfied with the performance of the generated ROM, it can be utilized as needed. Some general examples include multidisciplinary design and optimization of aerospace systems, 35 design space exploration for building design, 36 and decision space exploration for technology assessment. 37 A more specific example of how a predictive model like a trained RONIN model could be implemented comes from the concept of trade space exploration and exploitation, where an analyst or decision maker could be interested in using an evolutionary algorithm to characterize the Pareto Frontier of a multidisciplinary problem that utilizes a computational expensive simulation to calculate a variety of objective metrics. 22 The Pareto Frontier is made up of “non-dominated” solutions which are considered the best performing solutions in terms of one or more objectives. When it comes to multi-attribute/multi-discipline problems, it can be very insightful to characterize trends in this group of good solutions rather than just picking a single solution. A RONIN model in this context would work to replace the expensive full-order simulation such that the response behavior is captured given the bounds of the potential input vectors. Then, once the predictive RONIN model is trained and validated that it captures the required behavior correctly, it is used to generate responses as the evolutionary algorithm explores the trade space through changing in the input vector and comparing estimated responses. An example of this type of analysis can be found in the study by Felten et al. 38 who worked to efficiently explore and optimize different Space Domain Awareness architectures using a Genetic Algorithm (a specific type of evolutionary algorithm) and DOD supercomputers. The authors utilized a full-order orbital propagation simulation and a set of look-up tables to calculate multiple objective metrics, such as cost and revisit rate, among others, which is what drove the requirement to leverage a supercomputer to complete the analysis in a reasonable amount of time. Alternately, if a predictive RONIN model had been trained from the full-order orbital propagation simulation, the need for a supercomputer may not have been warranted, or at the very least, the trade space explored could have been exponentially larger.
In addition to optimal space architecture exploration, the RONIN methodology can be applied to mission sets to accelerate campaign-level analyses or wargaming exercises. Application of the RONIN methodology will next be explored, where instead of modeling the entire full-order simulation, the analyst selects a specific subset of missions, events, or engagements to train a RONIN model. Then, the trained RONIN model is called by the high-level simulation.
4. Application of RONIN methodology
To demonstrate a specific application of the RONIN methodology, a specific level of analysis within the simulation hierarchy was selected—the mission level model. This level of M&S is focused on modeling systems and SoS working together to accomplish a particular mission. 6 An example of a military application related to a SoS working together would be having a strike package of different types of aircraft working in concert to suppress enemy air defenses. This level of simulation typically involves timescales that measure hours in length and is focused on questions such as “How best to employ a system?.” Some examples of mission level simulations include Suppressor, Extended Air Defense Simulation (EADSIM), and Advanced Framework for Simulation, Integration, and Modeling (AFSIM).6,13
4.1. AFSIM
There is currently emphasis from the US Air Force to embrace simulation frameworks that provide a wide variety of capabilities and are accessible to a large community through open modular architectures. At the mission level, AFSIM is one preferred framework for many organizations. 39 The DOD owns the technical baseline for AFSIM, where Air Force Research Laboratory (AFRL) is charged with managing the support and continuous improvements that are most needed by the military analysis community.
AFSIM is an agent-based modeling framework that simulates a wide range of systems, from surface assets (facilities and ground vehicles) to aircraft to space assets. The behaviors and interactions between the agents and other agents or the environment are guided by behavior trees, which can be as simple or complex as needed. 40 The simulations that are run in AFSIM can be either deterministic or stochastic based on the data used to define an agent’s behaviors and the governing equations that determine the outcomes of interactions. The versatility and widespread adoption by the military analysis community made AFSIM a suitable framework for demonstrating the RONIN methodology on relevant DOD use cases.
4.2. Use case scenario
With the framework selected, next a specific mission was identified to generate the use case for demonstration. A Suppression of Enemy Air Defenses (SEAD) mission was selected to create a representative scenario that is of interest to capability development planners. USAF doctrine, AFDP 3-01 Counterair Operations, describes SEAD as an offensive counterair operation which is meant to neutralize, destroy, or degrade enemy ground-based air defenses. 2 The SEAD mission implemented in this demonstration consists of a blue force attacking a red Integrated Air Defense System (IADS), shown in Figure 5. The objective is to target high-payoff air defense resources, to then allow greater freedom of movement for other aircraft. 41

Use case scenario depiction.
This demonstration involved varying the composition of the blue force structure, which comprised three different fighter aircraft archetypes (Strike, SEAD, and Escort), each of which had different capabilities and behaviors. The red IADS was comprised of up to three Surface-to-Air Missile (SAM) battalions with varying amounts of missile launchers and supporting radars, along with up to four fighter aircraft which are defending the airspace above the IADS. Finally, the last component of the scenario was the attack decision by the blue forces; the space was discretized into nine grids as shown in Figure 5. The utility of a grid definition to geographically divide a battlespace coupled with various encoding methods was demonstrated by Bateman et al. 42 For this research, each defined grid is referred to as an area of interest (AOI), and the blue forces were directed to attack a specific AOI and destroy any red forces in that AOI along with any red fighters which tried to attack them. During the attack phase, the blue fighters would only engage with threats that they were specific equipped to engage and that posed a threat, for example, the SEAD fighters would only engage the SAM battalions they were in the specific AOI directed or in route to the AOI directed.
This mission set up was selected because it represents a combined packaged approach. The concept of using a combined package leverages different types of specialized aircraft to work together to achieve a favorable mission outcome. For this mission, a favorable outcome is where the number of enemy assets destroyed is increased, while keeping the number of friendly assets destroyed low. Each of the specific aircraft types could not achieve a favorable mission outcome on their own, for example, the blue air superiority aircraft, designated “Escort” for this scenario, have specific capabilities and behaviors which allow them to engage the red intercepting aircraft more effectively. However, they are not properly equipped to engage the ground-based threats presented by the SAM battalions and must rely on the SEAD aircraft to deal with that threat. This type of complementary strategy is seen in a wide variety of current military doctrine and tactics, so it is often conceptualized through the lens of SoS engineering, 43 where many different systems are brought together to achieve an objective that none of the individual systems could achieve on its own. This conceptualization helps provide a suitable framework not only to analyze how to best implement a system of systems (SoS) to achieve these campaign objectives but also to highlight the challenges in engineering a SoS.
4.3. Demonstration set up
The purpose of the demonstration was to show how a predictive RONIN model could be constructed from a DOD standard agent-based simulation and quantify its performance to emulate a set of output distributions. To set up the demonstration for this SEAD use case, a DOE was used to generate 1000 different input configurations which varied the blue and red force structures, along with the AOI that blue would attack. Table 1 provides a breakdown of the input parameters for the scenario with their minimum and maximum values. Each of these configurations was run through AFSIM 65 times to create a set of output response distributions which captured the number of remaining agents after the mission was completed. Due to the limited sampling budget available, selecting 65 runs increased the likelihood of capturing rare events or outcomes in the response distributions, while still attempting to explore a larger variety of input combinations. In contrast to more traditional scalar approaches which would rely on aggregation of the response distributions, to something like a percent survivability value, the RONIN model captures the underlying stochastic behavior without the need for aggregation or abstraction, thus, better preserving the full model’s variability. This demonstrated approach shows how a stochastic predictive model could be generated which could be used to rapidly play out multiple decision threads in a non-deterministic way without needing to rerun the original model. Alternatively, the predictive RONIN model could maintain knowledge of the full scenario at each decision point to provide context for events as they occur and not just at the conclusion of the mission.
Scenario parameter values.
4.4. Results
From the AFSIM scenario, the outputs included the remaining number of agents at the time of the conclusion of the scenario; this information was then translated into distributions of specific agents remaining in specific AOIs. For example, for a specific input configuration, the output distribution of Blue SEAD aircraft remaining in AOI 1 was listed as “AOI_1_BSEAD,” shown on the x-axis of Figures 6 and 7. This output data from the scenario was split into training and validation data sets, such that the predictive RONIN model was generated using the training data set and the validation data was set aside to confirm model performance later. Once the predictive RONIN model was generated, it could then be used to generate samples from the estimated response distributions, effectively emulating the behavior of running the full order of the AFSIM model in a fraction of the time. These estimated sample distributions of remaining agents were then compared to the training and validation sample distributions through the K-S test and JSD calculation. For this demonstration, the 1000 sample distributions were split using an 80/20 training to validation ratio, where each collection of 65 replications was considered a single sample distribution and was kept together as a distribution that would be used to either train or validate the RONIN model.

K-S p-value heatmaps: training (top) and validation (bottom) data sets.

JSD value heatmaps: training (top) and validation (bottom) data sets.
Overall, the predictive RONIN model performed very well in terms of creating estimated response distributions which matched the original response distributions. Table 2 lists the average percent of grids which passed the K-S test with a 95% confidence level, along with the associated variance for both the training and validation data sets. As shown by a large percentage of grids passing the K-S test, the CDFs are very similar. In addition, the average JSD value with its variance for both the training and validation sets are also listed. The very small JSD values show that the PDFs are similar. With the predictive RONIN model being stochastic in nature, each time the model is queried the specific estimated sample might not specifically match a training or validation sample. That is to say, for “Run 1” the number of remaining agents might not match specifically but because the good performance in terms of the stochastic modeling metrics, there is a high degree of confidence that both samples are from the same underlying distribution and if more samples were taken then they would align overall. This good performance was enabled by lessons learned through a set of foundational experiments, which are further described in Bateman, 44 along with having a large enough sample size (n = 1000) to allow for better training of the embedded regression model. The RONIN approach to capturing the stochastic nature of the full-order model provides a powerful tool in rapidly modeling a variety of inherently non-deterministic attributes, such as uncertainty, in more traceable way, as compared to deep learning methods like neural nets, while still being computationally efficient like more simplistic deterministic scalar models. In conducting this demonstration, we observed the computational efficiency but did not measure it empirically. Anecdotally, the AFSIM scenario needed approximately 5–10 s per run, thus it took minutes to get a distribution of outcomes, and the RONIN model generated a similar outcome distribution in less than 1 s.
Summarized demonstration results.
A closer look into the performance of the model was conducted using the following heat maps, which provide a visual representation of the K-S test p-values (Figure 6) and JSD values (Figure 7) in terms of a given grid location and input setting. The K-S test p-value quantifies the evidence, in this case the K-S test statistic, against a null hypothesis. Consequently, a low p-value suggests the data is inconsistent with the null, potentially favoring an alternative hypothesis. This visual inspection of the data allows for identification of potential problematic grid locations, which could have highly non-linear behavior in the original data set. Also, this inspection allows identification of problematic input settings that were difficult for the RONIN model to capture. For this specific demonstration, a low p-value would indicate that there is a significant difference between the RONIN output distribution predictions and the original AFSIM output distributions. This could suggest that the ROM is not accurately capturing the essential behavior of the model for a given grid and input combination. Conversely, a high p-value would indicate that there is not enough evidence to conclude a significant difference, suggesting that the ROM is performing well in replicating the behavior of the model for a given grid and input combination. Therefore, the primarily light-colored p-value heat map for the K-S test in Figure 6 indicates the demonstration model performed very well in estimating the original response CDFs, and the error appears to be randomly distributed for the training and validation data sets. Thus, there appeared to be no localized decreases in model performance in terms of estimating the response distribution based on either the grid location or input setting.
From the JSD heat maps in Figure 7, the dominance of the darker colors shows that many of the values are close to 0. As discussed previously, the JSD value is a normalized measurement of the difference in the PDFs, such that as the JSD approaches 0, the smaller the difference in the PDFs. Thus, demonstrating the model performed very well in estimating the original response PDFs and the error appears to be randomly distributed for the training and validation data sets. Once again, there appeared to be no localized decrease in model performance in terms of estimating the response distribution based on either the grid location or input setting.
The error in the model’s ability to emulate the original response distributions originates from slight inaccuracies when the embedded regression model, the Matérn 3/2 kernel, estimated the hyperparameters for the BGMM. This error in turn propagated to slightly inaccurate latent space distributions, which ultimately led to the measured differences in the full-order distributions. While each step does introduce new sources of error that will propagate through steps further along the process, a key source of error is the embedded regression model, because of its location in relation to other parts of the process. As such, further investigation into different regression or interpolation methods has been identified as area for future research and refinement.
While this demonstration did showcase the potential for the RONIN modeling methodology to effectively generate distributions which were found to be stochastically similar to the FOM distributions, there is a key limitation related to the setup of the original scenario. Like most surrogate modeling techniques, extrapolation outside of the original training and validation data is a prominent cause of poor performance in models created through the RONIN process. For this particular demonstration, if there is a change to the scenario outside the changes made through the DOE used, then the RONIN surrogate model will not create output distributions that would match the FOM. For example, if the laydown of the SAM sites is changed or the capabilities blue forces are modified then a new set training and validation must be generated to generate a new RONIN surrogate model.
5. Conclusion and future work
In conclusion, this demonstration was able to describe in detail how to implement the full RONIN modeling methodology to create a predictive surrogate model from a relevant use case. This predictive model was able to accurately estimate distributions of responses which were originally generated from AFSIM, a simulation framework that is commonly used by the DOD. These results effectively demonstrate that the original research objective, which was the formulation and demonstration of a methodology which leverages reduced order modeling methods for traceable model abstraction that effectively and efficiently captures complex system-of-systems behaviors within current military operations M&S, was met. For this demonstration, the computational efficiency was noted but not measured empirically, however anecdotally, the AFSIM scenario needed approximately 5–10 s per run, thus it took minutes to get a distribution of outcomes. Where the RONIN surrogate model needed less than 1 s to do the same thing; thus, for larger trade spaces with potentially thousands of distributions, the evaluation time goes from hours of simulation time for the FOM to minutes with the RONIN surrogate model. In summary, the RONIN modeling methodology works to create surrogate models which emulate stochastic FOMs by leveraging order reduction through either linear or non-linear methods, stochastic modeling methods such as BGMM, and regression modeling methods such as a Matérn 3/2 kernel. The generated RONIN surrogate model can take in the inputs of a given scenario which would be run numerous times in a stochastic simulation and output a distribution of responses that closely matches the output distribution from the full-order simulation. This application of reduced order and stochastic approaches gains the advantage of computational efficiency while still preserving the stochastic nature of advanced simulation frameworks. These advantages are key to enabling many query analyses such as trade space exploration or design space optimization.
By introducing and exploring advanced approaches for order reduction coupled with stochastic modeling and how they can be leveraged with current DOD simulation frameworks, this work is meant to serve as another step toward more effective MRM. While this work built upon current ROM approaches by addressing concerns related to accounting for non-linear behavior in discrete fields, stochastic inter-dependencies, and physical interactions, it did not focus on optimization of the embedded regression approach. As previously noted, the regression model utilized through these experiments and demonstration was selected based on its performance related estimating latent space representations which included abrupt discontinuities. From the error analysis conducted throughout this research, it was observed that the regression model did introduce some error depending on the characteristics of the data set being used for training but overall, it performed well.
In future work, a generalized approach to the “Regression” step in the RONIN process should be formulated such that the error in this critical step is reduced as much as possible. To improve the performance of the methodology, a generalized approach to the regression modeling could be implemented that leverages concepts like kernel density estimation, 45 kriging/Gaussian process interpolation, 46 response surface methodologies, 22 radial basis functions, 47 or polynomial chaos expansion, 48 just to name a few. The generalization would be in the selection of the most effective regression approach based on predictive performance and model efficiency. Said another way, the generalized approach would work to find the simplest model which meets the required level of accuracy. A key part of that will be the metrics used for model comparison, from literature, Mallows’ Cp, 49 Akaike’s Information Criterion, 50 or Bayesian Information Criterion 51 were all identified as standard means to comparing model performance. A dedicated research effort in regression model selection and implementation as it applies the RONIN methodology and DOD applications would provide military operations analysts with even better capabilities in terms of providing support to senior decision makers.
Footnotes
Acknowledgements
The views expressed in this article are those of the authors and do not reflect the official policy or position of the United States Air Force, United State Space Force, or the Department of Defense. This publication has been reviewed and deemed cleared for public release, case number: 88ABW-2024-0533.
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
