Abstract
Traffic safety research will continue to create new and improved methods for the analysis of safety data. Even if these models perform well, the precise underlying crash mechanism remains unknown. The missing piece is a tool that may be used to evaluate how well a method identifies the cause-and-effect relationship in the data. To meet these safety analysis needs, a high-resolution disaggregate data generating process called realistic artificial data (RAD) was developed. This tool simulates crash incidence on transportation facilities, capturing real-world causal links between individual roadway characteristics and crashes. The objective of this study was to check if the stochasticity embedded in the RAD generation process will be consistent for different random seeds and miles of data generated from the tool. To accomplish this, 10 different datasets were generated from the RAD tool and estimated using the negative binomial model; parameter estimates from the model were checked using a revised Wald statistic. The t-statistic estimates showed that the differences among the parameter values across the dataset are within a statistically acceptable level. Given the stability of the tool, the RAD framework can be useful in addressing the known limitations and knowledge gap in assessing the extent to which a statistical method succeeds in identifying the cause-and-effect relationship in the data. This in return can help guide and improve the practical application of statistical methods and eventually lead to more effective safety countermeasures that can reduce highway-related injuries and fatalities.
Keywords
Every year, approximately 1.3 million people’s lives are cut short as a result of highway crashes. Additionally, between 20 and 50 million people suffer non-fatal injuries, with many of them resulting in disability ( 1 ). The entire economic cost of fatal and non-fatal preventable incidents in 2020 was over 1 billion dollars, including vehicle damage, lost wages and productivity, and medical expenses ( 2 ). Because of the enormous societal costs associated with these crashes, research for many decades has focused on developing qualitative and quantitative information on how to make roads safer ( 3 ). Despite the strides made in the direction of the societal goal of Vision Zero through targeted legislation and the implementation of relevant safety-related measures, much work remains to be done in the field of safety analysis ( 4 ).
The goal of any traffic safety analysis is to identify and quantify the influence of factors contributing to the occurrence of traffic crashes and their associated consequences. Highway safety crash data have long been used to analyze safety problems, with uses ranging from identification of safety problems to determining their extent and creating models that are used to predict crashes ( 5 ). Even though the availability of safety crash data overall has expanded over the years, this does not necessarily mean that the quality of crash data is keeping up with the methodological advances ( 4 ). Earlier studies have tended to concentrate on the modeling component of the entire crash prediction process by developing new modeling approaches that offer superior fit, and the acquired data is implicitly believed to be a sufficient representation of reality.
Traditional statistical methods use archived safety data to guide the development of a model functional form, which is then used to estimate coefficients for variables, identify and compare significant factors, and finally compare which statistical method is best suited for a given scenario ( 6 ). Even if a model performs well, it may not accurately represent or identify the cause-and-effect relationship, because the intention of any developed model is prediction rather than identifying causal factors. Unfortunately, the detailed driving data and crash data that would better enable identification of cause-and-effect relationships in crash probabilities are typically not available ( 7 ). Most researchers have addressed this problem by framing their analytic approaches to study the factors that affect the number of crashes occurring on a roadway segment or intersection over some specified period.
Small sample size, time interval variation, and temporal and spatial autocorrelation are some of the most well-documented model estimation issues that have been raised in the literature ( 8 ). These problems are a potential source of error in modeling crash data that may cause incorrect estimates and inferences. Crash data, as previously mentioned in the literature ( 7 , 9–12), are frequently characterized by a sparse number of observations which can produce a low sample mean. This characteristic is attributed to the possibility that crash data for some roadway entities may include few observed crashes which results in a preponderance of zeros. Although it is believed that the size of the sample will have an impact on the crash prediction model’s performance, some have suggested a rule of thumb for data size requirements ( 7 , 8 , 10 ). A study reported that the dispersion parameter of Poisson-gamma models estimated from data characterized by low sample mean values and small sample size can be significantly biased (the value is likely to be mis-estimated) and negatively affect analyses commonly performed in highway safety ( 8 ). There is a significant increase in the probability that the dispersion parameter cannot be reliably estimated when the sample mean and sample size decrease.
The issue of time interval variance in crash data was also reported by other studies ( 7 , 10 ). Crash data are often collected over a certain time period. Over the collection period, some explanatory variables and their relationship to the crash incidents may change, a reality that is not usually considered because of the lack of detailed data within the collection period ( 10 ). Ignoring within-period variation in explanatory variables may result in biased estimation of parameters and incorrect prediction of crashes as a result of unobserved heterogeneity.
While the continuing march of methodological innovation has increased our understanding of the factors that affect crash frequencies, the potential of integrating improving methodology with significantly more detailed crash data offers the most promise for the future ( 3 ). To meet these safety analysis needs, a high-resolution disaggregate data generating process called realistic artificial data (RAD) was developed, which simulates crash incidence on transportation facilities. The tool was created to capture the real-world causal link between individual route parameters and crash statistics. The crash counts from the tool will be based on a set of (secret) “causal rules” that represent predetermined correlations between crash frequency/severity and specific roadway geometric characteristics ( 4 ).
Since the data generation process is now known, the RAD can serve as a testbed which will help us determine whether a statistical model developed indeed captures the underlying relationship between the independent variables and the resultant crashes, and this in turn will help guide and improve the practical application of statistical methods that will influence highway safety policy and eventually lead to more effective safety countermeasures that can reduce highway-related injuries and fatalities. In addition, there are many other questions which the RAD can help us answer, such as how many mile-years of data are necessary to produce valid results when a large amount of data is available.
The purpose of this study is to examine the stability of the parameters from different datasets generated from a RAD tool. The tool as stated earlier is a combination of a roadway generator and a crash generator with a predetermined causal relationship between roadway descriptors and crash frequency and severity. Estimation of crash prediction models for rural two-lane undivided highway segments was selected as a case study. Ten different datasets—two sets each of 150, 300, 500, 750, and 1000 mi—with different random seeds were generated from the tool. Negative binomial (NB) models were estimated with each of the datasets generated. Then we employed revised Wald statistics on the parameter estimates from the models to determine whether the estimates were stable across the different generated datasets. Examining the stability of the multiple datasets generated is important so we know that the randomness embedded in the RAD generation process is reliable.
Previous Work
The idea of RAD for traffic safety is not new; the early conception can be traced back to Dr Ezra Hauer who presented this idea at a 2008 TRB workshop titled “Future Directions in Highway Crash Data Modeling” ( 6 ). In a project funded by FHWA, Council et al. ( 13 ) investigated the effectiveness of various modeling approaches for cross-sectional studies. Highway Safety Information System (HSIS) data from Washington State were used to construct a dataset with 2400 mi of homogeneous segments that were each 0.02 mi long. They looked at single-vehicle lane-departure crashes on two-lane roadways. A modeler who was not aware of the presumed causal relationships was then given the crash and roadway data. The goal of the modeler was to estimate regression models and identify the causal relationships. The model’s outcomes were then compared with the assumed relationships. Another study by Lan and Srinivasan ( 14 ) utilized datasets produced by RAD for rural two-lane roads to examine the effectiveness of various regression models for calculating the crash modification factor (CMF) of horizontal curvature. To achieve this, three volunteers without prior knowledge of the embedded safety relationship used the RAD to estimate CMFs for horizontal curvature. This comparison was conducted for different levels of annual average daily traffic (AADT) and terrain. The CMFs estimated by the volunteers were then compared with the embedded safety relationship between horizontal curvature and crashes within the RAD. Horizontal curve radii that were higher rather than lower generally resulted in estimated CMFs that were closer to the true CMF values. Models that used site characteristics apart from AADT and curve radius usually performed better.
Miaou ( 15 ) studied the relationship between highway geometric characteristics and crashes using NB regression. Miaou suggested that the Poisson regression model should be used to establish the relationship between highway geometry and crashes. If overdispersion exists and is found to be moderate or high, the NB model can be explored. Another study by Abdel-Aty and Radwan ( 16 ) used NB modeling to model the frequency of crash occurrence, showing that high traffic volume, speeding, narrow shoulder width, and narrow lane width increase the likelihood of a crash. Some variations of count models have also been developed, such as zero-inflated NB models.
Anastasopoulos and Mannering ( 17 ) explored the use of random parameter count models as another methodological alternative in analyzing accident frequencies. Their findings showed that ignoring the possibility of random parameters when estimating count-data models can result in substantially different marginal effects and subsequent inferences relating to the magnitude of the effect of factors on accident frequencies. Shankar et al. ( 18 ) suggest that simple Poisson and NB modeling efforts do not address the possibility that some roadway sections observed to have no accidents during a specified time period may be qualitatively different from Poisson- or NB-distributed accident frequency counts.
A study by Malyshkina et al. ( 19 ) proposed a two-state Markov switching count-data model as an alternative to zero-inflated models to account for the preponderance of zeros sometimes observed in transportation count data. They proposed to overcome some of the criticism associated with the zero-accident state of the zero-inflated model by allowing individual roadway segments to switch between zero and normal-count states over time. They showed that the Markov switching model is a viable alternative and results in a superior statistical fit relative to the zero-inflated models. However, Lord et al. ( 20 ) in their study provided defensible guidance on how to appropriate model crash data. They suggested that carefully selecting the time/space scales for analysis, including an improved set of explanatory variables and unobserved heterogeneity effects in count regression models, or applying small-area statistical methods (observations with low exposure) represent the most defensible modeling approaches for datasets with a preponderance of zeros.
Other models that have been applied to crash frequency analysis based on their strength include multivariate models, which can model different crash types simultaneously ( 21 , 22 ); Poisson-lognormal models, which are more flexible than Poisson-gamma at handling overdispersion ( 23 , 24 ); and a generalized estimating equation for its ability to handle temporal correlation ( 25 – 27 ).
A Brief Overview of the RAD Generation Process
RAD Generation Process
The roadway generator and crash data generator are the two parts of the RAD framework used in this work. The roadway generator generates homogeneous segments with realistic road characteristics that represent variables typically found in the inventories of state transportation agencies. The Markov chain Monte Carlo principle is used in the road generator to assign specific values for the road characteristics. Markov chain is a systematic method for generating a sequence of random variables for which the current value is probabilistically dependent on the value of the prior variable. Specifically, selecting the next variable is only dependent on the last variable in the chain. The assignment characteristics of the developed roadway generator incorporate various Markov chain transition tables that determine probabilities of a descriptor changing based on data from various states ( 4 ). A random variable indicating how well the roadway characteristics change based on segment length according to the type of facility was also incorporated into the framework. That is, if you are trying to assign the shoulder width on a given segment, it has some high probability of being the same as the preceding segment but there is also some low probability that it will change.
The crash data generator is embedded with secret causal rules defining the true relationship between roadway descriptors and crash frequency and severities, that is, specifying how each roadway descriptor will affect a given crash type. Crash counts by crash type and severity are generated using well-known model structures and realistic relationships between crash counts and roadway characteristics; with this process, randomness in crashes per segment is generated for each user’s data ( 4 ). A system then combines the roadway and crash generator, producing a roadway file that attaches crash counts to each segment. The overall framework for generating the RAD is presented in the flow chart shown in Figure 1.

Realistic artificial data (RAD) generation flow chart.
The RAD generation tool is in the form of a database stand-alone software application that can be customized and run to prepare multiple realizations of data for various facility types under different random seeds and dataset specifications, such as total dataset mileage, using various combinations of inputs. This tool will be owned and operated by an entity where researchers who have no knowledge of the safety relationships that were embedded into the system can compare the results from their analysis. The degree to which they succeed will then be made evident by having the entity that owns the RAD generator compare the estimated parameters to the known relationships embedded in the data. The tool will serve as a testbed to help determine whether a statistical model indeed captures the underlying cause-and-effect relationship.
Crash Model Estimation Method
Although different mixed-Poisson distributions have been developed to model crash data (e.g., Poisson-lognormal, Poisson-inverse Gaussian, etc.), the most common distribution used for modeling crash data remains the Poisson-gamma, aka the NB distribution. The NB distribution offers a simple way to accommodate the overdispersion, especially since the final equation has a closed form and the mathematics to manipulate the relationship between the mean and the variance structures is relatively simple ( 12 ). The NB/Poisson-gamma model assumes that the Poisson parameter follows a gamma probability distribution. The model results in a closed-form equation and the mathematics to manipulate the relationship between the mean and the variance structures is relatively simple. The NB model is derived by rewriting the Poisson parameter for each observation i as
where EXP (
Data Generation
The data for the analysis was drawn from the RAD tool. The two-lane undivided roadway data contained horizontal curve data, crash data, and roadway data as described in Tables 1 to 3. To accomplish the objective of this study, all that was needed was to generate several RADs differing in mile-years of data and then to run various statistical models on all the datasets with the goal of showing the stability of the parameters from the dataset generated from the tool by examining the differences between the estimated and the assumed parameter values. Only the NB model for total crashes will be illustrated in this study because of space limitations. Ten different datasets—two sets each of 150, 300, 500, 750, 1000 mi—with different random seeds were generated. It should be noted that the same sized dataset (e.g., the two sets of 150 mi) resulted in different numbers of observations because of the data being randomly generated. In addition, the roadway characteristics generated had the same distribution as the sample size increase, which is a result of the effects described by the central limit theorem.
Descriptive Statistics for Continuous Variables (Datasets 1–5)
Note: SD = standard deviation; K- Fatal Crash; A-Incapacitating Crash; B-Non- incapacitating Crash; C- Possible Crash; PDO = property damage only; AADT = annual average daily traffic.
Descriptive Statistics for Continuous Variables (Datasets 6–10)
Note: SD = standard deviation; K = Fatal Crash; A = Incapacitating Crash; B = Non- incapacitating Crash; C = Possible Crash; PDO = property damage only; AADT = annual average daily traffic.
Descriptive Statistics for Categorical Variables (Datasets 1–10)
Empirical Analysis
An NB regression model was estimated for each dataset. The objective was to check whether the stochasticity embedded in the RAD generation process will be consistent for different random seed of data generated from the tool. To achieve this, the parameter estimate from the NB models for each dataset was examined using the revised Wald test statistics created by Hoover et al. ( 28 ) as shown below.
To check the differences in the parameters across the datasets, the t-statistics for all the parameters across all the datasets samples were computed using the computation of the test statistic mentioned above. Dataset 10 (1000 mi) was used as the benchmark to evaluate whether the parameters for other datasets were statistically different relative to this sample dataset. If the parameter test statistic is greater than the 90% t-statistic, it indicates that there is a significant difference between the datasets. On the other hand, if the parameter test statistics is below the 90% t-statistic, there is no significant difference between the datasets, and we can trust that the stochasticity in the RAD tool is consistent across different generations with different random seeds.
Model Results
Table 4 shows the NB parameter estimates for total crashes for all 10 databases. A visual examination reveals that the parameter estimates for each dataset model are relatively close to one another. Comparing the estimated parameters requires having the same variables in all models; to balance variable significance with identical variable sets, we dropped variables that were statistically insignificant based on the 90% significance level in more than six datasets. It was noted that the majority of the categories for lane width and speed limit in datasets 1 and 6 were statistically insignificant at 90% significance level; nonetheless, these variables were left in because the parameter estimate was rather logical in comparison with what has been published in other literature. Overall, the parameter estimates are consistent with what would normally be obtained when using conventional crash data, showing that the RAD tool generates data with reasonable relationships among the variables. However, the main objective of the study was focused on checking for parameter stability across the datasets with different random seeds and varying sizes generated using the tool.
Negative Binomial Estimates for Total Crashes
Note: AADT = annual average daily traffic; AIC = Akaike information criterion; NA = not available.
Parameter estimate.
Standard error.
Variables insignificant at 90% significant level.
Dataset 10 (1000 mi) was used as the population benchmark to evaluate whether each parameter in any model was statistically different from the corresponding one in that dataset. Dataset 10 was used because it had the largest number of observations and was believed to be the most able to produce convincing parameter estimates. As previously mentioned, if the parameter test statistic computed is higher than the 90% t-statistic, the result would indicate significant difference between the corresponding dataset and dataset 10.
Table 5 shows the revised Wald test statistics on the model parameter estimates. The test statistics for segment length across the datasets were lower than the 90% confidence value of 1.65, which could mean that there is no significant difference between the estimated parameters in the corresponding dataset and dataset 10. The test statistics across the datasets for the AADT parameter were also lower than the 90% confidence value of 1.65, indicating that the variation across the different datasets is within a statistically acceptable level. Note that the test statistics for some of the speed limit category coefficients (i.e., 30 mph, 35 mph, and 55 mph) in dataset 6 exceeded the 90% confidence value, which might indicate that the parameter values for those categories could be significantly different relative to dataset 10. The possible reason for this could be that these parameter estimates were actually statistically insignificant on their own.
Revised Wald Test Statistics on Model Parameter Estimates (Relative to Dataset 10)
Note: AADT = annual average daily traffic.
Figures 2 to 4 show the boxplot summary of the test statistics variation for parameter estimates across the datasets. The figures clearly reveal that the range of the test statistics across all the parameters is quite narrow and does not exceed the 90% confidence value of 1.65 from the parameters discussed above. In Figure 2, segment length and shoulder width (6 ft) have a majority of t-statistics falling above the upper quartile (right skewed) but still reasonably within the test of 90% significance which supports the discussion above; the parameter estimate of the other datasets relative to dataset 10 is not statistically different. The plot in Figure 3 shows the t-statistics all falling below the lower quartile (left skewed) and within the test of 90% confidence value. Figure 4 shows the variability among the various category levels of speed limit, with the test statistics all falling below the lower quartile except for speed limit (55 mph). Overall, there was variability across the different variables, but they all reasonably fall within the chosen confidence value.

Test statistics for parameter estimates across datasets for segment length, AADT, and shoulder width.

Test statistics for parameter estimates across datasets for lane width, presence of lighting, and presence of horizontal curve.

Test statistics for parameter estimates across datasets for speed limit.
Discussion
The parameter coefficients estimated using data from the tool are comparable with the estimates reached in previous studies for the variables included in this paper. For example, a study by Gooch et al. ( 29 ) quantified the safety performance of horizontal curves on two-way two-lane rural roads and estimated an AADT coefficient for total crashes as 0.697. Other studies ( 21 , 30 ) estimated 0.622 and 0.656, respectively, these estimates being comparable with the range of AADT estimates from this study. Segment length had estimates of 0.889, 0.801, and 0.459 respectively ( 27 , 29 , 31 ) while the presence of horizontal curve had estimates of 0.053 and 0.050 ( 24 , 29 ). Overall, these estimates compared with ours show that the data from the tool can produce results comparable to those from the traditional data.
Practical Implications
The results from this paper could have significant implications for highway safety research, for the development of information that can be used to make roads safer and crashes less severe. Since the data generation process in the tool is completely known it will allow objective evaluation and validation of various safety analysis methods used to verify various assumptions related to safety performance and in turn lead to providing effective countermeasures to address crashes. The RAD can also be helpful in generating large datasets with consistent conditions in cases where we cannot go back many years because of changes in road characteristics or drivers; this makes it easier to estimate models for unusual and rare events like minor crashes. In addition to this, the tool can help determine sample sizes by determining the quantity of data necessary to produce convincing results.
Conclusions
The current research uses datasets generated from the RAD tool with the objective of assessing the stability of the resulting estimated parameters across the varying datasets and random seeds. Revised Wald test statistics were generated to check whether the variation across the different datasets is within a statistically acceptable level using dataset 10 as the benchmark. The result clearly highlights the stability in various parameter estimates across the datasets.
The resulting stability found across the datasets indicates that the parameter estimates using RAD will be consistent regardless of the miles of segment-related data generated using the tool. Having this knowledge of stability, the dataset from the RAD tool can be used for other possible purposes, such as estimating different prediction models or comparing the performance of various safety analysis methods, and can also help determine the adequate sample size to get convincing results especially for fatal crashes with low realizations. This study contributes to safety research by providing data that can be used by researchers who have no knowledge of the cause–effect structure and who would apply the method they wish to assess. The degree to which they succeed would then be made evident by having the entity which owns the RAD generator compare the estimated parameters to the known relationships embedded in the data. In the future, other statistical models that have been employed by researchers, such as Poisson regression, Poisson-lognormal, random parameters, and multivariate modeling, will be used on the RAD dataset for all roadway facility types including segments and intersections, roadway specifically included in the highway safety manual.
Footnotes
Author Contributions
The authors confirm contribution to the paper as follows: study conception and design: John Ivan, Shanshan Zhao, Naveen Eluru, and Kai Wang; data collection: Oluwaseun Olufowobi, John Ivan; analysis and interpretation of results: Oluwaseun Olufowobi, John Ivan; draft manuscript preparation: John Ivan, Shanshan Zhao, and Kai Wang. All authors reviewed the results and approved the final version of the manuscript.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: The authors would like to gratefully acknowledge FHWA’s sponsorship in this research under project number 693JJ31950017.
The ideas expressed in this work do not necessarily indicate acceptance by FHWA of the findings and conclusions expressed within.
