Abstract
Hotspot identification is an important step in the highway safety management process. Errors in hotspot identification (HSID) may result in an inefficient use of limited resources for safety improvements. The empirical Bayesian (EB) HSID has been widely applied as an effective approach in identifying hotspots. However, there are some limitations with the EB approach. It assumes that the parameter estimates of the safety performance function (SPF) are correct without any uncertainty, and does not consider temporal instability in crashes, which has been reported in recent studies. The Bayesian hierarchical model is an emerging technique that addresses the limitations of the EB method. Thus, the objective of this study is to compare the performance of the standard EB method and the Bayesian hierarchical model in identifying hotspots. Three methods (crash rate, EB, and the Bayesian hierarchical model) were applied to identify risky intersections with different significance levels. Four evaluation tests (site consistency, method consistency, total rank differences, and Poisson mean differences tests) were conducted to assess the performance of these three methods. The testing results suggest that: (1) the Bayesian hierarchical model outperforms the crash rate and the EB methods in most cases, and the Bayesian hierarchical model improves the accuracy of HSID significantly; and (2) hotspots identified with crash rates are generally unreliable. This is significant for roadway agencies and practitioners trying to accurately rank sites in the roadway network to effectively manage safety investments. Roadway agencies and practitioners are encouraged to consider the Bayesian hierarchical model in identifying hotspots.
The identification of crash hotspots (also known as prone sites, sites with promise, or black spots) is one of the most important tasks in the roadway safety management process. Errors in hotspot identification (HSID) can result in inefficient use of limited resources for safety improvements and cause additional loss of lives.
Various methods have been proposed for HSID ( 1 – 3 ), and researchers have been continuously improving the methods ( 4 – 8 ); unfortunately, sites identified by different methods and ranking criteria are not identical ( 9 ) (they are not discussed in detail here because of space limitations). Observed crash counts and crash rates were often used by roadway agencies, but analyses have shown that these two methods cannot account for the regression-to-the-mean (RTM) bias and are not reliable ( 10 , 11 ). Methods based on empirical Bayes (EB) have shown superiority in estimating safety as well as in identifying hotspots ( 8 , 10 ). The standard EB method combines the observed crash counts of one site and the predicted safety of similar sites. The latter is typically derived from a safety performance function (SPF). The EB method was included in the first edition of the Highway Safety Manual ( 12 ) and is widely used in HSID for its ability in correcting the RTM bias and increasing estimation precision. Although many studies have shown that the EB method always performs better than other common HSID methods, it is not without limitations. One critical issue with the EB method is the implementation of the SPF to predict crashes. The SPF is usually modeled using crash data occurring at a similar “reference” pool of sites. In the conventional EB method, the SPF assumes that the number of crashes at each site follows a Poisson distribution, and the number of crashes in each year is independent. In other words, there is no yearly variation in safety at each site when assuming that the traffic volume and other key roadway features remain at the same level. However, this is often not true. With the evolution in vehicles, driving behavior, roadway design standards, and so forth, crash data can be unstable over time ( 13 , 14 ). Without accounting for the instability of crashes, the HSID results estimated through the EB method may be inaccurate under certain situations. In addition, the EB method assumes that the parameters in a fitted SPF are the true estimates, which is also a problematic assumption, especially when the sample size of reference sites is low ( 15 – 17 ).
Recently, Fawcett et al. proposed a novel Bayesian hierarchical model (denoted as Bayesian hierarchical model hereafter) for estimating and predicting roadway safety ( 18 ). The main advantage of the Bayesian hierarchical model is incorporating crash counts from multiple time points, with the counts in more recent years lending more weight to safety estimates than the counts from time points further in the past. The proposed model is able to capture the temporal trend of crashes at a site. A previous study has discussed the structure of the Bayesian hierarchical model and its application to real crash data. It was also noted that the standard EB method is a special case of the Bayesian hierarchical model ( 18 ). However, the HSID results have not been examined to compare the two methods. It is unknown whether or not the novel Bayesian hierarchical model outperforms the commonly used EB method in HSID. As a result, the primary purpose of this paper is to comparatively analyze the performance of the EB method and the Bayesian hierarchical model in HSID. To achieve the objective, this study identifies hotspots of intersections using the two methods, separately. Evaluation tests are conducted to assess the performance of each method.
Methodology
The following two sections introduce the EB method and Bayesian hierarchical model for identifying hotspots, separately. For comparison purposes, the crash rate-based HSID method is also discussed in the first section.
Crash Rate- and EB-Based HSID Methods
As the name implies, the crash rate-based method mainly calculates the rate of crashes at each site, and ranks the sites by their rates. It is usually calculated by dividing the observed crash number by exposure (e.g., vehicle miles of traveling, or VMT). There are a few types of crash rates: target crash rates (considering target collisions only), equivalent crash rate (converting crashes into the same severity level by different weight), and so forth. For the purpose of this study, the total number of crashes and the exposure, as traffic volume traveling through an intersection, are considered (see the “Data Description” section for more details). Thus, the crash rate for a site is calculated as
where
N = number of years in period j; and
Because the crash rate simply relies on the observed number of crashes and exposure at the sites, the randomness of crashes and RTM bias are not well addressed. It has been pointed out that the HSID results using crash rates are not reliable ( 10 , 19 , 20 ). Safety researchers proposed using an EB approach to correct for the RTM bias ( 21 – 23 ). With the EB approach, an estimate of the long-term safety of an entity is obtained from two sources, as described above. Let K be the observed number of crashes, which is Poisson-distributed, and let k be the expected crash count; the EB estimator of k is given as
where
The rate parameter
The NB distribution can be viewed as a mixture of Poisson distributions where the Poisson rate
where
y = response variable;
With the NB model structure, the weight factor w in the EB method is given as ( 22 ):
For the detailed procedures of estimating roadway safety and ranking sites using the EB method, readers are referred to ( 21 , 22 , 37 ).
Bayesian Hierarchical Model
Rather than estimating from a single before period, the Bayesian hierarchical model proposed by Fawcett et al. (
18
) incorporates counts from multiple past periods. This adjusts the variations over the past crash counts by an SPF [note that the researchers used an accident prediction model (APM), and it is essentially the same as a SPF]. For evaluating the model, a discrete time indicator
The SPF is used to overcome the effects of global trend and RTM, and it is assumed that
where
Here,
Evaluation Methods
The crash rate method, the EB method, and the Bayesian hierarchical model are all techniques to identify hotspots. Evaluation tests on their performance are needed as rules to measure their performances in HSID. In most previous studies related to HSID, the percentage of false positives (FP) and percentage of false negatives (FN) have been the only measures. FP is the percentage of safe sites claimed to be unsafe, whereas FN occurs when an unsafe site has failed to be identified. Because the feedback from FP and FN is in binary format, this does not provide insight into the relative performance of HSID methods. For this study, four evaluation tests are implemented, namely: (1) the site consistency test; (2) the method consistency test; (3) the total rank differences test; and (4) the Poisson mean differences test. These tests not only include the consideration of FP and FN, but also take site ranking into account for evaluating the HSID method. These tests follow and reference the evaluation tests proposed by Cheng and Washington ( 19 ).
The site consistency test is to evaluate the ability of HSID methods by measuring the consistent appearance of a high-risk site over a subsequent observation period. The high-risk section includes sites
The site consistency test evaluates the better method, with a higher output indicating a particular method is consistently good over sites.
The method consistency test looks at the consistency of the method over time periods, rather than being site based. For each method
The method consistency test justifies a better method if a method scores higher than other methods, meaning such a method has larger numbers of similar high-risk sites identified throughout the time periods.
The third test, the total rank differences test, compares the summation of ranking differences between high-risk sites identified from base period
Because it is a measure of variations of ranking, a smaller output indicates a better performance, that is
Last, the Poisson mean differences test is used as an evaluation test in this study. It is an extension of the widely used false identification (FI) test. The FI test determines FN and FP as they appear in each HSID method. A site is justified as FN if it is truly hazardous, but mistakenly identified by a method as safe. On the other hand, if a site is truly safe but wrongly assessed as hazardous, it is then considered as FP. Therefore, it is important to have a knowledge of the truly hazardous sites and truly safe sites over time and space. However, the observed crash rate is site specific and time specific. The true Poisson mean (TPM)
The Poisson mean differences test is improved from the FI test by setting the associated TPM difference as the weight. In this test, a relatively small output for
It is worth mentioning that researchers and practitioners have been using various measurements in ranking sites or network screening—for example, observed crash number or rate, EB method, difference between the EB and SPF (also known as potential for safety improvement, PSI), and ratio between the EB and SPF. Some previous studies ( 10 ) have shown that the EB method performs better than others in identifying hotspots. Thus, the EB method was utilized in this study.
Data Description
This study utilized the same data used by Fawcett et al. (
18
), but some filters were applied to exclude outliers and incomplete cases. The dataset includes annual accident counts from 2004 to 2012 at 734 intersections in the city of Halle, Germany. There are two types of observations in the dataset: numerical and binary. The numerical observations
Summary Statistics of Accident Counts and Observations (186 Sites)
Note: min. = minimum; max. = maximum; SD = standard deviation; vpd = vehicles per day.
Results
Crash Rate and EB HSID
Crash rates over the 186 intersections for every 2 years were calculated using the definition documented in the previous section (i.e., Equation 1). A few examples are illustrated in Table 2. As can be seen, Site 74 ranked first in the period of 2004 and 2005, and the crash rate was 3,947.1 per 100,000 vehicles traveling per year. It varied from 1,879.7 to 2,255.6 in the following three periods. It ranked the sixth, tenth, and eighth in 2006 and 2007, 2008 and 2009, 2010 and 2011, respectively. Table 2 also lists the sites with relatively lower crash rates (see the bottom rows).
Example Sites of Crash Rate-Based HSID Results
Note: Number in parentheses indicates ranking. HSID = hotspot identification; no. = number.
Number of total crashes per 100,000 vehicles of traveling per year.
As has been discussed, the SPF is an important part of the EB method. This study modeled the data with an NB distribution, and the modeling results are shown in Table 3. It is worth mentioning that an attempt was made to include all the variables in the model, because the main purpose of the SPF is prediction rather than inference. A few variables (e.g., speed limit 50, logarithm of major volume) are not significant at the 5% significance level, but they have been kept in the model. It is possible that some variables, such as volume and minor volume, are correlated, making the parameter estimates relatively unstable. However, this does not affect the prediction, which is of more interest in the EB method.
SPF of the NB Model
Note: SE = standard error; na = not applicable; AIC = Akaike information criterion; SPF = safety performance function; NB = negative binomial.
Not significant at the 95% level.
Using the SPF, the number of crashes for each site can be predicted, and the EB estimate is then calculated. Taking Site 9 as an example, the predicted number of crashes in 2004 is 8.39. The observed number of crashes in the year is 28. The weight is calculated as
Example Sites of EB-Based and Bayesian Hierarchical HSID Results
Note: Number in parentheses indicates ranking. no. = number; EB = empirical Bayesian; HSID = hotspot identification.
Expected number of crashes (in two years).
Bayesian Hierarchical Model
The Bayesian hierarchical model is coded in R for this study. The SPF model defined in Equation 9 estimates
Regression Results of SPF in the Bayesian Hierarchical Model
Note: SE = standard error; na = not applicable; AIC = Akaike information criterion; SPF = safety performance function.
Not significant at the 95% level.
Moreover, parameters are determined by their priors for MCMC methods. According to Fawcett et al. (
18
), the prior for the time-dependent inflation parameter
Observed Crash Count, SPF Prediction, and MCMC Results of Example Sites
Note: SPF = safety performance function; MCMC = Markov chain Monte Carlo; no. = number; y = observed crash count; µ = SPF Prediction; λ = MCMC results.
Like the crash rate and EB results, Table 4 illustrates the example results for the Bayesian hierarchical method. Taking Site 9 as an example, the predicted number of crashes in 2004 and 2005 by SPF is 16.5. From MCMC, the associated
Evaluation Results
As stated in the “Data Description” section, there are a total of 186 sites. With significance levels of 0.025, 0.050, and 0.075, there are 5, 9, and 14 sites considered as higher risk, respectively, for each HSID method and period. Four tests introduced in methodology are implemented to evaluate the HSID results from the crash rate method, EB method, and Bayesian hierarchical method. Table 7 presents the results of the four tests.
Results of Four Evaluation Tests on Crash Rate, EB, and Bayesian Hierarchical Method
Note: Bold and underlined indicate the highest performance. CR = crash rate; EB = empirical Bayesian; BH = Bayesian hierarchical.
The results of the four evaluation tests under
Conclusions and Discussions
Hotspot identification is the first step for improving traffic safety, and it is an important component of the highway safety management process. Errors in HSID lead to inefficient use of limited resources for safety improvements. Initially, roadway agencies used observed crash numbers and crash rates for identifying sites with promise. But this method does not account for the RTM bias. Safety analysts proposed using statistical models for estimating safety as well as HSID. So far, various models have been extensively applied to identify hotspots. The EB technique has been shown to be an effective approach for identifying hotspots, and it has been widely used in recent decades ( 12 ). However, there are also some limitations with the EB approach. It assumes the parameters in the SPF for predicting the number of crashes are correct without any variation, and also that the safety of a site is temporally independent and stable. Recent studies have revealed that this is not always true ( 13 ). These disadvantages with the EB approach may result in errors in hotspot identification. Fawcett et al. ( 18 ) proposed a novel Bayesian hierarchical model structure for estimating roadway safety. This model is able to capture the temporal trend of crashes at a site, and improves the accuracy of safety estimates. Thus, the study utilized the Bayesian hierarchical model in HSID and compared it to the results using the standard EB method. The purpose was to examine whether the Bayesian hierarchical model can identify hotspots more accurately.
To achieve the objective, this study analyzed crash data from 2004 to 2011 at 186 urban unsignalized intersections in the city of Halle, Germany. A certain number of intersections were identified as hotspots, using three methods: crash rate, EB, and the Bayesian hierarchical model, on a 2-year basis. The identification results, safety estimates, and ranking were examined using four evaluation measurements: (1) a site consistency test; (2) a method consistency test; (3) a total rank differences test; and (4) a Poisson mean differences test. All the test results indicate that the Bayesian hierarchical model performed the best among the three models. The crash rate-based HSID method was overall the worst (which did not come as a surprise), and the EB method was much better than the crash rate method. This is in line with previous studies [e.g., ( 2 , 11 , 38 )]. In short, crash rate-based HSID is not recommended, because it produces unreliable HSID results. Although the differences between the EB and Bayesian hierarchical model in HSID results were not as obvious as the differences between the EB and the crash rate-based model, the Bayesian hierarchical model outperformed the EB approach in almost all the tests and periods. This study found that the Bayesian hierarchical model provides more accurate identification results than the other two methods. Considering the high costs associated with false identification of collision-prone sites, safety analysts and practitioners are encouraged to consider the Bayesian hierarchical model for HSID to reasonably distribute funds for roadway safety improvements.
There are a few limitations to this study. First, a limited number of variables were included in the SPF development because of data availability. It may suffer from the omitted variable problem, as discussed in Wu et al. ( 39 , 40 ). Second, the Bayesian hierarchical model approach for estimating safety and hotspot identification needs MCMC, which may not be feasible for most transportation engineers. Software packages ( 41 ) with an interface for deploying the Bayesian hierarchical analysis are needed for safety analysts and roadway agencies. Finally, all the analyses in this study were based on historical crash data, which is passive. Proactive safety has gained more attention in recent years. Both the Bayesian hierarchical model and the EB method have the ability to predict crashes in future time periods. This is not tested in this study and needs further examination in the future.
Footnotes
Acknowledgements
The authors would like to thank Newcastle Research Data Service for providing the dataset.
Author Contributions
The authors confirm contribution to the paper as follows: study conception and design: XG, LW; data preparation: XG; analysis and interpretation of results: LF, XG, LW, YZ; draft manuscript preparation: LF, XG, LW, YZ. All authors reviewed the results and approved the final version of the manuscript.
The Standing Committee on Highway Safety Performance (ANB25) peer-reviewed this paper (19-03519).
