Abstract
High-frequency short delays in the urban rail transit system result in a cumulative delay effect within the network, which in turn affects the daily operational organization. Most studies focus more on long-term disruption, but there is less research on high-frequency short delays, lacking detailed classification and definition. Based on the 13-week operation data of Nanjing Metro and the detailed division of short-delay frequency, this study constructed four panel-regression models and compared and explained the influencing factors from multiple perspectives. The results show that the negative binomial fixed-effects model has the best goodness of fit. The two-way fixed-effects model increased complexity but does not improve fitting performance compared with the fixed-effects model. There are no significant time fixed effects in the three types of frequency data. New routes, remote stations with low passenger volume, stations with more train trips, and stations with surrounding land use of commercial service facilities have a higher frequency of minor delays, highlighting the need to deploy adequate facilities at these stations.
Keywords
China’s urban rail transit (URT) network has expanded rapidly in recent years, with 55 cities having opened 306 URT lines by the end of 2023 ( 1 ). With such unprecedented URT network growth, more cities face the challenge of managing high volumes of passengers under normal conditions and, more importantly, during unplanned disruptions and delays. Because of the lack of reliable methods to predict URT operation failures, most response measures still adopt the traditional reactive approach of posting many operations and maintenance personnel on standby 24 h a day to implement immediate response measures in the event of URT operational disruption. Other large cities with URT in countries other than China face similar network management problems ( 2 ).
The operation department of Nanjing Metro has made a detailed division of disruption based on the impact degree and failure handling process recorded in the failure report. These URT disruptions are divided into four categories (see Table 1) ( 3 ). Type I major delays significantly affect train operations with low frequency, whereas types II and III are minor delays with a relatively low impact but occurring frequently. Type I disruptions can result in delays >2 min and have a significant impact on operations, requiring adjustments to train operation plans. Type II and type III disruptions cause delays of <2 min and have a relatively minor impact on operations, as trains can utilize buffer times to recover from short delays, thereby avoiding the need for schedule adjustments. Type IV disruptions refer to in-station facility failures and have a minimal impact on train operations. For minor delays of <2 min, there are no recorded delay times in the failure data. Table 1 selects the 2019 data for Nanjing Metro as an example to demonstrate the classification criteria and frequency of disruptions of different types. As the COVID-19 pandemic affected the train operation plan and residents’ travel from 2020, the 2019 data represent a normal situation that has not been affected by the pandemic.
Type and Frequency of Disruptions in Nanjing Metro
URT systems are subject to major and minor delays during operation. At present, major delays have received more attention than minor delays. Most existing disruption research focuses on analyzing delay duration ( 4 ), train rescheduling ( 5 , 6 ), and bus bridging strategies ( 7 ) after long disruptions. Very few studies have focused on the frequency of minor URT disruptions. Although the effects of minor delays are not as severe as major delays, their frequent occurrence can result in a sizeable cumulative number of delays across the network. If minor delays are not addressed promptly and effectively, they can potentially escalate into major delays. The high frequency of minor delays increases the workload for dispatchers. However, most incident frequency models have focused on bus transit ( 8 ) and road-traffic collisions ( 9 ).
Most studies have not defined and subdivided minor delays of <2 min. However, from the statistical data, these minor delays occur frequently. The detailed definition of minor delays by Nanjing Metro can provide a reference for other regions. This study can also serve URT systems in other regions based on the exploration of minor delays. Mastering the frequency characteristics of disruptions helps managers to prevent and handle various types of disruptions in a targeted manner. Given the high frequency of minor delays, they can have a cumulative effect on the network and affect normal operation. Therefore, it is crucial to explore the frequency characteristics and influencing factors of minor delays to implement targeted control.
This study employs panel-regression models to explain the delay data from different aspects to facilitate a comprehensive understanding of the factors influencing the frequency of minor delays. The purpose is to enhance the research on the frequency of minor delays by exploring them through various models and more detailed classifications. This research seeks to supplement the existing analysis perspectives and case studies related to minor delays and to contribute to further analysis in the field of research on minor delays. The research results identified the factors influencing disruption frequency, helping to identify stations with a higher frequency of minor delays, provide a basis for personnel and facility deployment, strengthen control over high-risk stations, and reduce the possibility of minor delays escalating into major delays, thereby enhancing operation safety.
Literature Review
URT Disruption and Delay
Different regions have different standards for dividing and defining delays of varying degrees. Nanjing Metro makes a relatively detailed definition of minor delays ( 3 ). Some studies refer to minor delays as minor disturbances and explore the resilience of the URT system after such disturbances. The disturbances referred to here are events caused by weather or human factors ( 10 ). Researchers have studied the vulnerability of URT networks under different levels of disturbances. However, in this case, even minor disturbances have a significant impact on normal operations, causing partial station failures ( 11 ). The analysis of Singapore’s Mass Rapid Transport system focused on disruptions of >10 min ( 2 ). Some researchers have investigated the incident frequency of Melbourne’s light rail system and divided incidents into types A and B, depending on whether there were casualties involved. Type B incidents are defined as minor incidents with no casualties. The researchers did not separately investigate the influencing factors of the frequency of these two types ( 12 ).
Most studies on URT disruption and delay have focused on the impact of disruption and the treatment measures after delay. However, these studies often concentrate on longer, more impactful disruptions, paying little attention to shorter disruptions that have less impact on operations. In one study, the propagation of initial delay under the effect of node connectivity was studied, and the duration of disruption delay was predicted ( 13 ). Other researchers have analyzed factors such as the time, location, and train type of Toronto subway disruptions, constructing accelerated failure time models to predict the impact of the incident on the duration of delays ( 14 ), but the frequency of these delays has not been studied. The delay propagation rules have also been analyzed based on passenger congestion handling time and dissipation time ( 15 ). Some researchers have proposed a quantitative method for identifying passengers affected by disruptions based on the arrival time of passengers, which helps evaluate the scale and origin distribution of affected passengers ( 16 ). With regard to bus bridging studies after disruptions, some studies have focused on dispatching buses from a predesignated set of locations, but demand forecasts were rarely made, leading to overestimates of passenger demand ( 17 ). Some researchers have taken the Shanghai Metro as an example and constructed a system to simulate the spatiotemporal evolution of the network’s evacuation demand under disruption ( 18 ). A resilience-based optimization model has been proposed to maximize the network’s global average efficiency to optimize the restoration sequence scheme after disruption ( 19 ).
Although these studies have conducted in-depth discussions on issues such as delay prediction, delay duration, delay propagation, and evacuation after disruption, they have paid less attention to the frequency characteristics of delays, and different regions have different standards for dividing delays. More thorough research is still needed on the frequency of delays and influencing factors, especially on minor delays. Although there is limited research on the frequency of minor URT delays, the modeling of URT delay frequency can benefit immensely from the rich literature on road traffic collisions concerning the types of statistical models, model specifications, and explanatory variables used.
Existing Research Gap
It is clear from the literature review that existing studies on URT disruption frequency need to be more comprehensive. Very little research has been conducted to build models that analyze the frequency of minor URT delays and to understand the key influencing factors. The characteristics of time series and cross-sections can be explored using panel regression. The discussion in the introduction and literature review leads to the following observations. (1) Most studies have focused on investigating the characteristics of major delays in URT systems and developing effective emergency response methods, with relatively less attention given to short but high-frequency minor delays. There is a lack of research on URT disruption frequency, limiting our understanding of the influencing factors of the delay frequency in URT systems. (2) The definition of minor delays varies in existing research, and different regions have different standards. A more detailed and comprehensive definition is urgently needed. (3) The current research on URT disruption mainly utilizes failure data to explore the failure effect after the delay, with the time correlation of delays and the use of panel data rarely being considered.
Methods
Data
The data used in the present study are from the Nanjing Metro system. The network operation of the Nanjing Metro system makes it a representative case in China, with potential applicability to large-scale subway systems in other cities and countries. The network of this system in 2019 had 10 lines, including 159 stations, 13 of which were transfer stations, as shown in Figure 1.

Nanjing Metro network.
The study period spanned from March 4 to June 2, 2019, before the COVID-19 pandemic to provide a more accurate representation of normal conditions. During this period, the records of type II and type III disruptions were selected from the train failure data for 2019.
Figure 2 illustrates the characteristics of the data pertaining to two types of delays. For type II delays during the 13-week study period, 127 stations experienced fewer than 10 delays, 22 stations had between 10 and 20 delays, and only 10 stations experienced more than 20 delays. For type III delays during the study period, 101 stations experienced fewer than 10 delays, 33 stations had between 10 and 20 delays, 12 stations had between 20 and 30 delays, and only 13 stations had more than 30 delays. Therefore, most stations had fewer than 20 type II delays and fewer than 30 type III delays.

Frequency characteristics of the stations during the study period (March 4–June 2, 2019).
The passenger volume and train operating schedule from March 4 to June 2, 2019, were obtained. The years of operation of each line and the land-use types of each station were determined. Weather conditions in 2019 were queried from the historical records of the meteorological bureau. Since the frequency of minor disruption and delays per day per station were relatively small and affected commuting on weekdays and weekends, this research elected to study the frequency of station delays per week to explore the characteristics of their influencing factors. So, the time series is organized in weekly time units, whereas cross-sectional data are defined on a station-by-station basis.
In summary, six indicators related to URT operations were identified. These indicators were selected based on previous studies and the actual situation in Nanjing. Passengers and trains are the two most critical components of the URT system and directly affect the frequency of delays. Since they vary across different stations and weeks, they were chosen as the main explanatory variables.
Temperature changes have an indirect effect on disruption frequency, as cooler temperatures may mean precipitation or fog, which, in turn, affects operations in ground sections and makes trains more susceptible to skidding and other issues. The years of operation represent the level of aging in facilities. Newer lines may experience more disruptions during the initial run-in period, whereas older lines may be prone to failures because of equipment aging. Betweenness centrality reflects the position of stations in the network. Stations with higher betweenness centrality are typically transfer stations located at the center of the network, whereas those with lower betweenness centrality are usually on remote lines. Location information is an important indicator for each individual station, and similarly, land-use type is also a key factor for each individual station. Variables that vary only over individual stations or time were selected as secondary explanatory variables.
The study period consisted of 13 weeks. Each week’s average temperature was calculated, and the numbers of weekly inbound passengers and passing trains at each station were calculated. The definitions of the variables are shown in Table 2.
Definition of Variables
Betweenness centrality for a node refers to the number of times that node acts as the shortest bridge between two other nodes. The formula for betweenness centrality is shown in Equation 1 ( 20 ):
where
The three dependent variables investigated were the frequency of type II delay events, the frequency of type III delay events, and the frequency of type II and III delay events. Because type II and type III minor delays are very similar, this study aimed to investigate whether they have similar occurrence characteristics and whether they can be combined. Table 3 presents the descriptive statistics of each explanatory variable.
Descriptive Statistics of Variables
Note: SD = standard deviation; Min. = minimum; Max. = maximum.
Data Modeling
The data used in this study are panel data, including the delay frequency of all stations in the network over a 13-week period, the stations representing different sections, and the weeks representing different time series. Specifically, the data of all the stations in the network during the same week are cross-sectional data, whereas the data of different weeks at the same station are time-series data. These panel data capture both the time dimension and the cross-sectional dimension, as the two main explanatory variables vary over time and across different stations. Additionally, there are other secondary explanatory variables that vary either over time or across individuals, reflecting the effects of time and individuals. Given the nature of the data, panel regression analysis is deemed suitable for this study. Various tests were performed, and the most suitable model was determined by adding or subtracting explanatory variables appropriately.
Regression models based on panel data generally have three forms: the pooled model, the fixed-effects model, and the random-effects model ( 21 ). In the pooled model, when differences in time and individuals are not considered, the coefficients and intercept terms of the regression equation are constant, and ordinary least squares regression is used to estimate the parameters. The fixed-effects model considers the effects of time and cross-section to be fixed and the error terms to be related to the explanatory variables. The random-effects model considers the effects to be random and the error terms to be unrelated to the explanatory variables.
The F-test ( 22 ) is used to decide whether to select the pooled model or the fixed-effects model, and the Hausman test ( 23 , 24 ) is used to determine whether to build the random-effects model or the fixed-effects model. To explore the characteristics of the frequency of type II events, type III events, and type II and type III events separately, pooled models, fixed-effects models, and random-effects models were constructed for these three kinds of data.
The null hypothesis of the F-test is that all the individual-effect terms are equal to 0, which indicates that the fixed effects do not exist and the pooled model should be chosen. If
The Hausman test is used to explore the endogeneity of models. The null hypothesis is that the individual-effect terms are not correlated with the covariates. If
Table 4 presents the values and probabilities of the F-test and Hausman test of some models. The F-test and Hausman test suggested rejecting the null hypothesis and selecting the fixed-effects model. Therefore, this study conducted the remaining analysis using the fixed-effects models. Different explanatory variables were selected to construct panel-regression models. The individual fixed-effects model, time fixed-effects model, two-way fixed-effects model, and negative binomial fixed-effects model were examined, and the characteristics of the two kinds of short delays and their combination are discussed. The goodness of fit of the four models was evaluated using the Akaike information criterion (AIC) and Bayesian information criterion (BIC) measures. Figure 3 presents the research methods.
Values and Probabilities of F-Test and Hausman Test in Different Models
Note: Value and p denote the test results and significance.

Research methods.
Considering that some explanatory variables only vary among individuals (i.e., stations) or over time, exploring the fixed effects of individuals or time is necessary. This study elected to estimate the fixed-effects model considering individual fixed effects and time fixed effects separately and the two-way fixed-effects model considering both individual fixed effects and time fixed effects, then compare these models with the negative binomial fixed-effects model.
Individual Fixed-Effects Model
The individual fixed-effects model has different intercept terms for different individuals (time series). However, if the number of individuals is large, many dummy variables need to be included, causing too much loss of degree of freedom. Therefore, the intercept term can be divided into two parts,
According to this model, except for the explanatory variables, the effects of all other variables (not included in the regression model or not observable) affecting the dependent variable only change across individuals, representing the individual fixed effects. Thus, this model cannot include independent variables that do not change over time. The formula for the model is presented in Equation 2:
where
Time Fixed-Effects Model
The time fixed-effects model has different intercepts for different cross-sections. If the intercepts of the model are significantly different for different cross-sections but the intercepts are the same for different time series, then the time fixed-effects model should be used. Cross-sections will be included in the model as dummy variables. Time fixed-effects models cannot include independent variables that do not vary among individuals. The formula for the model is presented in Equation 3:
where
Two-Way Fixed-Effects Model
The two-way fixed-effects model is adopted to consider individual and time fixed effects concurrently. It is a model with different intercepts for different cross-sections and different time series. If it is confirmed that the intercepts of different cross-sections and different time series are significantly different, then the two-way fixed-effects model should be used. The formula for the model is shown in Equation 4:
where
Negative Binomial Fixed-Effects Model
When the distribution of the dependent variable is over-dispersed and the data mean and variance are not equal, the Poisson regression model is not recommended for use because it assumes that the mean is equal to the variance. If the Poisson regression is used, its regression coefficients are still consistent and unbiased, but its standard errors are underestimated.
In this case, the negative binomial regression model fits the data better. The negative binomial distribution can be used as an alternative to the Poisson distribution because it is a long-tail distribution, with emphasis given to the value of the extra parameter, which allows variance to be greater than the mean ( 25 ). At the same time, we still consider the fixed effects of individuals. The negative binomial mass function can be written according to Equation 5 ( 26 ):
where
Model Testing
In the process of constructing the fixed-effects models, the intragroup correlation is considered through robust estimation. Multiple tests ( 27 ) are used to determine the absence of heteroscedasticity and multicollinearity in the model.
For parameter estimation, the model fitting is optimized by maximizing the likelihood function. However, the model with the highest fit does not mean optimal, as some complex models have overfitting problems. In this case, criteria that can balance the accuracy and complexity of the model are required. The goodness of fit is assessed using AIC and BIC. The formulae for the AIC and BIC are shown in Equation 7 ( 28 ):
where
Results and Discussion
The F-values of all models suggest rejecting the null hypothesis, which means there are likely fixed effects, and accordingly, the pooled model should not be used. The Hausman test of each model rejected the null hypothesis, indicating that the fixed-effects model was more suitable than the random-effects model. Thus, the focus of the analysis is the fixed-effects model. The individual fixed-effects model, time fixed-effects model, two-way fixed-effects model, and negative binomial fixed-effects model are constructed for comparative analysis.
The comparison of the fitting performance of these four models is shown in Table 5. According to the results of AIC and BIC, the individual fixed-effects model has an excellent fitting performance on type II delays followed by type III delays. For the time fixed-effects model, the fitting performance of type II delay is better. The time fixed-effects model is slightly inferior to the individual fixed-effects model but reveals more effects of individual-related explanatory variables. For the two-way fixed-effects model, the model test results are close to the individual fixed-effects model. The fitting performance of the negative binomial fixed-effects model is the best. At the same time, the fitting of each model to the type II data is better than that of the type III and combined models. Moreover, the fitting performance indicates that although the two-way fixed-effects model has increased the capture of time fixed effects compared with the individual fixed-effects model, the goodness of fit has not improved but rather has increased the complexity of the model. Although the two-way fixed-effects model can also fit the data, it is not as concise as the individual fixed-effects model. The time dummy variables in the time fixed-effects model are not significant, indicating that the three types of data have little difference between each week and there is no obvious fixed time effect. Therefore, the time fixed-effects model is relatively poor.
Comparison of Model Testing
Note: AIC = Akaike information criterion; BIC = Bayesian information criterion.
Detailed model estimation results are shown in Tables 6–9. In Table 6, the main explanatory variables, weekly passenger volume and weekly number of train trips, are both significant for type II and type III delays. Although the weekly number of trains passing through a station is positively correlated with the frequency of type II and type III delays, the weekly passenger volume is negatively correlated. This is because emergency management personnel and facilities are more fully deployed at stations with larger passenger flows to mitigate the impact of short delays. The number of train dispatchers and passenger duty officers is proportional to the station’s size. The number of maintenance workers and security inspectors increases with the number of ticket booths and automatic gate machines. Moreover, disruptions at large passenger volume stations are likely to cause delays of >2 min, which are type I disruptions. Thus, these stations with higher weekly passenger volume are prone to longer-lasting and more influential type I disruptions, whereas those with minor type II and III disruptions are relatively less frequent. The disruptions occurring at stations with low passenger volume generally cause short delays. Therefore, more minor delay records exist. In addition, stations with low weekly passenger volume are mostly the remote stations of suburban lines that have a shorter operation year, and the minor delay occurs more frequently in the run-in period. For type II and type III data, the weekly passenger volume and number of trains are insignificant. In none of the models was the weekly mean temperature significant.
Estimation Results of the Individual Fixed-Effects Model
Note: Significance confidence levels: ***99%, **95%, *90%.
Estimation Results of the Time Fixed-Effects Model
Note: Significance confidence levels: ***99%, **95%, *90%. C = commercial service facilities land; CR = mixed residential and commercial land; G = green space and square land; P = public service facilities land; R = residential land; T = transportation facilities land; W = logistics and warehousing land.
Estimation Results of the Two-Way Fixed-Effects Model
Note: Significance confidence levels: ***99%, *90%.
Estimation Results of the Negative Binomial Fixed-Effects model
Note: Significance confidence levels: ***99%.
In Table 7, the time fixed-effects model includes the additional covariates of years of operation and betweenness centrality, which vary across individuals while removing the temperature covariate, which varies only over time. The week number and land-use type were included in the model in the form of dummy variables. The first week and commercial service facilities land were taken as reference terms.
For time dummy variables, subsequent weeks are not significantly different from the first week. This indicates that the time fixed effect of the data is not significant—that is, there is no significant difference between each week. From the results of the model testing, the fitting performance of the time fixed-effects model is slightly poor. Therefore, the time fixed effects can be ignored, but at the same time, attention can continue to be paid to the estimation results on other explanatory variables. The two main explanatory variables are still significant for type II data. However, only the weekly train trip is significant for type III and combined data. This indicates that after considering the time fixed effect, the weekly passenger volume is not statistically significant for these two types of data. Compared with land-use type C, other land-use types are significantly different and negatively correlated with type II delays. This shows that the commercial service facilities land is more prone to type II delays.
The years of operation and betweenness centrality are significant in type II and type III models. The regression coefficients are all negative, indicating that with the increase in the years of operation and betweenness centrality, the frequency of type II and type III delays decreases. This suggests that recent lines are more prone to experiencing more frequent minor delays than older lines. Remote stations are more prone to experiencing frequent minor delays, whereas stations located in the center of the network structure are less likely to experience such delays. For the Nanjing Metro, the new lines are mostly suburban lines, whereas the old lines are in the city center. There are two main reasons for the high frequency of short delays on newly built lines and remote stations: (1) the new line has a run-in period in the initial operation, which is more likely to cause more failures and delays; and (2) suburban lines have longer station spacing and longer operation lengths than urban lines. Moreover, unlike urban lines, which are primarily underground, suburban lines have more elevated sections and ground sections, which are more susceptible to precipitation and foreign objects, so they are more prone to skid and catenary failures, causing more short delays. This result is consistent with a similar finding in a recent study ( 29 ). To minimize the impact of external factors such as precipitation and foreign objects on suburban lines, operational management departments can make improvements from two aspects. First, more physical protection measures, such as longer rain shelters and higher protective nets, can be aded in areas prone to minor delays. Second, the “wet track” mode ( 30 ) can be adopted in ground sections during precipitation. In this mode, the braking force of the trains is increased, and the acceleration and deceleration rates are reduced to prevent trains from slipping in rainy and snowy weather, ensuring safe operation.
Table 8 shows the two-way fixed-effects model. The two main explanatory variables are shown to be significant in type II and type III models but not for the combined data. The week dummy variables are insignificant, suggesting that the time fixed effects are statistically insignificant. In the two-way fixed-effects model, the model includes both time fixed effects and individual fixed effects. If an explanatory variable changes only with time or only with individuals, it will be perfectly correlated with fixed effects, so it cannot be identified and estimated in the two-way fixed-effects model. Therefore, this model only includes the two main explanatory variables and does not include the four secondary explanatory variables. This means that the explanatory variables in the model are limited, as the focus is on capturing the variation over time and across individuals.
Table 9 shows the negative binomial fixed-effects model. The features of type III data are not suitable for negative binomial models, so only the other two types of data will be explored. Although the model can estimate the sum data, multiple explanatory variables are not significant and do not show good fitting performance. The two main explanatory variables are significant for type II. The frequency of type II delays increases with the weekly number of train trips and decreases with weekly passenger volume. These results are consistent with the other fixed-effects model. The weekly mean temperature is statistically insignificant in this model. This indicates that the change in weekly mean temperature has no significant effect on the frequency of short delays. The influence of weather factors cannot be reflected solely through average temperature.
Conclusions
By comparing the results of type II, type III, and combined frequency, it was found that the significance of some essential explanatory variables after combination could not pass the t-test. Therefore, it can be concluded that although type II and type III are both minor delays, their influencing factors and fixed effects are different, and they cannot be combined in regression modeling. These two types of delays should be analyzed separately. Different types of delays have different effects and significance of covariates under different models. The weekly passenger volume and weekly train trips, the main explanatory variables, are consistently shown to be statistically significant in different models. The greater the weekly passenger volume at a station, the lower the type II delay frequency. In contrast, the higher the number of weekly train trips passing through a station, the greater the occurrence of type II delays. In addition, the years of operation, betweenness centrality, and land-use types are also shown to be significant factors of minor delay frequency. In general, recent lines, remote stations, and stations situated in commercial service land use have a higher frequency of minor delays. The negative binomial fixed-effects model exhibited the best goodness of fit among the four models. The time fixed-effects model revealed more information about the explanatory variables, but there are no significant time fixed effects of frequency. The individual and two-way fixed-effects models achieved similar fitting performances, but the two-way fixed-effects model increased complexity. These models can well explain the direction and degree of the influencing factors of minor delay frequency and can be used to provide insights related to URT delay in other cities.
These findings can guide the operational management department, with regard to personnel deployment, train configuration, facility upgrade, and maintenance support. At present, the number of train dispatchers and passenger duty officers configured in the station is proportional to the size of the station. Recent lines and remote stations are the weak points that are easily ignored in emergency management. The research results show that in addition to deploying enough emergency personnel and facilities at stations with large passenger volumes and of critical importance (as measured by betweenness centrality), more complete emergency measures are also needed for newly built lines and less critical stations to mitigate the effects of frequent minor delays—for example, increasing the number of train dispatchers and their working hours at stations with frequent minor delays. Once a minor delay occurs, the dispatchers can promptly and effectively arrange for train drivers to adjust the train speed, avoiding the development of minor delays into longer type I delays, which can cause passenger congestion and changes in subsequent train schedules. Additionally, reducing the number of trains passing through low-traffic stations at the end of the line could be considered. By storing trains in sidings, the number of trains can be minimized while still meeting passenger demand. This helps to reduce the frequency of minor delays. Meanwhile, shelter facilities such as longer rain shelters and higher protective nets could be added to the ground section of the suburban line to reduce the impact of precipitation and foreign objects on operation and reduce minor delays caused by catenary failure and train slippage. The “wet track” mode could be adopted more frequently in the ground section of the suburban line during precipitation.
This study improves understanding of the characteristics and influencing factors of minor delay occurrence in URT networks, which can inform decision-making by management departments to devise and implement preventive and reactive operational strategies. The limitation of this study is that the limited explanatory variables considered in the model may affect the fitting performance, which will be the focus of subsequent research. If the failure data and operational data affected by the COVID-19 pandemic and safety protocols from 2020 to 2022 can be obtained in future research, comparative research on minor delay frequency between normal and pandemic conditions can be carried out. In the following research phase, we will consider more indirect influencing factors besides the direct influencing factors of frequency and include the service level provided by the station, train maintenance schedules, driver experience, and weather conditions in the explanatory variables. When adding new explanatory variables, multiple tests will be conducted to ensure that the new explanatory variable does not have issues such as autocorrelation or multicollinearity with the original explanatory variable. In addition, it is necessary to explore how adding more explanatory variables can improve the goodness of fit and avoid situations where increasing model complexity does not improve fitting performance. Meanwhile, frequency panel data using day as the unit of time will be considered to explore the differences between holidays and workdays.
Footnotes
Acknowledgements
The authors thank the Nanjing Metro Group Co. Ltd. for providing the relevant data. We also acknowledge the support of Nanjing Rail Transit Smart Transportation Research Station.
Author Contributions
The authors confirm contribution to the paper as follows: study conception and design: Jinyi Chen, Amer Shalaby; data collection: Tiezhu Li, Hui Liu, Yiyong Bo, Fei Lin, Jinyi Chen; analysis and interpretation of results: Jinyi Chen, Amer Shalaby; draft manuscript preparation: Jinyi Chen, Amer Shalaby, Tiezhu Li. All authors reviewed the results and approved the final version of the manuscript.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This research has been supported by Jiangsu Rail Transit Industry Development Collaborative Innovation Base Open Fund (No. GCXC2103 and No. GCXC2104); and the China Scholarship Council (No. 202006090134).
Data Accessibility Statement
The data that support the findings of this study are available from Nanjing Metro Group Co. Ltd., but restrictions apply to the availability of these data, which were used under license for the current study, and so are not publicly available. Data are, however, available from the authors on reasonable request and with permission of Nanjing Metro Group Co. Ltd.
