Abstract
With the advancement of automotive energy-saving and emission-reduction technologies, constructing multi-parameter driving cycles that accurately represent real-world driving behavior has become crucial for vehicle development, performance evaluation, and energy management strategy optimization. These cycles typically incorporate key parameters such as vehicle velocity, acceleration, and road grade, providing a comprehensive description of both the vehicle operational states and the external environment. However, constructing such multi-parameter driving cycles faces challenges including high-dimensional state spaces, complex constraints, and high computational costs. To address these challenges, this study proposes a dimensionality reduction-based multi-step optimization method designed to significantly reduce modeling complexity and improve generation efficiency. The method achieves effective dimensionality reduction through a joint low-dimensional representation of state transitions and conditional constraints, establishes a comprehensive evaluation framework integrating distributional consistency and statistical features, and employs a multi-step genetic algorithm for efficient iterative optimization of cycle sequences. Validation results demonstrate that the proposed method can reliably generate highly representative multi-parameter driving cycles, exhibiting substantial advantages over conventional baseline methods as regards modeling efficiency, sequence generation speed, and accuracy. The framework shows good extensibility and adaptability, offering an effective solution for constructing multi-parameter driving cycles across different vehicle types and driving scenarios.
Keywords
Introduction
Background and Motivation
Energy conservation and emission reduction are crucial societal issues ( 1 ). Governments worldwide are tightening emission regulations and driving technological progress toward sustainability ( 2 ). Concurrently, automakers are enhancing data collection and analysis systems to optimize vehicle design ( 3 ). The driving cycle is a key representation describing the operational state of a vehicle within its specific traffic and road environment ( 4 ). As a core benchmark for vehicle development, performance testing, and regulatory certification, the fidelity and representativeness of driving cycles directly determine the evaluation accuracy of vehicle fuel consumption, emission levels, and various performance indicators ( 5 , 6 ). Consequently, driving cycles constitute a foundational supporting technology driving technological advancement and product iteration in the automotive industry ( 7 , 8 ).
However, with changes in policies, technologies, road infrastructure, and usage scenarios, traditional standard driving cycles can no longer accurately reflect the transient performance characteristics of actual driving ( 9 , 10 ). Existing standard cycles, such as the Federal Test Procedure 75 (FTP-75), New European Driving Cycle (NEDC), Worldwide Harmonized Light Vehicles Test Cycle (WLTC), China Light-duty Vehicle Test Cycle (CLTC), and China Heavy-duty Commercial Vehicle Test Cycle (CHTC), consist solely of velocity sequences (11–13). Extensive empirical studies have revealed significant and systematic deviations between laboratory certification results based on fixed velocity–time profiles and the actual on-road performance of vehicles ( 14 , 15 ). The root cause of this “certification gap” lies in the simplified modeling paradigm relied on by traditional standard cycles, which centers on a single velocity sequence and is inadequate for capturing and reflecting the complexity of multi-parameter coupling in real-world driving scenarios ( 16 ). To more authentically replicate actual driving conditions and provide a more reliable data foundation for engineering applications such as vehicle development and energy management strategy validation, developing multi-parameter driving cycles has become an industry trend ( 17 ).
Acceleration, a key parameter reflecting vehicle dynamics and response capability, offers critical insights into vehicle performance, aiding optimization across design, manufacturing, and usage stages. Additionally, road grade significantly affects vehicle speed variation, acceleration requirements, and energy consumption; overlooking this factor can distort energy and emissions assessments ( 18 ). Therefore, incorporating parameters like acceleration and road grade to construct a multi-parameter driving cycle model can enhance the accuracy of vehicle state representation and real-world scenario simulation ( 19 ).
Despite this, research on multi-parameter driving cycles remains limited, primarily owing to the lack of a simple yet effective method to capture the relationships between velocity, acceleration, and road grade. This complexity hinders high-precision cycle design. Therefore, developing an efficient method to model these relationships is essential for reducing design complexity and improving the precision and efficiency of cycle development. To address this challenge, this study aims to systematically construct a complete methodological framework for multi-parameter driving cycle design. This framework spans from the dimensionality-reduction modeling of the multi-parameter state transition model, to the formulation of driving cycle evaluation metrics, then to the efficient generation of driving cycle random sequences with controlled start and end points, and further establishes a multi-step optimization framework. The core contributions of this study will be systematically validated using real-world vehicle operational data to ensure the reliability of the proposed method as regards effectiveness and generalizability.
Brief Review on Existing Multi-Parameter Driving Cycle Design Methods
Existing driving cycle design methods can be classified based on underlying logic into micro-trip (MT) and discrete-state methods. MT methods segment vehicle speed data into MTs according to specific rules and classify them using techniques like cluster analysis. A small number of MTs are randomly sampled to create the driving cycle ( 20 , 21 ). Standard driving cycles such as WLTC, CLTC, and CHTC are based on MT methods. While widely used in urban driving cycle development, MT methods have limitations: first, owing to the urban speed characteristics that favor MT segmentation, MT is more suited for urban driving cycles ( 22 ), but less effective for high-speed, long-distance driving ( 23 ); second, they primarily address velocity profiles and perform poorly with road grade profiles, potentially causing numerical discontinuities ( 24 ); third, these method rely on clustering algorithms, which require pre-defined classification numbers, introducing significant subjectivity into the results ( 25 ).
To more comprehensively capture the dynamic characteristics of vehicle operation, discrete-state methods are advantageous. These methods discretize vehicle data into states per second ( 26 ), and use Markov chain (MC) models to generate transient driving cycles ( 27 ). However, as parameters increase, the state space grows exponentially, and random state transitions make it challenging to obtain satisfactory driving cycles within a limited period. To enhance accuracy, various optimization algorithms have been applied. Cui et al. ( 28 ) combined MC with simulated annealing (SA) to better match velocity and acceleration patterns with real-world driving characteristics. They also proposed a combined method based on min–max ant colony optimization (ACO), effectively reducing the fuel consumption discrepancy between the real world and estimates ( 29 ). Yang et al. ( 30 ) applied Metropolis Hastings Sampling to random transitions of velocity and acceleration states and combined it with genetic algorithm (GA) to propose a fuel consumption-driven, high-accuracy driving cycle construction method. Existing studies have compared the applicability of optimization algorithms such as SA, ACO, and GA in engineering applications. The SA algorithm is straightforward to implement and excels in scenarios involving local optimization or when a good initial solution is available; however, it converges slowly in large and irregular solution spaces ( 31 ). ACO is well-suited for discrete optimization problems such as path planning, yet it suffers from high computational cost and a prolonged initial exploration phase in high-dimensional and complex spaces ( 32 ). GA has become a mainstream choice in driving cycle optimization owing to its robustness and efficiency in handling large-scale, multi-objective, and complex optimization problems ( 33 , 34 ), though its performance depends to some extent on the careful design of crossover, mutation, and other operators.
In the design of multi-parameter driving cycles incorporating velocity, acceleration, and road grade, Zhang et al. ( 35 , 36 ) adopted a joint encoding approach and combined MC with GA for driving cycle construction, demonstrating higher efficiency compared with purely random sampling. Addressing the challenge that low-frequency GPS data (sampling intervals of 3–5 s), widely available in practice, are insufficient to capture transient driving features, Shi et al. ( 37 ) proposed a solution that integrates low-frequency data with prior information from a typical driving cycle database. Using a GA-based MC evolutionary algorithm, they successfully reconstructed driving cycles with high-frequency characteristics. Furthermore, Shi et al. ( 38 ) introduced an evolutionary driving cycle design method based on forward and inverse MC models combined with GA, which shows significant advantages in the efficiency of generating random state sequences. However, these methods encounter challenges related to state space expansion when applied to multi-parameter joint encoding, resulting in increased computational complexity and constrained optimization efficiency.
This study focuses on the multi-parameter combined representation of velocity, acceleration, and road grade. Three primary approaches for incorporating road grade into driving cycle design can be summarized. The first method involves the joint encoding of velocity, acceleration, and road grade at the same time step ( 35 ). Figure 1 illustrates a schematic of the state transition probability matrix constructed after jointly encoding velocity, acceleration, and road grade (denoted as VAG). In this matrix, both the rows and columns correspond to the encoded VAG states, with the total number of states denoted as N. Each row of the matrix represents the current VAG state at time t, while each column represents the possible VAG state at the next time step t+1. The value in each cell of the matrix indicates the probability of transitioning from the state in the corresponding row to the state in the corresponding column, with the sum of probabilities in each row equal to 1. For clarity, this model is referred to as the original three-parameter VAG model. Although this approach features a compact structure, reducing the step size leads to a sharp increase in the number of states, substantially increasing computational cost and optimization difficulty.

Schematic diagram of the VAG (joint encoding velocity, acceleration, and road grade) state transition matrix.
The second method, based on the first, jointly encodes velocity and road grade, while omitting acceleration ( 39 ). For clarity, this model is denoted as the VG model. Although this simplification streamlines the driving cycle design process, it undermines the accuracy of vehicle condition representation. Acceleration is a crucial parameter that captures phases of vehicle acceleration and deceleration, directly affecting fuel economy and energy consumption. Its exclusion may lead to inaccurate estimations of fuel usage and energy efficiency.
The third method, based on average velocity, matches road grade with typical driving cycles ( 40 ). The procedure consists of three steps: 1) Construct a typical driving cycle without road grade; 2) Divide the velocity sequence of this cycle into segments, each labeled by its average velocity. Similarly, segment the observed road grade data and label each interval with its corresponding average observed velocity; 3) Match segments from the typical driving cycle to road grade intervals by selecting the closest average velocity, then concatenate the matched road grade segments to form a complete sequence. However, this matching approach has inherent limitations that may compromise reliability. The average velocity labels for the driving cycle and the road grade are derived from independent datasets with no direct relationship. Although average velocities may appear similar, the actual velocity profiles can differ significantly, leading to inconsistent driving behaviors and reducing the validity of the matched road grade.
In addition, current evaluations of driving cycle performance mainly rely on descriptive statistics, with limited attention paid to the distributional consistency between driving cycles and actual driving data ( 41 ). In distribution analysis, kernel density estimation (KDE) offers a smoother representation of data distributions compared with histogram-based binning methods ( 30 ). However, traditional KDE approaches, such as Gaussian KDE (GKDE), assume normal distribution, which contradicts the data-driven intent of non-parametric estimation. These methods also face limitations, such as inadequate local adaptability and boundary bias. In contrast, the diffusion-based KDE (DKDE) addresses these shortcomings by removing the dependence on normal distribution for bandwidth selection and effectively mitigating boundary bias. As a result, DKDE more accurately captures the underlying structure of data distributions ( 42 ). Unlike classical methods such as the chi-squared test, which are sensitive to binning choices and rely on theoretical distributions, DKDE offers greater flexibility. It requires no pre-specified distributional form and directly estimates the true density, making it particularly suitable for assessing distributional consistency between driving cycles and real-world data, especially in the presence of multi-modal, skewed, or unknown distributions ( 43 ).
In summary, two major challenges persist in current multi-parameter driving cycle design: first, accurately representing state information leads to an exponential increase in state space dimensionality, making it difficult to balance design efficiency and precision; second, omitting critical features such as acceleration undermines the representativeness of driving cycles. Moreover, to improve the evaluation framework, it is essential to incorporate distributional consistency metrics alongside descriptive statistics, thereby establishing a more comprehensive assessment system for driving cycle performance.
Contribution
In response to existing challenges in multi-parameter driving cycle construction—such as high model dimensionality, incomplete evaluation metrics, and low optimization efficiency—this study proposes a dimensionality reduction-based multi-step optimization method, aiming to efficiently construct multi-parameter driving cycles that better align with real-world vehicle operation. The approach first constructs conditional constraint matrices to establish low-dimensional state transition rules, decomposing the high-dimensional joint state into multiple subspaces. This significantly reduces modeling complexity while preserving physical relationships among variables. Subsequently, a comprehensive evaluation framework integrating distribution consistency is developed by combining DKDE with traditional descriptive statistical features. Finally, an initial population is efficiently generated using a forward-inverse segment replacement strategy, followed by iterative optimization via a multi-step GA to obtain highly representative multi-parameter driving cycles. The brief workflow is shown in Figure 2. The main contributions are summarized as follows:
Proposing a MC modeling method based on conditional constraint dimensionality reduction. This approach maintains parameter correlations while reducing the model complexity by four orders of magnitude, effectively overcoming the high dimensionality and computational complexity associated with conventional models.
Developing distribution consistency evaluation metrics based on DKDE. Integrated with traditional descriptive statistics, these metrics forms a multi-dimensional and multi-scale comprehensive evaluation framework, enabling a holistic assessment of the representativeness of multi-parameter driving cycles.
Designing a multi-step genetic optimization strategy. By decomposing the overall optimization process into sequential phases, this strategy achieves superior optimization efficiency and solution quality compared with traditional single-step optimization methods, making it particularly suitable for generating multi-parameter driving cycles.

Brief workflow.
Markov Chain Dimensionality Reduction Modeling
Markov Chain and Multi-Parameter Joint State Definition
A random process
To accurately characterize driving behavior in real-road environments while reducing state-space dimensionality, this study defines two key state variables:
1)
2)
The joint state at time n can thus be expressed as:
Probability Decomposition of Joint State Transitions
In the joint Markov model, the transition probability from the current state to the next state can be expressed as:
To reduce computational complexity, the state transition process is divided into two sequential steps according to Equation 3:
1) Determine the next VA state
2) Determine the next G state
Note that both
where:
I: Precise feasible set of VA—all possible values of the next VA state under the current VA and G conditions;
J: VA transition set—possible values of the next VA state based solely on the VA transition law;
K: Environmental constraint set—possible values of the next VA state constrained only by the current G state.
From fundamental probability principles, it is obvious that:
Similarly, for
where:
W: Precise feasible set of G—all possible values of the next G state given the current VA, current G, and the next VA state;
X: VA constraint set—possible values of the next G state constrained only by the current VA state;
Y: G transition set—possible values of the next G state based solely on the G transition law;
Z: VA intent constraint set—possible values of the next G state constrained only by the next VA state.
It is also obvious that:
Equations 7 and 12 indicate that any feasible state transition must satisfy all relevant constraints. This insight suggests that by approximating the high-dimensional sets I and W with the intersections of low-dimensional sets
The set
Based on set K, the conditional constraint matrix
The updated transition probability for the VA state is defined as:
where
Similarly, the set
Furthermore, based on sets X and Z, two conditional constraint matrices are defined:
The set of allowed G target states is defined as:
The updated transition probability for the G state is given by:
where
In summary, this study proposes a novel VA-G modeling framework that dynamically couples the VA and G subsystems via multiple low-dimensional conditional constraint matrices, enabling efficient and interpretable state evolution modeling, as illustrated in Figure 3. Departing from conventional high-dimensional joint probability representations, the model employs initial state transition matrices P and Q, along with three conditional constraint matrices

VA-G model structure.
State inference occurs in two sequential stages: first, the next VA state is inferred from the current VA and G states using P and
Verification of Model Validity
To evaluate the similarity between the VA-G model and the original three-parameter VAG model, the proportion of states in the random state sequence generated by the VA-G model that adhere to the state transition relationships of the VAG model is calculated. A higher value of this proportion indicates that the VA-G model more accurately captures the state transition characteristics of the original model. The analysis is conducted using real vehicle operation data of varying scales, all obtained from the vehicle networking big data platform of a major Chinese automotive manufacturer. The experimental procedure consisted of the following steps:
1) State Discretization
First, state discretization is performed on velocity, acceleration, and road grade. Based on practical conditions, the velocity range is set to [0, 36] m/s, the acceleration range to [−2, 2] m/s2, and the road grade range to [−10, 10] %. To analyze the impact of the discretization step size on state representation accuracy, a comparative experiment is conducted using three discretization granularities: coarse, medium, and fine. The specific step sizes and the corresponding number of states for each granularity are presented in Table 1.
State Discretization Step Sizes and State Counts
Note: V = velocity; A = acceleration; G = road grade.
To validate the accuracy of different discretization granularities, a segment of real-world vehicle operating data is selected for comparative analysis. Figure 4 illustrates the reconstruction of the actual operating data by the state sequences under the three granularities. It can be observed that the medium and coarse granularities exhibit significant deviations from the actual data, while the fine granularity demonstrates superior representational capability. Furthermore, Figure 5 presents the root mean square error (RMSE) between the state sequences and the actual data for the three granularities. The fine granularity yields the smallest RMSE, indicating the lowest deviation from the actual data. Therefore, the fine discretization granularity is selected for state discretization in this study.
2) Joint Encoding
VA joint encoding is performed as:
where
VAG joint encoding is defined as:
where
3) Statistics and Modeling
Data are extracted based on mileage thresholds. State transition matrices for VA, G, and VAG are constructed according to the state encodings (
4) Similarity Verification
The proportion of states in the VA-G-generated sequences that satisfy the transition relations of the original VAG model is computed using:
where

Comparison of state sequences under three discretization granularities with real-world operating data.

Comparison of root mean square error (RMSE) between state sequences and actual data under three discretization granularities.
For each mileage threshold, 500 random VA and G state sequences, each lasting 1,800 s, are generated using the corresponding VA-G model. The pseudocode for generating these sequences is presented in Algorithm 1. The VA and G sequences are converted into a joint three-dimensional state encoding format while respecting temporal continuity. States in each sequence meeting the original model’s transition conditions (i.e., with transition probabilities greater than zero) are counted, and the proportion r is calculated using Equation 21.
Figure 6 shows the proportion of states in sequences generated by the VA-G model that adhere to the transition relations of the original model. The results demonstrate that the proportion reaches nearly 0.9 at 20,000 km, exceeds 0.98 at 1,000,000 km and approaching 1. These findings confirm that, under large-scale data conditions, the VA-G model accurately captures the transition characteristics of the original VAG model. Thus, with sufficient data support, the VA-G model effectively represents variable interactions and state transition dynamics.

Proportion of states that satisfy the original model’s transition relations.
Comparative Analysis of Modeling Efficiency
To systematically evaluate the modeling efficiency of the proposed VA-G model, we compare it with two existing models: the VAG model ( 35 ) and the VG model ( 39 ). The real-world vehicle operational data were sourced from a vehicle networking big data platform of a major Chinese automotive manufacturer. The dataset comprises fully loaded highway driving data from 50 express delivery box trucks operating nationwide, with a total accumulated mileage of 1.32 × 106 km. All experiments are performed on a computer equipped with a 13th Gen Intel® Core™ i9-13900KF 3.00 GHz CPU and 32 GB of RAM. Modeling and computations are carried out in MATLAB® 2024b, and the same configuration is used for all subsequent tests.
Under a unified fine-grained discrete encoding scheme, Table 2 compares the three models as regards state-space size, model complexity, memory usage, and computational time required for modeling. The VA-G model decomposes the system state into two low-dimensional subspaces with dimensions of 14,801 and 101, respectively. Compared with the VAG model and the VG model, the proposed VA-G model reduces the state space size to only 1.0% of that of the VAG model and 40.9% of that of the VG model. With reference to model complexity, the VA-G model achieves a reduction of four orders of magnitude relative to the VAG model and one order of magnitude relative to the VG model. With reference to memory usage, the VA-G model occupies only 22.8 % of the memory required by the VAG model and 60.5 % of that required by the VG model. As for computational time for modeling, the VA-G model requires only 0.6% of the time needed by the VAG model and 12.2% of that required by the VG model. In summary, the VA-G model significantly reduces computational resource demands and substantially improves modeling feasibility and efficiency.
Comparison of State-Space Size, Complexity, Memory Usage, and Computational Time
Note: VAG = joint encoding velocity, acceleration, and road grade; VG = joint encoding velocity and road grade; VA-G = dimensionality reduction model proposed in this study.
Multi-Step Genetic Optimization Strategy
This section presents a multi-step genetic optimization strategy for the multi-parameter driving cycles. First, a comprehensive evaluation index system is established to ensure precise assessment of the driving cycles. Then, the VA-G model is utilized to generate random state sequences, which form the initial population. Finally, multi-step optimization is performed through the GA to obtain high-precision representative multi-parameter driving cycles that accurately reflect real-world driving characteristics.
Evaluation Criteria Construction
The accuracy of a driving cycle in representing real-world driving behavior relies on a comprehensive evaluation system. This section introduces distribution consistency evaluation metrics, extending traditional methods, to better assess the alignment between the driving cycle and actual driving data. Specifically, the one- and two-dimensional probability density distributions of velocity, acceleration, and road grade are estimated for both the database and the driving cycle using DKDE. Consistency evaluation metrics are then constructed based on these distributions.
Given independent samples
where
is a Gaussian probability density function with location
With initial condition:
where
Neumann boundary condition:
By leveraging the relationship between GKDE and the Fourier heat equation, finding the optimal bandwidth for Equation 22 is equivalent to identifying the optimal diffusion time t for the diffusion process governed by Equation 24. Considering the initial and Neumann boundary conditions, the analytical solution for Equation 24 on the finite domain
where the kernel function is:
Unlike conventional methods, DKDE determines the optimal bandwidth directly from the data using an improved plug-in method ( 42 ), without assuming normality. This makes the approach fully non-parametric and data-driven.
The optimal bandwidth is given by:
where
For
Using DKDE, the probability density distributions for both the database and the driving cycle are obtained. The distributions are discretized, and based on the concept of the overlapping coefficient ( 46 ), their consistency is quantified by calculating the overlapping area between the two probability density distributions ( 47 ):
where
Distribution Consistency Evaluation Metrics
In addition to distribution consistency, this study incorporates road grade statistics alongside velocity and acceleration to form a comprehensive set of evaluation metrics. Descriptive statistics for both the database and the driving cycle are computed as shown in Table 4. The relative deviation between them is calculated as:
where
Statistical Characteristic Parameters
Random State Sequence Generation
The standard driving cycles typically begin and end with an idle state. Accordingly, in this study, the random state sequences are also defined to start and terminate under idle conditions. Generating sequences that satisfy this requirement through direct random simulation is computationally inefficient. To address this, we extend the forward and inverse modeling approach proposed in previous work (
38
) to the current context by constructing both forward and inverse state transition matrices for VA and G, along with the conditional constraint matrices

Segment replacement between forward and inverse state sequences.
To evaluate the efficiency of the proposed VA-G model in generating random state sequences, a comparative analysis is conducted against the existing VAG model ( 35 ) and VG model ( 39 ). The tests are based on real-world vehicle operation data covering 1.32 × 106km. Each random sequence is set to a duration of 1,800 s, and multiple repeated trials are performed to ensure statistical reliability.
As summarized in Table 5, the VA-G model demonstrates a significant advantage in generation efficiency, with an average time of only 2.2 s per sequence—54.3 times faster than the pure random method used in the VAG model. This improvement is attributed to the decoupling of state variables and the incorporation of low-dimensional constraint mechanisms, which substantially reduce computational complexity while maintaining sequence diversity. Therefore, the proposed method is well suited for efficiently generating initial populations, providing high-quality initialization support for subsequent optimization problems.
Comparison of State Sequence Generation Efficiency
Note: VAG = joint encoding velocity, acceleration, and road grade; VG = joint encoding velocity and road grade; VA-G = dimensionality reduction model proposed in this study.
Multi-step Optimization Algorithm
Under time-limited conditions, random simulation methods struggle to efficiently generate representative driving cycles. The GA, with its strong stochastic search and exploration capabilities, is well-suited to address such challenges ( 48 ). This study adopts an iterative GA-based evolution mechanism to progressively optimize representative driving cycles that better reflect real-world driving characteristics. The overall optimization process is illustrated in Figure 8.

Optimization process.
Owing to the multi-step optimization strategy, each step requires a specifically designed objective function to ensure effectiveness and directionality throughout the process. Drawing on the satisfaction criterion model presented in Shi et al. ( 37 ), this study integrates the established evaluation metrics to formulate the objective function, thereby transforming the multi-metric optimization problem into a single-objective convergence criterion function.
where
where
A constraint function is then designed based on the number of states
where
Building on previous work ( 38 ), this study adopts the roulette wheel selection strategy and introduces an elitism retention mechanism. The two-point crossover operator is adopted to ensure compliance with MC state transition constraints, as shown in Figure 9. The mutation operator utilizes a strategy based on inverse state sequences, consistent with the principle in Figure 7, guaranteeing both mutational effectiveness and adherence to MC state transition constraints.

Crossover operation diagram.
Results and Discussion
This section aims to validate the practicality and generalizability of the proposed method by conducting representative driving cycle generation analysis using data from different vehicle types and driving scenarios, followed by performance comparisons with traditional joint modeling approaches. Heavy-duty trucks are vital in China’s freight industry owing to their high payload capacity. Currently, approximately 94% of these trucks are powered by high-power diesel engines, resulting in notable fuel consumption and emissions ( 3 ). Therefore, this study focuses on this category of commercial vehicles.
Data Description
All data used in this study were sourced from a vehicle networking big data platform of a major Chinese automotive manufacturer, which enables real-time collection and monitoring of driving data. This approach allows for the acquisition of sufficient, authentic driving data covering various vehicle types and operating scenarios without the need for extensive additional testing. Two databases are employed in this study: the first consists of fully-loaded expressway operation data from 50 delivery box trucks operating nationwide, as previously mentioned; the second comprises newly added one-month continuous urban plain road operation data from 3,059 fully loaded heavy-duty tractor trucks across China.
All data were sampled at 1 Hz and contain multiple dimensions of information, including timestamp, vehicle model, geographic coordinates, velocity, acceleration, road grade, road type, terrain attribute, and gross vehicle weight. To ensure data quality, systematic preprocessing of the raw data was performed. First, sensor faults and extreme outliers (e.g., signal loss in tunnels) were identified and removed based on the 3-sigma criterion. For short-term missing data, linear interpolation between the preceding and subsequent normal breakpoints was applied. Finally, a second-order Butterworth filter was used to smooth signals such as velocity, and zero-phase forward-backward filtering was employed to eliminate phase delay. An illustrative data format is provided in Table 6. A short segment of actual vehicle operation data is shown in Appendix A to illustrate the pattern of these data more clearly.
Data Format Example
Note: Road type: 1 = Highway; 2 = Arterial road; 3 = Urban road; Terrain type: 1 =Plain; 2 =Hilly; 3 = Mountainous.
After the above processing, the final valid datasets are as follows: the total mileage of the delivery box truck data amounts to 1.32 × 106 km, and that of the heavy-duty tractor truck data reaches 9.76 × 105 km. Detailed information is presented in Table 7. The two datasets differ significantly as regards vehicle type, driving environment, and operational characteristics, thereby effectively enabling the examination of the proposed method’s adaptability and robustness across different vehicle models and driving scenarios.
Database Details
Figure 10 illustrates the distribution of representative vehicle operating trajectories. Figure 10a displays the trajectory distribution in the core area of the eastern coastal economic belt—a region characterized by a dense road network and high trajectory coverage, representing one of the most economically developed and highly urbanized areas in China. Figure 10b provides a color-coded legend that distinguishes terrain features and road attributes using different colors, offering an intuitive basis for analyzing vehicle operating characteristics under varying geographical environments. This nationwide coverage of trajectories across typical geographical regions ensures the representativeness of the research data relating to terrain representation and road conditions.

Vehicle driving trajectories: (a) eastern coastal areas, (b) trajectory color-coded legend. (Color online only.)
VA-G models are constructed based on databases D1 and D2, respectively. Taking D1 as an example, the forward and inverse state transition matrices of VA and G are shown in Figure 11. It can be observed that the state transitions of VA are highly concentrated near the diagonal, indicating that vehicle states tend to transition between adjacent values, reflecting the smooth driving behavior typical in real-world conditions. Similarly, the state transitions of G also show strong diagonal concentration, suggesting that changes in road grade are generally small in magnitude, which aligns with the continuity assumption in real road environments. The structures of the conditional constraint matrices

Forward and inverse state transition matrices of velocity–acceleration (VA) and road grade (G) for database D1: (a) VA forward, (b) VA inverse, (c) G forward, (d) G inverse.

Conditional constraint matrices of velocity–acceleration (VA) and road grade (G) for database D1: (a) C1, (b) C2, (c) C3.
Goodness-of-Fit Test
To demonstrate the advantages of DKDE over conventional KDE methods, a comparative analysis is conducted using Database D1, evaluating DKDE against both GKDE and Epanechnikov KDE (EKDE) in estimating one- and two-dimensional distributions of velocity, acceleration, and road grade. This comparison aims to validate the effectiveness of DKDE in modeling driving cycle data.
The bandwidths for GKDE and EKDE are calculated using Silverman’s rule of thumb ( 49 ):
where σ is the standard deviation of the sample, T is the sample size, and
Bandwidth for One- and Two-dimensional Distributions
Note: DKDE = diffusion-based kernel density estimator; GKDE = Gaussian kernel density estimation; EKDE = Epanechnikov kernel density estimation.
The kernel density estimates for one-dimensional distributions are presented in Figure 13. The comparison shows that in the peak regions of the distributions, the DKDE estimate aligns the closest match to the shape of the actual data. In contrast, both GKDE and EKDE display noticeable deviations around the peaks of the acceleration and road-grade distributions. For the velocity distribution, the conventional methods exhibit boundary bias near zero velocity, whereas DKDE demonstrates superior fitting accuracy in this region owing to its boundary adaptive capability.

Kernel density estimation of one-dimensional distribution: (a) velocity, (b) acceleration, (c) road grade.
To quantitatively evaluate the fitting performance of each method, the Kolmogorov–Smirnov (KS) statistic and the RMSE are adopted as goodness-of-fit metrics. Values of the KS statistic and RMSE closer to zero indicate a higher agreement between the density estimate and the true data distribution. The KS statistic measures the maximum deviation between two cumulative distribution functions, while RMSE quantifies the average difference over the entire domain. Together, these two metrics provide a comprehensive assessment of goodness-of-fit, ranging from the worst-case to the average-case scenario. As illustrated in Figure 14, KS and RMSE consistently indicate that the DKDE method outperforms all other compared approaches, confirming its effectiveness in density estimation for driving cycle data. Furthermore, the two-dimensional kernel density estimates are presented in Figure 15. Compared with GKDE and EKDE, DKDE also provides a closer approximation to the true joint distributions, further validating its superior performance in high-dimensional density estimation.

Goodness-of-fit test for one-dimensional distribution: (a) Kolmogorov–Smirnov (KS), (b) root mean square error (RMSE).

Kernel density estimation of two-dimensional joint distribution: (a) velocity–acceleration, (b) velocity–road grade, (c) acceleration–road grade.
Representative Driving Cycles Construction
The parameters used in the VA and G optimization processes are summarized in Table 9. To facilitate analysis and reproducibility, the maximum number of iterations is adopted as the stopping condition for the algorithm. For each driving scenario, 50 multi-parameter representative driving cycles are generated through repeated experiments to evaluate the effectiveness and generalizability of the algorithm.
Optimization Process Parameters
Note: VA = velocity–acceleration; G = road grade.
Concerning statistical characteristic evaluation metrics, as shown in Figures 16 and 17, among the 50 highway driving cycles of Database D1, three exceed the 10% threshold for VA-related metrics, while the remaining 47 meet the predefined requirements. The majority of these show relative deviations around 5%. For the 50 urban driving cycles of Database D2, two exceed the 10% threshold for G-related metrics, with the other 48 complying with the criteria; most relative errors of the statistical parameters in these cases are also within approximately 5%. Overall, the success rate for obtaining well-representative driving cycles exceeds 90%. Considering distribution consistency, illustrated in Figure 18, all driving cycles meet the satisfactory threshold, with one-dimensional consistency exceeding 90% and two-dimensional consistency above 80%.

Velocity–acceleration-related statistical characteristic evaluation metrics for 50 driving cycles.

Road grade-related statistical characteristic evaluation metrics for 50 driving cycles.

Distribution consistency evaluation metrics for 50 driving cycles.
It is noteworthy that among all distribution metrics, the one-dimensional consistency of the velocity distribution is comparatively lower, while the two-dimensional consistency of the velocity–road grade joint distribution is the lowest across all two-dimensional metrics. This outcome is largely attributable to the complexity of the state space: the velocity state occupies a larger state space than the acceleration and road grade states, and the joint velocity–road grade state possesses the highest dimensionality and largest state space among all two-dimensional joint states. Consequently, optimizing distribution consistency for this joint state is more challenging, leading to its relatively weaker performance.
To further examine the evolution of distribution consistency during optimization, the velocity distribution and the velocity–road grade joint distribution from the D1 highway cycle are analyzed as representative cases. As shown in Figures 19 and 20, both distributions gradually converge toward their corresponding target database distributions as the number of iterations increases. This convergent behavior clearly demonstrates the effectiveness of the optimization algorithm and confirms the guiding role of the evaluation index system in steering the optimization process.

Evolution of the velocity distribution of the driving cycle with increasing iteration count.

Evolution of the joint velocity–road grade distribution of the driving cycle with increasing iteration count.
The convergence curves of the VA and G objective functions are shown in Figures 21 and 22, respectively. Both objective functions converge stably, with highway driving cycles achieving better convergence performance than urban driving cycles, attributable to their smaller deviations in distribution consistency. Figures 23 and 24 present the multi-parameter representative driving cycles for databases D1 and D2. These cycles exhibit smooth profiles and clearly discernible features, accurately reflecting real-world driving conditions. The results demonstrate that the proposed method generalizes well across different vehicle types and driving scenarios. The representative driving cycles capture authentic driving characteristics effectively, confirming the robustness and practical applicability of the method.

Convergence curves of the velocity–acceleration (VA) objective function.

Convergence curves of the road grade (G) objective function.

Highway driving cycle of D1.

Urban driving cycle of D2.
Performance Comparison with Conventional Joint Modeling Methods
This study focuses on the joint representation of velocity, acceleration, and road grade parameters, and systematically compares the performance of different joint modeling approaches in constructing driving cycles. Given that traditional purely random MC methods exhibit clear limitations in both efficiency and accuracy, and the existing literature ( 30 , 35 , 38 ) has fully demonstrated this, the relevant comparisons will not be repeated in this section.
Using database D1, the VAG model ( 35 ) and the VG model ( 39 ) are selected as benchmarks. To ensure fairness and reproducibility, all methods adopt the same evaluation metrics, encoding resolution, evolutionary strategies, and optimization parameters. Since both VAG and VG are well-established methods, 15 highway driving cycles are generated for each model for performance validation. The average absolute relative deviation across all evaluation metrics is used as the primary criterion for comprehensively comparing the performance of the VAG, VG, and the proposed VA-G models.
The representative driving cycles generated by the three models are labeled as DC_VAG, DC_VG, and DC_VA-G, respectively. Figure 25 compares the relative deviations in statistical characteristics among the driving cycles produced by each model, while Figure 26 presents the comparison of distribution consistency metrics. Specifically, DC_VAG shows maximum and average relative deviations of 12.3% and 7.3%, respectively; DC_VG exhibits 13.7% and 7.5%. In contrast, the proposed VA-G model yields DC_VA-G with significantly lower deviations of only 6.6% and 4.1%, respectively. Moreover, all distribution consistency values of DC_VA-G exceed 85%, consistently outperforming both reference methods. These results indicate that the VA-G model generates driving cycles that better align with real-world driving data, not only achieving smaller deviations in statistical features but also demonstrating superior consistency in distribution morphology.

Comparison of relative deviations in statistical characteristic parameters.

Comparison of distribution consistency metrics.
The driving cycles DC_VAG and DC_VG are shown in Figures 27 and 28, respectively. It can be observed that the road grade sequence in DC_VAG exhibits abrupt variations. DC_VG shows delayed velocity increase and frequent fluctuations during initial acceleration, failing to accurately capture real highway driving behavior. Its road grade sequence also displays noticeable oscillations. Moreover, since the VG model does not explicitly represent vehicle acceleration, the acceleration profile derived from numerical differentiation of velocity lacks physical constraints, leading to significant spurious fluctuations.

Driving cycle DC_VAG.

Driving cycle DC_VG.
The average optimization computation times for DC_VAG, DC_VG, and DC_VA-G are 1492.5 s, 1120.3 s, and 926.7 s, respectively. Within these totals, the average durations attributed to DKDE are 127.9 s, 66.7 s, and 62.2 s, each accounting for less than 10% of its corresponding total optimization time. This indicates that the computational overhead introduced by DKDE during actual execution is minor. Although the proposed VA-G model uses a multi-step optimization strategy, it still requires the shortest total computation time among the three. Compared with DC_VAG and DC_VG, DC_VA-G improves generation efficiency by 37.9% and 17.3%, respectively. If DC_VAG and DC_VG were further constrained to fully satisfy all evaluation metric thresholds, their computation times would increase substantially, making the efficiency advantage of the proposed method even more pronounced.
In summary, the proposed VA-G model effectively overcomes the inherent limitations of existing methods, such as excessively large state spaces or insufficient representation of key features. The proposed approach enables the generation of more accurate representative driving cycles, providing a more reliable data foundation for developing and validating vehicle energy management strategies.
Conclusion
To reduce the complexity involved in designing high-precision multi-parameter driving cycles, this study proposes a dimensionality reduction-based multi-step optimization method. By constructing a low-dimensional VA-G model, the high-dimensional joint representation is decomposed into tractable subproblems. Combined with a comprehensive evaluation system and a multi-step optimization strategy, the method efficiently optimizes driving cycles incorporating velocity, acceleration, and road grade. Validation using data from different driving scenarios and comparisons with existing methods lead to the following conclusions:
1) The VA-G model effectively incorporates conditional constraints among velocity, acceleration, and road grade. It reduces state-space dimensionality while preserving physical relationships, significantly improving computational efficiency and providing a scalable framework for integrating additional environmental and vehicle parameters.
2) The VA-G model outperforms existing VAG and VG models in modeling efficiency, state sequence generation speed, and driving cycle accuracy, demonstrating enhanced practicality and generalizability.
3) The multi-step optimization framework exhibits strong adaptability and reliability across various vehicle types and driving scenarios. The generated driving cycles faithfully represent real-world driving behavior and environmental conditions, offering a systematic methodology for constructing complex multi-parameter driving cycles.
Although the method yields satisfactory results, it currently depends on large volumes of high-quality real driving data, limiting its application in data-scarce contexts. Future work will focus on: 1) Extend validation to more vehicle types (e.g., electric and special-purpose vehicles) and complex operating scenarios (e.g., extreme climates and challenging terrains) to assess generality and robustness. 2) Enrich the model by incorporating additional parameters such as road curvature, gear-shift operations, and ambient temperature, enabling higher-dimensional and more representative driving-cycle modeling. 3) Develop adaptive cycle-generation strategies that accommodate diverse constraints across vehicle types and driving contexts, enhancing practical applicability and broader utility.
Supplemental Material
sj-docx-1-trr-10.1177_03611981261437053 – Supplemental material for Dimensionality Reduction-Based Multi-Step Optimization Method for Multi-Parameter Driving Cycles
Supplemental material, sj-docx-1-trr-10.1177_03611981261437053 for Dimensionality Reduction-Based Multi-Step Optimization Method for Multi-Parameter Driving Cycles by Suhua Jia, Shuming Shi, Nan Lin, Laiyi Zhang, Boan Chen and Bingjian Yue in Transportation Research Record
Footnotes
Notation
ACO Ant colony optimization
CHTC China heavy-duty commercial vehicle test cycle
CLTC China Light-Duty Vehicle Test Cycle
DC_VAG Representative driving cycles generated by the VAG model
DC_VA-G Representative driving cycles generated by the VA-G model
DC_VG Representative driving cycles generated by the VG model
DKDE Diffusion-based kernel density estimator
EKDE Epanechnikov KDE
FTP-75 Federal Test Procedure 75
G Road grade
GA Genetic algorithms
GKDE Gaussian KDE
KDE Kernel density estimation
KS Kolmogorov–Smirnov statistic
MC Markov chain
MT Micro-trips
NEDC New European Driving Cycle
RMSE Root mean square error
SA Simulated annealing
VA Velocity–acceleration
VA-G Dimensionality reduction model proposed in this study
VAG Joint encoding velocity, acceleration, and road grade
VG Joint encoding velocity and road grade
WLTC Worldwide Harmonized Light Vehicles Test Cycle
Author Contributions
The authors confirm contribution to the paper as follows: study conception and design: Suhua Jia, Shuming Shi; validation: Nan Lin; data pre-processing: Laiyi Zhang; investigation: Boan Chen; data collection: Bingjian Yue; analysis and interpretation of results: Suhua Jia; draft manuscript preparation: Suhua Jia, Shuming Shi. All authors reviewed the results and approved the final version of the manuscript.
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work is supported by the National Natural Science Foundation of China (Grant number 51975242).
Supplemental Material
Supplemental material for this article is available online.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
