Abstract
This study developed an indicator prediction model for steel processes adapting to dynamic operating condition deviations. By combining just-in-time learning, ensemble learning techniques, and target similarity extraction, the model improves prediction accuracy and robustness. Validated on industrial rolling and sintering data, the model achieves up to 18% improvement over traditional models. Utilizing multi-dimensional parameters, the 6% error margin hit rates for yield and tensile strength reached 86% and 99% in rolling data, while the Al2O3 hit rate reached 96% in sintering. Notably, for boundary samples (top/bottom 25% quantiles), the model achieved a 56% relative improvement in hit rate. Furthermore, the model significantly enhances prediction robustness under dynamic operating deviations, reducing the accuracy fluctuation across time segments from 12% to 5%.
Keywords
Introduction
The stage or end-point target variables involved in each process of steel production are called the key indicators of steel process, such as steel rolling quality, 1 sintering quality, 2 molten steel temperature, 3 and so on. The real-time information of key indicators has a significant impact on improving product quality, reducing energy consumption, and improving production efficiency. However, the traditional detection and measurement methods often have the problems of time lag and high cost. Therefore, the prediction technology of key indicators of a steel process flow based on production process data came into being.4,5 It can predict key indicators in advance, provide support for production decisions, and reduce maintenance and use costs.
The core challenges of the prediction technology for key indicators of an iron and steel process are focused on two aspects: data and models. One is to comprehensively and accurately collect the relevant influencing factors of indicators, and the other is to establish the model relationship between relevant factors and key indicators. At present, most researchers study and improve the model on the basis of high-quality data, and realise the accurate prediction of key indicators with the help of machine learning technology. 6 Furthermore, an increasing number of studies have extensively explored the complex physical properties, micro-structural evolution, physical metallurgy, and accurate classification of various steel products using advanced analytical and machine learning technologies.7–14 Li et al. 15 propose a novel dynamic time feature expanding and extracting framework integrated with recurrent neural network regression to predict sinter quality in industrial processes. Shi et al. 4 propose a data-driven furnace heat prediction and feedback model integrating genetic algorithm-enhanced machine learning with blast furnace metallurgical processes, achieving great accuracy for hot metal temperature and chemical heat prediction, while improving furnace temperature stability through industrial implementation.
In a broader context, advanced data-driven predictive frameworks have achieved widespread success in related metal processing and manufacturing fields, such as tube bending prediction and thermal error compensation.16,17 Furthermore, foundational research continues to deeply explore complex physical metallurgical phenomena, such as the thermal stability of metastable austenite in steels. 18 A large number of previous studies have directly established global indicators prediction models for historical datasets, and achieved good results in the experimental state. However, if it is applied to the practice of steel production, the model effect is likely to decline or drop sharply. The reason is that most of the production process parameters change with time in the industrial scenario, that is, the operating condition deviation. 19 Operating condition deviation will lead to unpredictable changes in the data distribution after the training model. The production process parameter data presented when the model is applied to production does not correspond to the data distribution provided during the training. At the same time, the mapping relationship between the production process parameters and the target indicators will also change, resulting in a decline in the prediction ability of the model. 20
To address the unpredictability of data distribution shifts, the concept of just-in-time learning (JITL) is widely employed. Originally established in the late 1990s as a powerful ‘model-on-demand’ framework for anticipating time-series behaviour and facilitating process control,21,22 JITL effectively mitigates the continuous drift of operating conditions by dynamically extracting a sample set similar to the current state from historical data for local modelling. 23 Furthermore, ensemble learning (EL) 24 has been introduced into these industrial scenarios to further enhance prediction accuracy by aggregating multiple models.
Recent advancements have further expanded JITL’s capabilities in industrial soft sensing by integrating it with robust representation learning and advanced autoencoders to handle complex, non-stationary data distributions.25,26 Building on these modern frameworks, Shao et al. 27 introduced JITL into industrial soft sensing and optimised its performance by addressing sample selection bias and non-Gaussian distribution modelling challenges, achieving improved prediction accuracy. Dong et al. 28 combined JITL and EL technology, proposed a soft sensing method based on integrated local modelling to realise the just-in-time prediction of mechanical properties of hot rolled strips, and verified the effect of the model in the experiment. To effectively track shifting trends in both stable and transitional conditions, Song et al. 29 introduce JITL technology and propose a data-driven progressive search and parallel model framework, which rapidly predicts temperature during transition states in continuous annealing processes. By leveraging an ensemble framework, de Matos et al. 3 achieved higher prediction accuracy for molten steel temperature than individual models, effectively modelling both linear and non-linear thermal loss patterns caused by dynamic variables. We designate this method as HLRFR and employ it in subsequent comparative experiments.
On the basis of existing research, this study combines EL with JITL to propose an ensemble learning-based just-in-time learning (EL-JITL) method. To improve it, inspired by recent optimisations of JITL similarity criteria in sintering processes, 30 we introduce the concept of target similarity. It predicts the target through the pseudo-label prediction model, limits the similar sample set of EL-JITL, and improves the prediction ability of the model for boundary samples. The main novelty and rationale of this work lie in overcoming the ‘feature–target decoupling’ phenomenon that inherently plagues standard global models and traditional distance-based JITL under dynamic operating condition deviations. Unlike existing models that rely solely on historical feature–space proximity, which can be highly misleading under industrial noise, our proposed framework introduces a pseudo-label-driven target similarity mechanism. By explicitly coupling this constraint with a multi-dimensional ensemble stacking strategy, the TS-EL-JITL method fundamentally reduces epistemic uncertainty and avoids over-smoothing. Ultimately, we propose the TS-EL-JITL method to further improve the robustness and accuracy of key indicator predictions in the steel process.
For clarity and ease of reference, all principal mathematical symbols and acronyms used throughout this paper are summarised in the Notation in the Appendix.
Methodology
The dynamic change and offset of production conditions cause the data distribution to change rapidly, which greatly restricts the effect of the traditional prediction model. The introduction of JITL for dynamic local modelling and prediction of current samples to be tested in the production process can alleviate the above problems to a certain extent. However, JITL uses a single similarity strategy to select the set of modelling samples from historical data, which has limitations for reflecting the characteristics of the current sample to be tested, and further limits the effect of the local prediction model.
In order to deal with the above technical difficulties, this study proposes an EL-JITL prediction method that integrates target similarity, named TS-EL-JITL. The method is divided into three stages: JITL, integrating target similarity and EL. Regarding the nomenclature, TS-EL-JITL is named following a descriptive modifier convention rather than strict chronological execution: JITL serves as the core foundational framework, while TS and EL are descriptive prefixes representing the two novel enhancement mechanisms applied to it. Firstly, a variety of similarity strategies are integrated to make up for the limitations of a single similarity strategy in capturing similarity and enhance the generalisation ability of the model. Secondly, the target similarity calculation module is introduced to evaluate the similarity between the target indicator and the historical data of the sample to be tested, and further optimise the set of similar samples, which plays a great role in improving the prediction effect of boundary samples. Finally, the EL idea is used to improve the overall accuracy and robustness of the model. This method not only improves the accuracy of prediction, but also enhances the adaptability of the model to the rapid change of production conditions. The overall model framework is shown in Figure 1.

Framework of TS-EL-JITL. TS-EL-JITL: target-similarity ensemble just-in-time learning.
Next, the three stages will be described in detail. The historical data samples are represented by
JITL stage
In the JITL stage, traditional methods usually only rely on a single similarity measurement strategy to screen similar samples. However, in iron and steel production, the evaluation of sample similarity involves multiple dimensions, such as composition similarity, control parameter adjustment similarity, and so on. Therefore, the multi-dimensional similarity strategy is adopted to overcome the limitation of single similarity and ensure that the model can extract more comprehensive and multi-dimensional information from historical data, to provide more accurate predictions for the samples to be tested.
In the TS-EL-JITL, two types of similarity measurement methods are used simultaneously to ensure the diversity of the subsequent ensemble model. All the following metric formulas are actively computed in our algorithm:
Distance Similarity/correlation coefficient
In the above formula,
For similarity measurement, the sample with the highest similarity to the sample to be tested is selected. Taking cosine similarity as an example, the formula for selecting samples is as follows:
By this method, we construct
Integrating target similarity stage
In the process of steel key indicator measurement, measurement errors occasionally lead to significant deviations between readings and their true values. Meanwhile, even if certain historical samples exhibit high similarity in the feature space, unobserved characteristics with significant differences may still result in large deviations in measurement outcomes. In addition, for test samples that may possess abnormal indicators, it is necessary to select samples with boundary indicator values from the historical data to construct a similar sample set. This helps to identify potential high risks associated with the current sample.
Traditional JITL methods, which only rely on the similarity comparison between sample features, are not enough to effectively deal with the above complex situations. In order to solve this problem, the idea of target similarity is integrated to further compare the similarity of samples on the indicator. Specifically, first, build a pseudo-label prediction model
31
based on all historical data. Here, the multiple linear regression model
32
is used to give the sample to be tested a pseudo-label. Although relatively simple, the linear base model is deliberately chosen to strictly satisfy the millisecond-level inference latency constraints of legacy industrial control systems and edge computing devices.
33
The formula is as follows:
By this method, we reconstruct
EL stage
In the field of EL, the bagging method focuses on improving the robustness of prediction, while the boosting method focuses on improving the accuracy of prediction. In view of the dual requirements of accuracy and robustness in the prediction of key indicators of steel process flow, the stacking method, which combines the advantages of the two methods, is selected in this study. The specific implementation steps are as follows:
Based on The The Taking
In this process, the
Industrial deployment and practical application
To deploy the TS-EL-JITL framework in real-world steelmaking, an offline-training and online-inference paradigm is adopted. During the initial offline phase, historical process data specific to targeted production scenarios are continuously collected directly from the production line to construct a reliable sample pool. The pseudo-label model is then pre-trained to meet the strict millisecond-level latency constraints of legacy industrial control systems. Online, upon receiving a new sample, the system executes multi-dimensional similarity calculations. Following the target similarity mechanisms, the local stacking ensemble dynamically generates the final prediction. Furthermore, to maintain long-term adaptability to continuous condition deviations (e.g. equipment degradation), the historical database is periodically updated via a sliding window mechanism, ensuring system evolution without disrupting production.
Experimental design and result analysis
Experimental set-up and configurations
Experimental data and pre-processing
Data pre-processing
The original data feature dimension is often high, and there is also significant correlation between features. Therefore, pre-processing the data, eliminating redundant information, and retaining the main information of the data are of great significance to improve the data quality and better train the model.
The raw dataset first undergoes systematic data cleaning. Illegal sensor readings identified from an industrial perspective are converted to nulls and subsequently imputed using mean filling. Next, the Density-Based Outliers (DB_outliers) algorithm34,35 is applied to detect and eliminate extreme anomalies while preserving genuine operating condition deviations. Following data cleaning, a comprehensive feature selection process is executed to reduce dimensionality and avoid ill-conditioned matrices during local modelling.36,37 This process comprises three steps: (1) low variance screening to remove stagnant variables; (2) collinearity processing via Pearson correlation to exclude highly redundant features; and (3) low contribution processing using tree-based importance evaluation to retain only the critical target-related parameters. This streamlined sub-set ensures both accuracy and computational efficiency for the subsequent JITL modelling.
Dataset
To comprehensively evaluate the proposed framework, we utilise two industrial datasets collected from typical steel manufacturing scenarios: the hot-rolling process and the sintering process. The feature variables in both datasets are directly acquired from the sensor networks of their corresponding production lines, encompassing material properties, chemical compositions, and various thermal-mechanical operating parameters. Detailed statistics of the two datasets, including sample amounts, target indicators with their distributions (mean
Datasets details.
Experiments and baselines
Experimental set-up
In order to comprehensively evaluate the performance of the prediction model of key indicators of steel process flow adapting to the dynamic changes of operating conditions, three experiments were set up to verify the performances of three dimensions: global prediction ability, boundary conditions prediction ability, and operating condition deviation adaptability.
The first is the comparative experiment of the global prediction ability of the model. The dataset was divided into a training set and a test set according to time sequence, and the overall prediction performance of each model on the same test set was compared.
The second is the comparative experiment of boundary conditions prediction ability. The dataset was divided into training set and test set according to time sequence, and the prediction performance of each model on boundary samples (the top and bottom 25% quantiles) within the same test set was compared.
The third is the comparative experiment of operating condition deviation adaptability. The dataset is divided into training set and test set according to time sequence, and then the test set is divided into three test sections to observe the accuracy and robustness of the prediction results of the model in different test sections.
Hyper-parameter settings
The number of base models (
Baseline models for comparison
blackTo comprehensively evaluate the TS-EL-JITL framework, six representative data-driven models were selected as baselines:
Evaluation metrics
For the global prediction and boundary conditions prediction experiments, the commonly used evaluation metrics of regression tasks such as
where
For the comparative experiment of operating condition deviation adaptability, the above evaluation metrics are calculated for three test sections to measure the prediction accuracy. On this basis, the mean value and range of these indexes are calculated to further measure the robustness of the prediction results of the model in different test sections.
Comparative experiment of global prediction ability
In this experiment, the first 80% of the dataset is used as the training set, and the remaining 20% is used as the test set. In order to ensure the fairness of model comparison and objectively show model performance, linear regression is used as the base model for all models, including a traditional machine learning model (ML), JITL, and TS-EL-JITL. For the hot rolling scenario, the yield strength and tensile strength are predicted, and for the sintering scenario, the Al2O3 sinter composition is predicted. The prediction results are shown in Table 2. The bold result in each row is the best among the compared models for each evaluation metric.
Comparison experiment of global predictive ability (%).
ML: machine learning; JITL: just-in-time learning; TS-EL-JITL: target-similarity ensemble just-in-time learning; MAPE: mean absolute percentage error.
As shown in Table 2, among the traditional global models, the state-of-the-art tree ensemble, LightGBM, notably outperforms standard ML and XGB, demonstrating strong non-linear fitting capabilities. Furthermore, the cutting-edge tabular foundation model, TabPFN, exhibits remarkable performance, even matching the best
Comparative experiment of boundary conditions prediction ability
In the iron and steel industry, the accurate prediction of boundary indicators is of great significance for timely early warning of boundary samples and the mitigation of production risks. For example, in the hot rolling scenario, boundary values of mechanical properties usually indicate high risk, requiring special attention. However, due to the long-tailed distribution of industrial data, this constitutes a challenging imbalanced regression problem where traditional global models often fail. 41 In order to explore the performance of the model proposed in this study on the boundary values prediction task, the test sample set comprising label values beyond the central 50% range (specifically, those in the top 25% and bottom 25% tail distributions) was selected as the boundary values evaluation set to make statistics on the prediction effect. The results are shown in Table 3.
Prediction performance of boundary samples (%).
ML: machine learning; JITL: just-in-time learning; TS-EL-JITL: target-similarity ensemble just-in-time learning; MAPE: mean absolute percentage error.
Table 3 presents the predictive performance on boundary samples (top and bottom 25% quantiles). Predicting these long-tailed boundary values is a challenging imbalanced regression problem in metallurgical processes. Consequently, global models like LightGBM and TabPFN experience significant performance degradation (e.g. stagnating at a 64.71% hit rate for yield strength). This occurs because their global optimisation objectives inherently tend to over-smooth rare boundary values to fit the majority of normal samples. Conversely, localised strategies like JITL demonstrate better resilience (70.59% hit rate) but remain limited by relying solely on feature–space distances. Our proposed TS-EL-JITL successfully overcomes this, achieving an outstanding hit rate of 76.47% for yield strength (a 56% relative improvement over traditional ML) and 98.04% for tensile strength. By explicitly integrating the target similarity mechanism, TS-EL-JITL enables the algorithm to retrieve historical samples with matching boundary target behaviours. Furthermore, the ensemble stacking framework effectively mitigates the high variance and overfitting risks associated with modelling rare events. These results prove our method’s superiority in providing reliable early warnings for high-risk boundary conditions.
In order to further intuitively display the prediction performance of boundary values, the predicted values and real values of boundary samples from ML and TS-EL-JITL are compared, and the results are shown in Figure 2.

Comparison chart of predicted value and real value. The shaded areas indicate the prediction deviation, highlighting the improved fitting accuracy of TS-EL-JITL at boundary values. TS-EL-JITL: target-similarity ensemble just-in-time learning.
As shown in Figure 2, from ML to TS-EL-JITL, the degree of fitting between the predicted results and the real values is significantly improved in the details of boundary low values (the shaded area is significantly smaller). In ML, where the fitting degree is not ideal, TS-EL-JITL has higher accuracy in both trend and numerical aspects. It is proven that TS-EL-JITL can significantly improve the prediction accuracy for boundary values and can more accurately predict high-risk situations.
To isolate the specific contribution of the target similarity module, an ablation experiment was conducted comparing TS-EL-JITL with EL-JITL (which excludes the target similarity mechanism). As shown in Figure 3, removing this module causes a notable performance decline across all metrics; for instance, the 6% hit rate drops from 76.47% to approximately 68.6%. This gap highlights that in complex industrial environments, feature–space similarity alone can be misleading when similar inputs yield divergent outputs, a phenomenon known as feature–target decoupling. By integrating target similarity, the model effectively constrains the local sample pool to those truly representative of the boundary target regime. This validates the module’s critical role in enhancing predictive accuracy and providing reliable early warnings for high risk.

Performance comparison of TS-EL-JITL and EL-JITL in boundary values. The results are based on a deterministic evaluation of the fixed test set. TS-EL-JITL: target-similarity ensemble just-in-time learning; EL-JITL: ensemble just-in-time learning.
Visual analysis of sample space shrinkage
To intuitively demonstrate how the proposed target similarity mechanism mitigates feature–target decoupling, a Three-dimensional (3D) visualisation experiment of the sample selection process was conducted. We selected a representative boundary test sample from the hot-rolling dataset with a true yield strength of 285 MPa. The historical sample pool was projected into a 3D space, where the
As illustrated in Figure 4, standard JITL initially selects a broad neighbourhood of 320 candidates (light blue dots) based solely on feature–space proximity. However, due to inherent measurement noise and unobserved industrial disturbances, these feature-similar samples exhibit severe divergence in the target domain, with many drifting towards the normal operational distribution. In contrast, by imposing the target similarity constraint, TS-EL-JITL effectively shrinks this candidate space. It filters out the ‘feature-similar but target-divergent’ false neighbours, retaining a highly refined sub-set of exactly 100 core samples (dark blue triangles) that tightly cluster around the 285 MPa boundary target. This visualisation explicitly proves that TS-EL-JITL successfully extracts a high-density, target-consistent sub-set from the noisy feature-similar domain, thereby fundamentally reducing pseudo-label uncertainty for boundary values prediction.

Three-dimensional visualisation of sample space shrinkage via target similarity.
Comparative experiment on adaptability of operating condition deviation
In this study, the population stability index (PSI) was used to evaluate the operating condition deviation in the production process. PSI is an index widely used in the field of financial risk control. In recent years, it has also been introduced into the industrial scene to evaluate the product quality of the production line and the stability of the production process control. It is of great significance to improve the product quality and the control ability of the production line. The lower the PSI value, the more stable the production process. Specifically, the PSI value between 0 and 0.1 indicates that the production process has no change or changes very little, the PSI value between 0.1 and 0.25 indicates that the production process is slightly unstable, and when the PSI value exceeds 0.25, it indicates that the characteristics of the production process show great instability. The application of PSI can effectively monitor and evaluate the stability of the production process, and provide a scientific basis for timely adjustment and optimisation of production. The PSI was calculated for key production parameters to quantify their distribution shifts over time, with the results illustrated in Figure 5. It can be observed from the figure that the PSI values of almost all important production process parameters are significantly higher than 0.25 (red scatters), which indicates that the operating condition deviation phenomenon is very serious with the passage of time.

PSI values of various production process parameters. The highly scattered distribution in the upper region (PSI > 0.25) explicitly demonstrates the severity of the operating condition deviation. PSI: population stability index.
To intuitively illustrate this dynamic deviation, the dataset was simply divided into four consecutive segments in chronological order to analyse the distribution of key features over time. As shown in Figure 6, the boxplots reveal that the statistical behaviours of these representative features (such as medians and data ranges) fluctuate significantly across the four segments. This visual evidence, consistent with the PSI results, intuitively confirms that operating condition drift continuously occurs within the dataset as time progresses.

Time-segmented distribution of key features, illustrating the operating condition deviation.
In order to evaluate the impact of this model on the prediction performance of the target value when dealing with the dynamic deviation of operating conditions, the datasets are sorted in chronological order, and the training set and test set are divided accordingly: specifically, the first 70% of the dataset is used as the training set, the next 10% (i.e. 70% to 80% of the parts) is used as the first test set, the next 10% (i.e. 80% to 90% of the parts) is used as the second test set, and the last 10% (i.e. 90% to 100% of the parts) is used as the third test set. Based on this division, experiments on different models are carried out on the overall test set and three segmented test sets.
In the hot rolling quality prediction, the 6% hit rate and

Adaptability experiment of operating condition deviation. The shaded bands represent the fluctuation range (deviation) between maximum and minimum prediction performances across different chronological segments; a narrower band indicates higher robustness against condition deviation.
The results show that TS-EL-JITL achieves the best adaptation effect in the case of operating condition deviation. In the three test sections, the maximum prediction effect of the research model is better than that of the comparison model, and the prediction mean and minimum prediction effect are also the best in almost all cases, indicating that it can still maintain high prediction accuracy when dealing with the deviation of operating conditions. In addition, it can be observed that in the two experiments of hot rolling yield strength 6% hit rate and hot rolling tensile strength
To sum up, the experiment proves that TS-EL-JITL has strong adaptability in the case of dynamic deviation of operating conditions, which not only has advantages in prediction effect, but also has strong robustness. The superior performance of TS-EL-JITL under deviations can be heavily attributed to the variance-reduction property of the ensemble stacking mechanism. In real-world steel manufacturing, measurement noise and pseudo-label uncertainty are inevitable. While individual pseudo-labels or local models may suffer from this epistemic uncertainty, our approach of aggregating predictions across multi-dimensional similarity metrics inherently acts as a structural filter. This aligns with recent findings in advanced metallurgical modelling, where combining domain-knowledge integrated methods or ensemble frameworks significantly buffers the system against high-dimensional industrial noise and sensor anomalies.12,42,43 By cross-validating the target similarity through multiple base learners, the final stacked output smooths out individual prediction spikes, thereby stabilising the indicator forecasting even when the underlying data drifts.
Conclusions
By integrating a target-similarity mechanism with EL-JITL, this study proposes the TS-EL-JITL framework to effectively overcome the feature–target decoupling problem in steel manufacturing. Quantitatively, the framework delivers a maximum overall accuracy improvement of 18%. More meaningfully, for high-risk boundary samples, it achieves a 56% relative improvement in hit rate over traditional global baselines, providing reliable early warnings for industrial anomalies and effectively preventing the over-smoothing of critical production shifts. The model demonstrates exceptional robustness against dynamic condition deviations. It successfully compresses the prediction accuracy fluctuation range across varying chronological segments from over 12% down to approximately 5%, ensuring highly stable performance in shifting industrial environments. Future research will focus on cross-plant validation and fine-tuning advanced tabular foundation models (e.g. TabPFN) to capture deeper multi-indicator correlations, while strictly maintaining the millisecond-level inference latency required by legacy control systems.
Footnotes
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
Declaration of conflicting interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
