Abstract
In the hot metal pretreatment desulphurisation process, the sulphur content needs to be controlled precisely. A novel soft sensing modelling method is proposed to address the problem that sulphur content cannot be directly measured online. Firstly, the main desulphurisation behaviour is represented with the proposed simplified mechanism model for boosting computational efficiency. Moreover, the complex influence of the changing operation condition on the parameters of the simplified mechanism model is described with neural networks to compose a hybrid structured hz. Then, for training this hybrid model, a conjoint learning scheme is designed to ensure stability and reliablity. The experimental results show that the proposed hybrid model obtains higher prediction accuracy and better robustness compared with the pure mechanism modelling technique and the existing benchmark data-driven modelling methods. In addition, the proposed method is able to provide real-time prediction ability, which is substantial to the optimisation control of the hot metal pretreatment process.
Introduction
With the rapid development of modern industrial production and technology, user requirements for steel quality are increasing. As one of the harmful elements in steel, sulphur has a great adverse effect on the performance of steel. Therefore, pre-desulphurisation of hot metal by pretreatment outside the furnace has become an important technology in the steel-making process. However, due to the lack of available chemical composition sensor, the desulphurisation process is controlled manually based on empirical estimates, leading to problems such as low accuracy and poor stability. In order to optimise the operation, it is quite necessary to know the status of desulphurisation process while in progress. Therefore, the development of the soft sensing model to forecast how the sulphur content changes becomes a serious concern.
To model a process, mechanism modelling approach is a widely used soft sensing method. At present, there are several mechanism models developed for sulphur content prediction.1–3 This methodology is also employed in other steelmaking processes, such as converter process modelling and ladle furnace steel refining process modelling.4–7 However, these mechanism models are commonly used for process analysis for their well-generalisation performance. There exist some shortcomings that include the following:
The determination of the parameters in these mechanism models, which has a great impact on model accuracy, is still a challenge due to the complex physical and chemical reactions. The shift in operation conditions among different hot metal treatment batches also causes a change in these parameters. It is commonly difficult to characterise the change rule of these parameters in a mechanistic way. Therefore, these mechanism models are rarely applied directly to an actual production for process monitoring and optimisation.
To achieve more accurate prediction, several data modelling methods have been introduced into the steelmaking process,8–15 such as partial least squares (PLS),8,9 support vector regression (SVR),10,11 neural network12–14 and random forest (RF).
15
Some methods also have been used for sulphur content prediction.16,17 In these methods, the functional relationship between input and output can be discovered by extracting information from process data. The advantages of these data models are low modelling cost and strong self-learning ability, so high-precision prediction models are easier to obtain. However, there are some common defects in practical application of such models: Firstly, the resulting model is usually a black box. As a result, it is hard to provide complete reliability due to the lack of interpretability; secondly, the performance of a data model depends entirely on the quality of data. The quality of data is a generalised concept that usually includes the number of samples in the dataset and the accuracy of the numerical values for each data point. However, for steelmaking process such as hot metal pretreatment process, the data affected by many complicated environmental factors and flexible operation conditions usually include strong noises and large-scale variations. Consequently, when directly applying data modelling methods on sulphur content prediction, it is difficult to ensure prediction accuracy and generalisation performance; in addition, being restricted by production technology, only the sulphur content of starting time point and ending time point of a pretreatment cycle are recorded in the process data. As a result, a data model established on the base of these data is unable to support real-time prediction due to insufficient dynamic information.
Hybrid modelling method integrating mechanism knowledge and data information is considered capable of enhancing the prediction accuracy of a pure mechanism model and improving the reliability of a data model. At present, the effectiveness of hybrid modelling method has been verified in chemical engineering, steelmaking process, weather forecast and other fields.18–21 However, the application of such method in hot metal pretreatment desulphurisation process is rarely reported at present.
Existing reports on hybrid modelling have provided positive evaluations of the performance of hybrid models. The hybrid model has better performance compared to the pure model because it can integrate the advantages of the mechanism model and the data model. Therefore, the hybrid modelling method is chosen to develop a sulphur content prediction model. In this article, a soft sensing method based on hybrid modelling is proposed and applied to model hot metal pretreatment desulphurisation process. A simplified sulphur content mechanism model is firstly established based on the theoretical derivation and simplification. Then, neural network is introduced to represent the complex influence of the changing operation condition on those unknown parameters in the simplified mechanism model, thereby forming a serial structured hybrid prediction model to ensure both accuracy and reliability. Besides, in order to fine tune the weights of the neural network part in the hybrid model, a conjoint learning scheme is designed. Combining the advantages of mechanism model and data model, the proposed hybrid model is able to achieve higher accuracy and better robustness. Moreover, it also has the ability for real-time prediction, which is helpful to achieve optimal control of the hot metal pretreatment desulphurisation process.
The remainder of this article is organised as follows. In the second section, the mechanism modelling of desulphurisation reaction is introduced. In the third section, a novel hybrid modelling method is proposed, and a conjoint learning scheme is used to optimise the parameters of the model. In the fourth section, the effectiveness of the proposed method is validated by actual production data, and the calculation results are analysed. Finally, conclusions and future work are provided in the fifth section.
Mechanism modelling of desulphurisation process
In this section, mechanism modelling, which is as a basis of the proposed hybrid modelling method, is carried out to represent the major dynamic of hot metal desulphurisation process.
Introduction of desulphurisation reaction principle
Sulphur is usually considered as one of the main impurities in hot metal, and injecting magnesium powder is a common technology to achieve desulphurisation. 22 The schematic picture of the process of desulphurisation by Mg injection is shown in Figure 1. Magnesium metal has a strong affinity for sulphur, and its melting point and boiling point are relatively low. Therefore, the method of injecting magnesium powder into hot metal with inert gas is usually used to achieve the purpose of desulphurisation. Under the action of high temperature, the magnesium powder particles are rapidly liquefied, vaporised and dissolved in hot metal to react with sulphur. Because the reaction product magnesium sulphide has a melting point higher than 2000°C and a small density, it is kept as solid slag and finally floating up to enter the slag on the top of hot metal at smelting temperature.

Diagrammatic sketch of desulphurisation reaction.
In summary, the whole desulphurisation process can be divided into three steps:
Step 1: Magnesium powder vaporises and dissolves at a high temperature.
Step 2: The dissolved magnesium reacts with the sulphur to produce magnesium sulphide. Step 3: The generated magnesium sulphide floats upward to enter the slag.
In addition, residual slag may cause the occurrence of resulphurisation: (MgS) + [O] = (MgO) + [S], which leads to sulphur returning to the hot metal. Therefore, it is necessary to carry out slag removal treatment on the site of hot metal pretreatment. Cleaning the slag generated by the reaction as much as possible helps to reduce the risk of resulphurisation and improve desulphurisation efficiency.
Mechanism modelling of desulphurisation process
There are the following assumptions for the desulphurisation reaction shown in Figure 1: It is assumed that the mass transfer of magnesium and sulphur controls the overall desulphurisation rate. The desulphurisation reaction follows first-order reaction kinetics. The sulphur concentration and temperature in hot metal are uniform. The mechanism model of desulphurisation process can be established according to the principles of thermodynamics and kinetics. According to the principle of reaction kinetics, the change rate of sulphur and magnesium content in the reaction process can be mathematically described as shown in equation (3).
[%S]* and [%Mg]* can be determined through the thermodynamic principle of desulphurisation reaction by equations (6)–(10). According to the equilibrium relations, the molar fluxes of the reactants to the reaction interface are equal, which can be expressed by equation (6),
Parameters determination methods for the mechanism model.
Based on the theoretical and experimental analysis, Wagner's equation is able to provide precise estimation of the activity coefficients of reactants (fS, fMg) listed in Table 1. However, it is difficult to accurately estimate the utilisation rate of magnesium (fd) and the reaction rate constant (kS, kMg) by the methods listed in Table 1. Because the condition under which these methods well behaved hardly matches the practical case due to the complexity and variability in reaction conditions. Therefore, it is necessary to explore an effective strategy for estimating these hardly determined parameters for accurate sulphur content prediction.
Hybrid modelling of desulphurisation process
In this section, a hybrid modelling method that integrates mechanism modelling approach with data modelling technique is proposed to realise real-time and accurate prediction of sulphur content during desulphurisation.
The framework of the proposed hybrid modelling method is shown in Figure 2. In the proposed hybrid modelling method, mechanism model is exploited to represent the main process dynamic behaviour for reliability. To further boost the prediction accuracy, data modelling technique using back propagation neural network (BPNN) is introduced to establish a parameter estimator to determine the reasonable values of the unknown parameters in the mechanism model that may change under different production conditions. These unknown parameters, that is, kMg, kS and fd, are denoted by

Framework of the proposed hybrid modelling method.
By integration of the self-learning capability of BPNN and the good generalisation ability of mechanism model, the proposed method enables the hybrid model to adapt to the flexible operation conditions, thereby satisfying the desire for high prediction accuracy and real-time prediction capability.
Formulation of the hybrid model
In this article, data modelling method is used to capture the influence of operating conditions on those undetermined parameters in the mechanism model, so as to optimally adjust these parameters to the changing operation conditions. BPNN, as one of the most popular data modelling methods, is selected for its good learning ability and generalisation property. BPNN with strong nonlinear description capability is an effective method to fit the input-output mapping relationship for a system with complex or unclear internal structure. 29
The formulation of the hybrid model can be expressed as follows:
Strategy for training the hybrid model
To train the proposed hybrid model, it is necessary to optimise the unknown coefficients of BPNN by learning from practical data. When both the input data and the output data are available, the BPNN can be directly trained from practical data by routine training algorithms such as error back propagation. However, with respect to the proposed hybrid model, BPNN serves as its parameter estimator part whose output is fed to the mechanism part in a serial way. Therefore, the output data of BPNN, the parameters to be estimated, are not available before training. As a result, the routine training algorithms can not be used.
For training of the BPNN embedded in the hybrid model, the overall prediction error is computed by evaluation of the deviation of the prediction values of sulphur content from the real values. A conjoint learning scheme is thereafter designed to optimise the unknown coefficients of BPNN to improve the overall prediction accuracy based on the gradient descent optimisation algorithm. The proposed scheme enables the overall prediction error to be back-propagated through both the mechanism part and BPNN part. In this way, the BPNN can be trained.
The output of the hybrid model is used to calculate the overall prediction error which is defined as follows:
The overall prediction error E can be reduced by the gradient descent (GD) method. According to the chain rule for composite function derivation, the partial derivative of E with respect to the unknown coefficients of BPNN are given by equations (19)–(22),
Let θ={
It is found that GD has the problem of slow convergence. As an improvement of GD, the momentum gradient descent (MGD) method is able to effectively speed up the convergence rate of updating by adding momentum factor items. Moreover, MGD makes the updating process insensitive to local details of the error surface to prevent the algorithm from falling into the local minimum.
30
The update rule based on MGD is described as follows:
Training algorithm optimisation
The major difficulty in implementing the GD method for training the proposed hybrid model is the inability to solve the analytic formula of the partial derivative term g’ in the chained equations (19)–(22). This is because g’, which represents the mechanism part in the hybrid model, is formulated with nonlinear partial differential equations. As a result, the analytic algebraic expression of g’ is unable to be directly solved by integration. For implementation of GD method, one way is to use numerical integration method instead. However, the numerical integration method may cause some problems such as low efficiency and lack of stability, especially when the interval of integration time is large or the volume of the data set for training is great.
In order to solve the above problem, a simplified analytical mechanism model of the desulphurisation process is designed for reducing model complexity and boosting efficiency. By means of the simplification technique, an approximate analytical representation of the mechanism part of the hybrid model is obtained and the training strategy proposed in the last section can be implemented.
In the practical desulphurisation process, the sulphur content in the hot metal is usually much greater than that of magnesium, that is, [%S] >> [%Mg]. Thus we obtain
It is important to point out that the main purpose to make the above simplification is to seek an analytic formula of the partial derivative term g’. For clarity, we make an explicit definition of g’ as follows:

Workflow of hybrid model.
Training process for hybrid model.
Experiments
In this section, the effectiveness of the proposed hybrid modelling method is validated. Firstly, the effectiveness of the simplified mechanism modelling method used in hybrid modelling is validated. Thereafter, feature selection and parameter tuning are carried out for optimising the performance of the proposed hybrid modelling method. Finally, the proposed hybrid modelling method is compared with some benchmark modelling methods on practical production data.
The hybrid modelling method proposed in this article has been validated in a hot metal pretreatment plant. This plant processes 100 tons of hot metal. The basic information of hot metal includes the following categories: initial temperature, composition, hot metal weight, injection duration and magnesium consumption. The sampling process is as follows:
Before processing, the hot metal is sampled, and its composition and temperature are measured. During the pretreatment process of hot metal, the weight and duration of magnesium powder injection are accumulated and recorded in real-time during the process. Finally, the processed hot metal is also sampled and tested.
A process computer is used to track data in real time and store information throughout the process. A data list containing this information is generated after each heat is processed.
In this article, 600 samples are used for method validation, where 500 samples are randomly selected for the training process, while the remaining 100 samples are used for testing. In each sample, there are some variables recorded, including injection duration (TimePC), magnesium consumption (MgSum), initial sulphur content (S0), hot metal weight (Wm), injection rate (MgV), initial temperature (T0), initial manganese content (Mn0), initial silicon content (Si0), initial phosphorus content (P0) and the endpoint sulphur content (Send). For this study, Send is used as the output variable, while the rest are used as the candidate input variables. The variation range of each feature of these samples is shown in Table 3.
Variation range of each feature of samples.
Validation of the simplified mechanism modelling
In this section, the effectiveness of the simplified mechanism model is validated by experiments. Calculation results of the original mechanism model and that of the proposed simplified mechanism model on each sample are recorded for comparison. For each sample, the change of sulphur content in the whole process are calculated every second. The calculation results of six samples that are randomly selected from the total 600 samples are shown in Figure 4, where the calculation result of the original mechanism model and that of the simplified mechanism model are computed on the sample production operation data. In order to quantitatively evaluate the precision loss due to simplification, the endpoint sulphur content calculation results of the original mechanism model and that of the proposed simplified mechanism model on all the 600 samples are compared, and analysed with a scatter diagram, as shown in Figure 5.

Comparison of calculation results of reactant content in whole desulphurisation process. (a) Prediction results of case 1. (b) Prediction results of case 2. (c) Prediction results of case 3. (d) Prediction results of case 4. (e) Prediction results of case 5. (f) Prediction results of case 6.

Comparison of calculation results of two mechanism models for endpoint sulphur content.
It can be found from Figure 4 that the calculation results of the simplified mechanism model closely follow the original mechanism model. Moreover, it is able to be found from Figure 5 that the calculation results of the simplified mechanism model and the original mechanism model on endpoint prediction are basically consistent. In most cases, the deviation is less than 0.00053%. The maximum deviation of endpoint sulphur content calculation occurs in the rare cases that the sulphur content is extremely low. The absolute error corresponding to that is about 0.0022%. Therefore, it can be concluded that the simplified mechanism model is able to approximate the original mechanism model with high accuracy.
Optimal tuning of the proposed hybrid modelling method
In this section, the proposed hybrid model is optimised. Especially, the input feature of the BPNN part of the hybrid model and the gradient descent strategy selection for parameter optimisation are studied for attaining optimal prediction performance.
Input feature selection
In the proposed hybrid model, the BPNN part is used to provide reasonable estimation of the parameters in the mechanism model to adapt to the changing operation conditions. Therefore, the performance of the BPNN part plays a key role for the improvement of prediction accuracy. To optimise the design of the BPNN part, the choice of input features is critical as including redundant features may degrade overall performance. Therefore, feature selection is conducted to get rid of this problem, so as to improve the endpoint sulphur content prediction performance.
There are several correlation evaluation methods available for feature selection. These methods commonly work by measuring the correlation between each input feature and the output variable. Within these methods, Pearson correlation coefficient (PCC) is commonly used, which describes the degree of linear correlation between variables,

Correlation coefficients between input features and the output.
When adopting the correlation evaluation method for feature selection, all seven available variables mentioned above are initially considered as candidate features. On the basis of the kinetic principle of desulphurisation reaction, it is able to be inferred that S0 is an important factor affecting the prediction results. As a comparison, it can be found from Figure 6 that TimePC and MgSum provide more relevant results. The reason may be that TimePC and MgSum reflect the influence of reaction kinetic condition which contributes more for the final results. Therefore, TimePC, MgSum and S0 are firstly selected as the input features, while additional input features are generated sequentially from the remaining candidate features according to correlation ranking. For optimally determining the cutoff point to compose the best input features subset, after each new input feature is generated, we calculate the overall prediction error and halt the iteration when the improvement of prediction accuracy becomes non-obvious. Table 4 shows the prediction errors of this iterative process. The prediction error is expressed by the mean absolute error (MAE),
Prediction errors of the hybrid model under different inputs features selection.
Gradient descent strategy selection for parameter optimisation
The gradient descent method for parameter optimisation is studied to further improve the performance of the proposed hybrid model. Two gradient descent methods, GD and MGD, are respectively selected to train the proposed hybrid model and thereafter a comparison is carried out. Originally, the learning rates of GD and MGD are both selected as 0.001. Figure 7(a) shows the results of the two algorithms. It can be found that the loss curve obtained by MGD at the early stage of training is oscillatory, while the overall convergence rate is faster than GD. When the learning rates increase to 0.005 of which the results are shown in Figure 7(b), it can be found that GD becomes sensitive to the learning rate, while MGD continues to converge and obtains a relatively faster convergence rate. Therefore, with respect to GD, inappropriate parameter selection is easy to cause oscillation during the iterative process. Meanwhile, comparing Figure 7(a) and (b), it can be found that the final convergence results of MGD under different learning rates are similar. To sum up, MGD shows better stability performance and faster convergence rate when the parameters are selected properly. Based on the above considerations, MGD is finally used as the parameter optimisation strategy for training the proposed hybrid model.

Comparison of training performance between GD and MGD. (a) Case 1: learning rate = 0.001. (b) Case 2: learning rate = 0.005.
Validation of the proposed hybrid modelling method
The performance of the proposed hybrid model is tested and compared with some benchmark modelling methods in this section. Three types of prediction models are used, including a mechanism model, three data models and a parallel hybrid model (PHM). All these models are tested on the endpoint sulphur content prediction.
The mechanism model used is established by directly adopting the mechanism modelling method mentioned in the second section. The three data models for comparison are developed by three benchmark data-driven algorithms that have been used in the modelling task in steelmaking process, including SVR, RF and BPNN. While the PHM has been applied to the prediction of sulphur content in steelmaking process. 32 In PHM, data modelling technique is used to make up the gap between mechanism model and reality in order to obtain higher prediction accuracy. Therefore, such a PHM method is developed for sulphur content prediction in PHM process for comparison. With respect to the three data models and PHM, the feature selection method in Section 4.2 is used to determine their input variables for optimal performance. Finally, five variables S0, TimePC, MgSum, MgV and Wm are selected as the input of BPNN and PHM. Four variables S0, TimePC, MgSum and Wm are selected as the input of SVR and RF.
For the mechanism model, we aim to search for the average optimal estimation of the three unknown parameters (kMg, kS and fd) in the way that the estimation result is able to provide the minimal overall prediction error. The final estimations are: kMg = 0.412, kS = 0.658 and fd = 0.506. In order to optimise the hyperparameters of the data models, 5-fold cross-validation method is used. For the SVR model, the penalty coefficient and kernel width are selected as 4 and 2, respectively. For the BPNN model, the number of hidden layer nodes is selected as 7. For the RF model, the number of trees is selected as 120. The statistical measures considered for performance comparison include MAE and the root mean square error (RMSE).
Overall prediction performance
Comparison is firstly carried out on the overall performance of endpoint sulphur content prediction. Considering the randomness in the experimental procedure, each model runs five times independently. Table 5 lists the average (AVE) and maximum (MAX) of the five running results of each model.
Performances of the different models for endpoint prediction.
It can be found from Table 5 that the average and maximum testing errors obtained by the proposed hybrid model are both smaller than those of other models. Table 5 shows that the average RMSE of the five runs of the proposed hybrid model is 10.31 ppm, which is 9% lower than that of SVR (11.35 ppm). Besides, the average RMSE of the proposed hybrid model is reduced by 9–62% compared to all benchmark models, and the maximum is reduced by 21–66% . Specifically, the mechanism model obtains the largest prediction error. The reason may be that the parameters of the mechanism model are fixed rather than changing with the actual operating conditions. As a result, the mismatch with the practical case leads to a large prediction error. Compared with the mechanism model, the three data models show smaller prediction errors, which may be related to the powerful nonlinear approximation capability of the data-driven algorithms. This nonlinear approximation capability is also reflected in the proposed hybrid model and PHM. Both of them obtain better prediction performance than the mechanism model. However, the overall prediction accuracy of PHM is worse than that of the proposed hybrid model. PHM shows a limited compensation ability for the deviation of the mechanism model. By contrast, the proposed hybrid model achieves better accuracy, which shows that this serial hybrid mode gives a better play to the complementary ability of data model and mechanism model for the modelling of hot metal desulphurisation process.
For clarity, the boxplots of the results of five independent runs are shown in Figure 8. It can be found from Figure 8 that the prediction results of the proposed hybrid model show the lowest fluctuation. Specifically, the MAEs obtained by the proposed hybrid model range from 7.65 to 8.50 ppm, with a variance of 8.40E-10. The range of RMSEs is 9.45 to 10.80 ppm, and the variance is 2.14E-09. It can be found from Figure 8 that, among the five benchmark models, SVR and PHM obtain predicted results with lower fluctuation. For SVR, the variance of MAEs obtained by five runs is 1.81E-09, which is about 2.2 times larger than the proposed hybrid model. The variance of RMSEs is 1.48E-08, about 6.9 times larger than the proposed hybrid model. For PHM, the variance of MAEs is 3.22E-09, about 3.8 times larger than the proposed hybrid model. The variance of RMSEs is 3.53E-09, about 1.6 times larger than the proposed hybrid model. The above experiments reflect that the proposed hybrid model not only ensures the prediction accuracy but also achieves more steady performance.

Comparison of prediction errors of endpoint prediction. (a) Boxplot of MAEs. (b) Boxplot of RMSEs.
In the experiment, it is also found that the dynamic parameters of different samples are indeed different. The distribution histogram of dynamic parameters for all samples is shown in Figure 9. It can be observed from the experimental results that there are relatively significant changes in parameters between different heats. The range of fd is 0.35∼0.63; The range of kMg is 0.52∼0.94.

Distribution histogram of dynamic parameters. (a) Distribution histogram of kMg. (b) Distribution histogram of fd.
The results of model training show that the changes in these parameters are related to different operating conditions. Pearson coefficient is used to evaluate the correlation between various input features and dynamic parameters to further analyse the patterns of parameter changes. The results of the correlation assessment are listed in Table 6.
Pearson coefficients between input features and dynamic parameters.
It is able to be found from Table 6 that fd is most significantly affected by S0, and there is a certain positive correlation between them. The reason may be that the higher the initial sulphur content, the greater the amount of magnesium required to reach equilibrium conditions, making the desulphurisation reaction easier to proceed. Therefore, the utilisation rate of magnesium is larger. Besides, TimePC and MgSum also have a certain correlation with fd. This may be caused by the correlation between input features. Objectively, a higher initial sulphur content is likely to result in greater injection duration and magnesium consumption. In addition, the kMg is most significantly affected by Wm, and there is a certain negative correlation between them. The reason may be that the larger the weight and volume of hot metal, the slower the overall mass transfer process.
Prediction performance under different operating conditions
In this section, the comparison on the prediction performance under different operating conditions is carried. Due to that the practical production environment is complex and changing, the prediction performance under different operating conditions is an important aspect of concern in the practical desulphurisation process.
In practical desulphurisation process, it is customary to divide the operating conditions into two types, namely high sulphur and normal sulphur, according to the magnitude of the value of initial sulphur content. Figure 10 is the frequency distribution histogram of all samples in terms of initial sulphur content. Customarily, the initial sulphur content greater than 0.06% is regarded as high sulphur, a special operating condition of concern in practice. Therefore, the prediction performance under this special operating condition is analysed in this section. The samples with initial sulphur content greater than 0.06% are taken as high-sulphur samples. The basic information of high-sulphur samples are shown in Table 7.

The frequency distribution of all samples in terms of initial sulphur content.
Range of various features of high-sulphur samples.
For these samples, the average and maximum of five runs of each model is listed in Table 8. The prediction results of five independent runs produced by different methods are used to draw the boxplots as shown in Figure 11.

Comparison of model prediction accuracy for high-sulphur samples. (a) Boxplot of MAEs. (b) Boxplot of RMSEs.
Prediction performances of models for high sulphur samples.
As shown in Table 8, compared with other models, the proposed hybrid model has the smallest prediction error, including RMSE and MAE. For RMSE, the average of the five runs produced by the proposed hybrid model is 46–84% lower than that of the other benchmark models, and the maximum is 64–85% lower than that of the other benchmark models. It can be found from Figure 11 that the proposed hybrid model shows the most stable performance. Specifically, the fluctuation ranges of MAE and RMSE obtained by the proposed hybrid model are 8.77 and 13.48 ppm and 9.33 and 14.76 ppm, respectively. In addition, PHM shows the lowest fluctuation among the five benchmark models. For MAE, the variance of five running results produced by PHM is 2.62E-07, about 7.2 times larger than the proposed hybrid model. For RMSE, the variance is 4.42E-07, about 12.1 times larger than the proposed hybrid model. Compared with the test results for normal samples, the proposed hybrid model shows more obvious advantage than other models under the condition of high sulphur.
The experimental results show that the proposed hybrid model is superior to other models in terms of prediction accuracy and stability for high-sulphur samples, which reflects its better adaptability to special operating conditions. This may be due to the fact that the proposed hybrid model only leaves the unmodelled dynamic characteristics to be learned by the data-driven algorithm. As a result, the proposed hybrid model obtains lower dependence on data, and shows good generalisation performance for high-sulphur conditions with few samples. By contrast, the pure data modelling methods need to use data information to model the whole dynamic behaviour. The complete dependence on the data may lead to large deviations in the predictions of these data models for samples with special operating conditions.
Prediction performance on small data sets
The learning ability under small-scale samples is another important index for the rapid construction and application of a model. Especially for the hot metal pretreatment process, the accumulation of samples of 600 batches usually takes several months. Therefore, the learning ability under small-scale samples is considered and tested in this section.
We randomly select 100 samples, half of which are used as the training set and the rest are used as the test set. Feature selection and parameter tuning are re-performed to adapt to small-scale training data. Each model is still run five times independently. The average and maximum results of five independent runs of each model are listed in Table 9. The boxplots of MAEs and RMSEs of the five independent runs produced by different methods are obtained, as shown in Figure 12.

Comparison of model prediction accuracy under small data set. (a) Boxplot of MAEs. (b) Boxplot of RMSEs.
Prediction performances of models on small data set.
It can be found from Table 9 that the proposed hybrid model obtained smaller prediction error on the small data set. Besides, Figure 12 shows that the proposed hybrid model has a more stable performance. The fluctuation ranges of MAE and RMSE obtained by the proposed hybrid model are 9.74 and 11.78 ppm and 12.30 and 15.30 ppm, respectively. Among the five benchmark models, SVR achieves better average accuracy and stability. For MAE, the variance of five running results produced by SVR is 2.64E-08, about 3.9 times the proposed hybrid model. For RMSE, the variance is 5.11E-08, about 3.3 times the proposed hybrid model. These results show that the proposed hybrid model also achieves better accuracy and stability with few samples.
Real-time prediction performance
Real-time prediction capability is a unique ability of the proposed hybrid model. Because the sulphur content information in the intermediate process of desulphurisation reaction is immeasurable, it makes the data set collected unable to provide sufficient dynamic information. As a result, the data models and PHM trained on this data set are not capable of real-time prediction.
For testing the real-time prediction performance of the proposed hybrid model, several samples are selected randomly from different operation conditions including different initial sulphur content, different injection duration and different magnesium consumption. As an example, the real-time prediction results for six samples are shown in Figure 13. Among them, subgraphs (a)–(b) represent cases with lower and higher initial sulphur content, respectively. Subgraphs (c)–(d) represent cases with more and less magnesium consumption, respectively. Subgraphs (e)–(f) represent cases with shorter and longer injection duration, respectively. The red curve in each subgraph of Figure 13 represents the sulphur content of the whole process predicted by the proposed hybrid model, and the black dotted line represents the practical endpoint sulphur content of the current furnace. In addition, the input features and dynamic parameters used for each sample in Figure 13 are shown in Table 10.

Real-time prediction of sulphur content by hybrid model. (a) Prediction results of case 1. (b) Prediction results of case 2. (c) Prediction results of case 3. (d) Prediction results of case 4. (e) Prediction results of case 5. (f) Prediction results of case 6.
Input features and dynamic parameters used for each sample in Figure 13.
Magnesium powder undergoes stages of mass transfer and reaction after entering hot metal, which are influenced by dynamic limiting conditions such as temperature, composition and injection rate of the hot metal. Therefore, the theoretical time to reach steady state is not necessarily consistent with the experiment. Considering that the entire desulphurisation kinetics process is affected by some restrictive conditions, this article models based on the thermodynamics and kinetics of the reaction, so that the impact of these conditions on the predicted results can be addressed. It can be found from experimental experience that the reaction time of most furnaces is relatively close to the injection operation time, as the reaction process is rapid. However, there are still some cases that are not very close, so it is necessary to consider the dynamic characteristics.
It can be found from Figure 13 that the prediction performance in the entire reaction process is relatively consistent with the experience of the dynamic behaviour of desulphurisation, showing that the proposed hybrid model is able to support real-time prediction despite the data set lacks adequate information of the reaction dynamics. The reason for this is that, in the proposed hybrid model, the dynamic behaviour of desulphurisation reaction has mainly represented by the mechanism part while the data set is only responsible for fine adjustment of the underdetermined parameters to enable the hybrid model to adapt to changing operation conditions. The real-time prediction capability makes the proposed hybrid model competent to process monitoring and process operation guidance.
Conclusion
In this work, a novel soft sensing method of sulphur content based on hybrid modelling is developed for the hot metal pretreatment process. In the proposed hybrid modelling scheme, the main process dynamics is represented by the simplified mechanism modelling technique while the potential unmodeled details are described by BPNN-based data modelling method. This hybrid modelling method is able to integrate the generalisation ability of mechanism modelling technique and the learning capability of data modelling method. It can be concluded from the experimental results that the proposed sulphur content hybrid prediction model achieves both high accuracy and good robustness. Moreover, the proposed hybrid model is able to provide real-time prediction ability in the whole desulphurisation process. Therefore, the proposed hybrid model is quite suitable for process monitoring and operation optimisation.
Footnotes
Declaration of conflicting interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by the Fundamental Research Funds for the Central Universities, National Natural Science Foundation of China (grant number N224001-8, U21A20117).
