Abstract
Gearboxes are critical transmission components in the drivetrain of wind turbine, which have a dominant failure rate and the highest downtime loss in all wind turbine subsystems. However, load variations of wind turbine gearbox are far from smooth and usually nondeterministic, which result in inconsistent data distributions. To solve the problem, a novel performance degradation assessment and prognosis method based on maximum mean discrepancy is proposed to test the difference between data distributions and extract the characteristics of multi-source working conditions data. Besides, the increase in sensors will bring more difficulties to establish prediction models in real-world scenarios due to different installation locations. In view of this, a transfer learning strategy called joint distribution adaptation is utilized to adapt data distribution between multi-sensor signals. Nevertheless, the presence of background noise of wind turbine signals restricts the applicability of these algorithms in practice. To further reduce the distribution difference, a novel criterion is proposed to evaluate and measure the data distribution difference between known and tested working conditions based on the witness function of maximum mean discrepancy. The application and superiority of proposed methodology are validated using a wind turbine gearbox life-cycle test data set. Meanwhile, model comparison and cross-verification are conducted between conventional and proposed prediction models. The results indicate that the proposed method has a better performance in performance degradation assessment for wind turbine gearbox.
Keywords
Introduction
Gearboxes are critical transmission components in the drivetrain of wind turbine (WT) as they withstand highly volatile and rough working conditions. Undetected faults in WT gearbox may cause degradation in the performance of the equipment, which have a serious impact on safe operation of wind farms and power grids. Fault prognostics and health management (PHM) for WT gearbox is an effective method to improve the service life of WT. 1 In general, a typical PHM of rotary machinery can be divided into four steps: data acquisition, feature extraction, health indicator (HI) construction, and remaining useful life (RUL) prediction. 2 Among these, selecting an effective feature for HI construction is considered as the most challenging step for performance degradation assessment of WT gearbox. 3 Also, challenges such as upcoming large data volumes, variability, and velocity will bring more difficulties for signal processing methods. 4
Rotating machinery and components such as shaft, bearing, and gear are inevitable parts of WT gearbox, which have specific vibration signature for their standard condition that changes with the development of damage. 5 Therefore, vibration analysis methods have been widely developed in fault detection of gearbox. In general, condition indicators which refer to the characteristics extracted from measured vibration signals are applied to gearbox degradation assessment. Wang et al. 6 utilized time-domain and frequency-domain statistical features to detect early gear faults and track gear performance degradation. However, spectrum leakage under variable speeds is a major drawback of time-domain and frequency-domain analyses. To overcome these drawbacks, Ni et al. 7 and Liu et al. 8 applied fault frequency band energy selected by time–frequency domain analysis to gearbox condition monitoring. However, investigated methods discussed in these literatures were mainly suitable for conventional gearboxes, not for WT gearboxes considering their high reliability, complicated load, and long service life requirement. 9 Igba et al.10,11 considered that signal correlation with root mean square (RMS) feature is superior for detecting progressive failures such as bearing pitting or shaft cracks in WT gearbox, while peak values have a high sensitivity in terms of identifying gear tooth fractures. However, both of them are easily affected by background noise unrelated to fault or normal operation. The residual signals and time synchronous averaged (TSA) signals attempt to overcome the limitations of time-domain features with its ability to easily identify and isolate fault frequencies for WT gearbox health condition monitoring.12–14 However, the kinematic scheme of WT gearbox includes parallel spur gear stage and planetary gear stage. Among these subcomponents, the high-speed parallel gear stage is found to be the most unreliable gearbox stage, 15 which has attracted many research works in this field.16–18 Conversely, vibration signals measured from planetary gear stage and their spectra are more complex than those obtained from parallel gear stage. Consequently, vibration-based methods working well for parallel gear stage often fail to detect and diagnose faults of planetary gear stage. To the best of our knowledge, the literature reported on planetary gearbox condition monitoring under nonstationary condition is deficient.19–21 To sum up, vibration-based condition monitoring should be a relatively tough task for multi-component WT gearbox. In addition to vibration-based methods, many research works have been conducted in WT gearbox degradation assessment based on different signals, such as temperature, current, and oil analysis.22–26 Among them, reference patterns-application based on similarity-based measure between normal and real-time operating condition is commonly used to evaluate the WT gearbox health condition.27–29 Mahalanobis distance is an effective tool to extract the characteristic difference of data distribution, better than Euclidean distance considering the correlation of different variables.30–32 However, conventional Mahout’s distance measuring methods perform poorly compared to alternatives when high data dimensionality is present. 33 Hence, the prediction capability of supervisory control and data acquisition (SCADA) data is posterior to vibration data in WT gearbox fault detection due to the response to fault. 34 Maximum mean discrepancy (MMD) is a modern unsupervised kernel-based pattern recognition method that, paired with the proper feature representations and kernels, may improve differentiation between two distinct populations. 35 Due to this heritage, the proposed assessment method can lead to enhanced group discrimination and accuracy based on raw time series vibration data directly. To the best of our knowledge, MMD-based method has not been previously reported in fault detection of WT gearbox.
Vibration signals generated from WT gearboxes are mixed signals from multiple blind sources, where signals generated from gear, bearing, and shaft are purely additive.36,37 Meanwhile, different gear stages suffer from different loads due to large transmission ratio and transmission efficiency, 38 which will result in inconsistent data distribution. Therefore, it is subjected to inevitable cross-term interferences and makes conventional prediction model inadaptable to the change with mismatch of probability distribution. Moreover, owing to the attenuation and interference induced by complicated signal transmission path, it is difficult to acquire effective fault information via a single sensor. 39 To obtain the reliable health condition of gearboxes, multi-sensor fusion has been conducted to obtain reliable health information on WT gearboxes.40–42 Hence, there must exist a potential correlation between different sensors since they are mounted on the same transmission path. 43 Therefore, how to mine potential correlation is a vital issue in performance degradation assessment of multi-stage WT gearbox. Joint distribution adaptation (JDA) was proposed to adapt training domain data and target domain data through the marginal and conditional distribution adaptation, 44 which is widely used in transfer learning.45–48 However, the presence of background noise restricts the applicability of these algorithms in real scenarios of WT gearbox. Varying loads or speeds aggravate data distribution difference, which may result in incomplete adaptation. To further reduce the difference, the integration of JDA and witness function of MMD (WF-MMD) is proposed to evaluate and measure the distribution difference of each known and tested working conditions.
In view of the above-mentioned issues, a multi-sensor transfer learning strategy is proposed to solve the data distribution mismatch due to the change in sensor locations and operating conditions. The multi-source domain ensemble learning mechanism is utilized to make full use of data from multiple modeling-related domains to improve the robustness of prediction model. Furthermore, a promising dissimilarity-based measurement is proposed to detect the WT gearbox health condition using MMD. To reduce incomplete adaptation caused by background noise and varying operations, a novel criterion based on WF-MMD is proposed to further reduce the data distribution difference. There are two steps in the proposed method. First, MMD is used to evaluate signal distribution difference and extract HI of WT gearbox. Second, JDA combined with WF-MMD is employed to predict future tendency. The effectiveness of the proposed methodology is validated using experimental data collected from a WT gearbox highly accelerated life test (HALT). Experimental results illustrate that proposed method provides a robust fault feature extraction and enhances health assessment capability.
The remainder of this article is organized as follows: Section “Methodology” outlines the systematic methodology of MMD and JDA. Univariate tendency prediction strategy is briefly described. Furthermore, WF-MMD-JDA is proposed for enhanced adaptation of data distribution. Section “Proposed performance degradation assessment methodology” shows the process of proposed method based on MMD and multi-sensor transfer learning in detail. In section “Experimental verification and analysis,” the robustness of proposed method is validated based on a WT gearbox HALT data set. The conclusion and recommendations for future research are presented in section “Conclusion and future work.”
Methodology
MMD
MMD-based distance is a new statistic to quantify the dissimilarity of probability measures and depends on finding a function from among the kernel space that can maximize the distance.
Health assessment indicator
Assume there are two sets with n observations from a domain
A reproducing kernel Hilbert space (RKHS)
An unbiased estimate of
where u is unbiased, h(·) is the kernel of U-statistic defined as h(xi, xj, yi, yj) = k(xi, xj) + k(yi, yj) − k(xi, yj) − k(xj, yi). Intuitively, the empirical test statistic
Although MMD has a nice unbiased and minimum variance U-statistic estimator (MMD
u
), it cannot be directly applied since MMD
u
costs O(n2) to compute based on a sample of n data points. To improve the computational efficiency and obtain an easy-to-compute threshold for hypothesis testing, an alternative statistic for
However, fault detection requires significant new derivations to obtain test threshold. Moreover, fault detection consists of a sum of highly correlated MMD statistics because these
where B is the fixed data length,
For all MMD statistics, a small value close to 0 means that the data X and Y are sampled from the same distribution, whereas a large value means that the embedded distributions in samples X and Y are significantly different.
Similarity assessment based on WF-MMD
WF-MMD is utilized to visualize where a statistical model most misrepresents the data. 51 Similarity evaluates whether operating conditions modes are measureable. In this article, WF-MMD is proposed for data-enhanced adaptation, which is presented by
where X is a set of samples from unidentified operating condition, Y is a set of samples from identified operating condition, and ϕ(x) is the kernel mapping k(·,
It is important to mention that x in equation (6) could be any point in x∈ x , and μX(x) and μY(x) is the expectation of the density function on two distributions X and Y, respectively. The proposed WF-MMD is more of a criterion to visualize and evaluate the similarity of two data clusters in kernel space.
JDA
JDA is aimed at adapting data distribution of different operating conditions in new space through a transformation method. The step-by-step procedure of JDA can be described by 44
Define a domain
Given domain
Given labeled source domain
Then, the key of JDA method is to find a transformation T so that the distance between Ps(Txs) and Pt(Txt) after transformation can be as close as possible, so is the distance between Qs(ls|Txs) and Qt(lt|Txt). Dimensionality reduction methods (principal component analysis (PCA)) can learn a transformed feature representation. Through denoting X = [x1, …, xn] as the input data matrix and H as the centering matrix, the covariance matrix can be computed as XHXT. The learning goal of PCA is to find an orthogonal transformation matrix A such that embedded data variance is maximized
where tr(·) denotes the trace of a matrix. This optimization problem can be solved by eigendecomposition XHXTA = AΦ, where Φ are the k largest eigenvalues. Then, the optimal k dimensional representation is presented by ATX.
However, the difference in both marginal and conditional distributions is still significantly large through feature transformation. Consequently, the next step of JDA after feature transformation is divided into two parts: marginal distribution adaptation and conditional distribution adaptation.
Marginal distribution adaptation
To reduce the difference between marginal distributions Ps(ATxs) and Pt(ATxt), MMD is adopted to compare different distributions, which computes the distance between sample means of source and target domain data in the k-dimensional embeddings 52
where M0 is a MMD matrix and is computed as follows
where ns and nt are the number of samples in the source domain and the target domain, respectively. By minimizing equation (8), marginal distributions between domains are drawn close under the new representation ATX.
Conditional distribution adaptation
However, reducing the difference in marginal distributions does not guarantee that conditional distributions between domains can also be drawn close. The second step is to adapt the conditional probability distribution of source domain and target domain. However, there is no labeled data in target domain, which means that conditional distribution Qt(lt|xt) cannot be obtained directly. Therefore, the solution is to train prediction model in source domain and to get pseudo-labels of target data. Then, the MMD distance between conditional distributions can be expressed as
where Mc is computed as follows
where
Iterative refinement and learning algorithm
JDA is aimed to simultaneously minimize the difference in both marginal and conditional distributions. Thus, equations (8) and (10) are incorporated, which forms the JDA optimization problem
where λ is the regularization parameter.
According to the constrained optimization theory, JDA optimization problem can be solved by Lagrange method. The Lagrange function of equation (12) can be defined as
where I is the unit matrix and Φ is the Lagrange multiplier.
Through the adaptation of JDA, data distribution difference between unknown conditions can be eliminated.
Univariate tendency prediction
Estimating WT gearbox performance degradation is to predict the tendency of degradation indicator via a univariate time series prediction method. The univariate time series prediction is defined as follows: consider HIs measured up to time t and a time series of observation Xt = {xt, xt−r, xt−2r, …, xt−(n−1)r} extracted from HIs, where r is the interval of the measure and (n − 1) is the length of data. The value of HI (xt+p) is required to be known at time (t + p), where p is the length of prediction. (xt+p) is obtained by following two steps:
Define how to get xt+p. It can be obtained directly from observation Xt, which requires to get a prediction model for each length p. A univariate prediction model is presented as follows
where f(Xt) represents the univariate prediction model of time series Xt.
2.Choose prediction model. Regression model, as one of the most popular data-driven methods for tendency prediction, attempts to find degradation process by regression functions and then predict future performance degradation until the predicted value reaches a pre-defined failure threshold.
Using historical run-to-failure HIs, we can build regression prediction models for degradation indicators. The objective of regression analysis is to find an empirical relation for predicting WT gearbox performance degradation.
Enhanced prediction adaptability based on WF-MMD and JDA
During the operation of WT gearbox, varying loads or speeds aggravate data distribution difference, which may result in incomplete adaptation. To further reduce the incomplete adaptation, we combine WF-MMD with JDA to evaluate and measure the distribution of each known and tested working conditions. The proposed multi-sensor adaptation methodology can be summarized as follows:
For known vibration signals x(1), …, x(n) (n = 3 in this article) and an unknown signal xt, merge known vibration signals (x(1), …, x(n)) into a source domain data set Xs and set unknown vibration signals xt as a target domain data set X t .
Reduce the distribution difference between source domain data set Xs and target domain data set X t through the JDA method as follows
where JDA(·) refers to equation (12), and X′s and X′t represent the source domain data and target domain data after the adaptation, respectively.
Then divide the X′s into single adaptation signals {x(1)′, …, x(i)′} and train adaptation signals to obtain the regression model { f1(·), …, fn(·)}.
Utilize WF-MMD to measure the distribution difference between X′s and X′t according to equation (17)
After obtaining the WF-MMD value, the weight value of each regression model can be calculated by
Integrated regression model ft(·) of unknown signal xt is represented by equation (19). The detailed multi-sensor adaptation process is shown in Figure 1

Flowchart of multi-sensor adaptation.
Eventually, the trend prediction can be carried out through integrated regression.
Proposed performance degradation assessment methodology
There are two steps in the proposed method. First,
Collect vibration signals from the entire life WT gearbox and extract
Merge n operating conditions into a huge source domain. Then, the inter-domain distribution difference between unknown working condition data (target domain) and source domain can be reduced by equation (16);
Utilize WF-MMD to further eliminate distribution difference;
Obtain the weight value of different models compared with unknown operating condition;
Merge regression models to obtain the integrated regression model of unknown operating condition, and the trend prediction can be carried out. The flowchart of proposed method is shown in Figure 2.

Flowchart of the performance degradation assessment.
The performance of different regression models is evaluated using three statistics: mean absolute error (MAE), root mean square error (RMSE), and mean absolute percentage error (MAPE). The mathematical expressions are as follows
where n is the number of data, and y(i) and
Experimental verification and analysis
WT gearbox test rig
The WT gearbox test rig is developed to simulate the entire life operation under actual working condition. Figure 3 shows the line diagram of designed WT gearbox test rig. The test system is composed of the following main components: (1) dual-linkage drive motor with a rated speed up to 1500 r/min and variable frequency controller, (2) dual-linkage load motor with a rated torque up to 21,134 N m, (3) main test WT gearbox linked with accompanied test WT gearbox through the swelling sleeve shown in Figure 4(a), (4) two torque and speed sensors at both sides of the rig, and (5) two universal couplings at the gearbox output shaft. The drive motor is utilized to substitute the blade to transmit mechanical power. The load motor, also called generator, is employed to generate electric power. Unlike the conventional gearbox test rig, the main tested WT gearbox is close to generator, while the accompanied tested gearbox is located near the motor. The data collection system is composed of vibration, speed, torque sensors, and acquisition module, which can be seen in Figure 4(b).

Diagram of WT gearbox test rig.

WT gearbox test platform: (a) tested WT gearbox and (b) dual-linkage motor and acquisition system.
Figure 5 depicts the test WT gearbox, which consists of two parallel gear stages and one planetary gear stage. Its nominal output power is 1.5 MW. It can be observed from Figure 5(b) that the planetary stage gear pair consists of sun gear 1, planetary gear 3, and ring gear 5. Ring gear is fixed, and planetary arm 4 is connected with the main shaft via the locking plate, constituting input terminal. Sun gear shaft drives low-speed shaft 2 through internal spline, forming output terminal. Low-speed stage gear 6 meshes with middle-speed stage pinion 8 which makes up the intermediate-speed stage gear pair. High-speed stage gear pair is formed by high-speed stage gear 11 and middle-speed gear 7. Planetary stage is internal bevel gear transmission. Intermediate-speed shaft 9 is external bevel gear transmission, so is the high-speed shaft 10, which is then coupled to generator drive end. Internal configuration of the WT gearbox is given in Figure 5(a). The total transmission ratio of the gearbox is 104.0778, and the key parameters of WT gearbox are listed in Table 1.

Tested WT gearbox configuration: (a) three-dimensional structure and (b) two-dimensional planar structure.
Properties of WT gearbox.
WT: wind turbine.
Referring the standard and prior experience of manufacturers in wind industry,53,54 four acceleration sensors are attached on the casing of the gearbox from low-speed to high-speed stages. The exact location of sensors can be seen in Figure 6: (1) low-speed shaft bearing (LSSB), (2) high-speed shaft (HSS), (3) planet inner gear ring (PIGR), and (4) torque arm and bearing (TAB). Vibration signals are collected continuously throughout the entire test. Considering fault characteristic frequency of WT gearbox and the long term of HALT, the sample rate of signals is set at 4096 Hz.

Orientation of vibration sensors marked as red elliptic region: (a) LSSB, (b) HSS, (c) PIGR, and (d) TAB.
HALT
It is well known that reducing the test time without changing the fatigue capacity of gearbox is a challenge to all manufacturers. To accelerate the test, the 135% extreme load with an overturning moment of 14,491 N m is applied to the rig. Hence, the test speed of 1440 r/min is suggested as rated working speed. After 430 h (approximately 25 days) of operation, HALT is stopped when the gearbox is stuck, and a failure occurs. It should be noted that WT gearbox is dismantled and maintained on 17th day due to the abnormal sound. Also, the status of gearbox is obtained to establish the relationship with monitoring signals. The condition of gear, bearing, and oil filter at the end of HALT is shown in Figure 7. The debris damage not only happens to low-speed stage bearing but also at ring gear. A contact hard line is generated on the surface of high-speed stage gear and spline. Metal debris is found on oil filter. Meanwhile, more stretch marks are formed with long-term cyclic contact load.

Damaged components of WT gearbox: (a) low-speed stage bearing roller, (b) high-speed stage gear, (c) spline, (d) ring gear, and (e) oil filter.
Considering the gearbox fault frequency and the volume of monitoring signals as much as possible, we intercept 1 s vibration data per hour for signal processing and tendency prediction as illustrated in Figure 8. It can be observed from Figure 8 that there exists an increasing trend in the amplitude of all vibration signals with the progress of time. The vibration amplitude of the second sensor, mounted close to HSS bearing (Figure 8(b)), is larger than the other three sensors. Owing to the complicated transformation path, the vibration amplitude of PIGR (Figure 8(c)) is the lowest. The trend of LSSB (Figure 8(a)) is smooth and there is a sharp rise in the tendency of TAB (Figure 8(d)). However, due to background noise in raw time series data, not much fault information can be extracted from vibration signals.

Measured life-cycle vibration signals: (a) LSSB,(b) HSS, (c) PIGR, and (d) TAB.
Feature extraction and health assessment
Consequently, 12 time-domain features are extracted from HSS vibration signal: root mean square (RMS), minimum (Min), maximum (Max), mean (M), kurtosis value (Kv), shape factor (SF), skewness (Sk), impulsive factor (IF), absolute amplitude (AA), root amplitude (RA), peak to peak (P-P), and clearance factor (Clf). The results are shown in Figure 9. It can be observed that there are many differences in each feature throughout the entire life of WT gearbox. Most of the extracted features depict a significant increase in amplitude over the complete life cycle and thus reflect its degradation trend. However, the problem of high-frequency fluctuations can be observed obviously owing to the operation interference noise so that early HSS degradation cannot be well known. It will increase the difficulty of building prognostic models. Furthermore, time-domain features are significantly divergent at the last failure period, which may increase difficulty to define a failure threshold. It can be concluded that the tendency performance of time-domain features is not satisfactory. Nevertheless, effective performance degradation evaluation can capture the change in performance at different life-cycle periods. Therefore, effective features that are in a trend and correlate to failure progression should be obtained.

Time-domain features of vibration signal: (a) RMS, (b) Min, (c) Max, (d) M, (e) Kv, (f) SF, (g) Sk, (h) IF, (i) AA, (j) RA,(k) P-P, and (l) Clf.

From Figure 10, we can observe that
To further verify the accuracy of M-statistics model, the other three vibration signals are selected to train M-statistics models. Figure 11 shows the indicators calculated by the other three signals. The trends of

As for the determination of life status in Figures 10 and 11, there is no uniform criterion for judging the life status of WT gearbox. Regular gearboxes also do not have relevant judgment standards. These critical points are usually determined by experienced experts according to the actual operation. In this article, we do not attempt to propose a new judgment standard but convert the experience of experts into a smart health self-assessment by realizing the self-prediction of degradation with the same type of WT gearbox.
To quantitatively quantify the suitableness of degradation assessment of different HIs, the monotonicity, prognosability, and trendability are utilized to carry out a comparison. 55 These three evaluation indexes are a type of metrics used to quantify the suitableness and are defined as follows:
Monotonicity
where n is the number of measurement time points. The monotonicity of indicator is calculated by the average difference of the fraction of positive and negative derivatives for each path.
Prognosability:
where h is the value of degradation indicator
Trendability:
where m is the number of monitoring systems.
Figure 12 shows the three evaluation indexes for each degradation indicator, respectively. Table 2 shows the suitableness for

Indices performance metrics for time-domain features and
Suitableness for
MMD: maximum mean discrepancy; LSSB: low-speed shaft bearing; HSS: high-speed shaft; PIGR: planet inner gear ring; TAB: torque arm and bearing.
Tendency prediction and adaptation comparison
Using historical run-to-failure vibration signal analysis and health assessment, a data-driven prognostic methodology is proposed for degradation tendency prediction. According to the preliminary experiment in our lab, least square support vector regression (LS-SVR) is chosen for univariate tendency prediction, that the involved regression model makes a satisfactory prediction performance. The model is trained and calibrated offline by utilizing monitoring status from the past lifelong histories and then predicts future time series, which are those data that do not participate in the training step.
In general, PHM for WT gearbox shall be carried out once monitoring signals are acquired. However, it should be noted that data samples obtained at a relatively early stage of HALT may not provide enough fault information to perform performance degradation prediction. Therefore, data after the 320th hour at the wear period are taken as the training and test samples for tendency prediction. The model data set consists of 110 samples. First, 55 training samples of the first sequence of the data set are utilized to obtain prediction model. The aim is to build the relationship between future indicator and previous few indicators of WT gearbox.
The first step is to verify the effectiveness of tendency prediction model, LS-SVR is utilized to establish four offline prediction models based on four different signals (LSSB, HSS, PIGR, and TAB), respectively. Then, four different prediction models are utilized to predict future tendency of corresponding signals. Finally, 55 test samples of the last sequence are fed to the above-trained prediction models to obtain real-time health condition. Figure 13 depicts actual performance degradation tendency (blue line) and predicted performance degradation tendency (red line), highlighting the performance of tendency prediction models.

Tendency prediction results by corresponding model: (a) LSSB, (b) HSS, (c) PIGR, and (d) TAB.
In addition to the visual depiction, three parameters have been considered to evaluate the performance of four different prediction models: RMSE, MAE, and MAPE. Table 3 lists the detailed tendency prediction results. It can be seen from Table 3 that RMSE, MAE, and MAPE of all prediction results are slightly lower, while the accuracy rate is higher, which means that the proposed method can accurately rely on past degradation patterns to project future status of WT gearbox. After the verification of different prediction models, it is observed that LS-SVR is a suitable model for tendency prediction of WT gearbox.
Test results of four models in terms of RMSE, MAE, and MAPE.
RMSE: root mean square error; MAE: mean absolute error; MAPE: mean absolute percentage error; LSSB: low-speed shaft bearing; HSS: high-speed shaft; PIGR: planet inner gear ring; TAB: torque arm and bearing.
To illustrate the essentiality of transfer learning strategy with the change in operating condition, fourfold cross-validation tendency predictions are carried out for a comparison, that is, one of four prediction models is utilized to perform the prediction of other three tendencies. Figure 14 depicts actual HIs (blue line) and predicted HIs (red line), highlighting the performance of different cross-validations. It can be intuitively observed that figures distributing on the diagonal line (Figure 14(a), (f), (k, and (p)) show the good prediction performance, while the performance of other models deteriorates in a sharp drop owing to inconsistent data distribution between different vibration signals.

Cross-validation tendency prediction results based on LS-SVR.
To quantitatively evaluate the predicted accuracy of tendency prediction models, three indexes mentioned above are calculated for a comparison. Table 4 lists the detailed information of these models. Furthermore, prediction error comparison histogram of cross-validation is illustrated in Figure 15. From Table 4 and Figure 15, two conclusions can be summed up. First, the tendency predicted using same prediction model has a high fitness. This validates the advantage of regression model for tendency prediction. Second, the performance of prediction model deteriorates sharply with the change in prediction target. The main reason to explain the deterioration is the inconsistent data distribution of different monitoring signals mounted at different locations. Consequently, considering the potential correlation between different monitoring signals generated from different measuring locations, 56 a transfer learning strategy is demanded to adapt data distribution difference.
Prediction accuracy of cross-validation tendency prediction.
LSSB: low-speed shaft bearing; HSS: high-speed shaft; PIGR: planet inner gear ring; TAB: torque arm and bearing; RMSE: root mean square error; MAE: mean absolute error; MAPE: mean absolute percentage error.

Prediction error comparison histogram of cross-validation: (a) RMSE, (b) MAE, and (c) MAPE.
To overcome the drawbacks of conventional prediction model under the multi-sensor measurement, we utilize transfer learning to adapt data distribution between different signals generated from different measuring locations. In the proposed prediction process, cross-validation prediction accuracy is utilized to explore the proposed WF-MMD-JDA transfer learning ability. For the better explanation, Figure 16(a) is the prediction result of vibration signal LSSB integrated from vibration signals HSS, PIGR, and TAB, that is, the vibration signal LSSB is taken as unidentified data, the other three vibration signals (HSS, PIGR, and TAB) are considered as the huge source domain. Then, data distribution difference can be reduced by equation (16). Moreover, WF-MMD-JDA is utilized to eliminate distribution difference. After obtaining weight values of different models compared with unknown monitoring signal, the integrated regression model of unknown monitoring signal can be carried out. Figure 13 shows the prediction results of cross-validation based on WF-MMD-JDA. From Figure 16, we can intuitively observe that WF-MMD-JDA has the ability to conduct the multi-source unknown tendency prediction of WT gearbox.

Tendency prediction results based on WF-MMD-JDA: (a) LSSB, (b) HSS, (c) PIGR, and (d) TAB.
Furthermore, RMSE, MAE, and MAPE have been considered to evaluate the performance of each cross-prediction model. Table 5 lists the detailed prediction results of WF-MMD-JDA. Compared with the prediction results given by the conventional PHM model in Table 4, it validates the advantage of WF-MMD-JDA for multi-sensor tendency prediction.
Test results of WF-MMD-JDA in terms of RMSE, MAE, and MAPE.
WF: witness function; MMD: maximum mean discrepancy; JDA: joint distribution adaptation; RMSE: root mean square error; MAE: mean absolute error; MAPE: mean absolute percentage error; LSSB: low-speed shaft bearing; HSS: high-speed shaft; PIGR: planet inner gear ring; TAB: torque arm and bearing.
To validate the superiority of improved transfer learning method, three different integrated prediction models are employed to predict health tendency: (1) JDA, (2) MMD-JDA, and (3) WF-MMD-JDA. For above three algorithms, they all use the same regression prediction model (LS-SVR) to improve the accuracy of comparison. In the proposed comparison process, cross-validation prediction accuracy is considered to explore WF-MMD-JDA maximum adaptation ability. Figures 17 and 18 are the prediction result of integrated cross-validation based on JDA and MMD-JDA, respectively. It can be observed from Figures 16 to 18 that when the prediction target changes, the performance of prediction model deteriorates, and the prediction accuracy of using JDA migration learning is improved. Compared with the conventional PHM model (seen in Figure 14), all three prediction methods based on transfer learning (JDA) can predict future tendency of unknown monitoring vibration signal generated from WT gearbox. It validates the advantage of JDA for multi-sensor tendency prediction.

Tendency prediction results based on JDA: (a) LSSB, (b) HSS, (c) PIGR, and (d) TAB.

Tendency prediction results based on MMD-JDA: (a) LSSB, (b) HSS, (c) PIGR, and (d) TAB.
From Figures 17 and 18, we can intuitively observe that the accuracy in tendency prediction of the WT gearbox improves by adding the JDA migration learning strategy. Tables 6 to 9 list the detailed tendency prediction results of conventional PHM model (LS-SVR), JDA, MMD-JDA, and WF-MMD-JDA. Furthermore, prediction error comparison histograms of four prediction models are illustrated in Figure 16. From Tables 6 to 9 and Figure 19, the following conclusions can be summarized: (1) the prediction model based on LS-SVR performs worse in comparison with the other models both in terms of accuracy and precision; (2) the prediction results of transfer learning are improved considering the potential correlation between different monitoring signals; (3) owing to the further reduction of distribution difference, the performance of MMD-JDA is better than JDA; and (4) nevertheless, the prediction results of proposed WF-MMD-JDA have the highest accuracy because it further increases the adaptation between identified and unidentified vibration signals. After the comparison of different tendency prediction models, it is observed that WF-MMD-JDA obtains the best model for tendency prediction of WT gearboxes. To sum up, the robustness of proposed prediction model is improved by migration learning strategy and multi-source domain integrated learning mechanism. Meanwhile, it should also be informed that collecting large volumes of heterogeneous data from multiple sensors on-board (like condition monitoring system [CMS], SCADA, and oil data) of WTs might be feasible in the near or not so near future. Although it may require a big data storage technology, there is no need to perform model retraining of the same type WT gearboxes.
Test results of LSSB in terms of RMSE, MAE, and MAPE.
LSSB: low-speed shaft bearing; RMSE: root mean square error; MAE: mean absolute error; MAPE: mean absolute percentage error; LS-SVR: least square support vector regression; JDA: joint distribution adaptation; MMD: maximum mean discrepancy; WF: witness function.
Test results of HSS in terms of RMSE, MAE, and MAPE.
HSS: high-speed shaft; RMSE: root mean square error; MAE: mean absolute error; MAPE: mean absolute percentage error; LS-SVR: least square support vector regression; JDA: joint distribution adaptation; MMD: maximum mean discrepancy; WF: witness function.
Test results of PIGR in terms of RMSE, MAE, and MAPE.
PIGR: planet inner gear ring; RMSE: root mean square error; MAE: mean absolute error; MAPE: mean absolute percentage error; LS-SVR: least square support vector regression; JDA: joint distribution adaptation; MMD: maximum mean discrepancy; WF: witness function.
Test results of TAB in terms of RMSE, MAE, and MAPE.
TAB: torque arm and bearing; RMSE: root mean square error; MAE: mean absolute error; MAPE: mean absolute percentage error; LS-SVR: least square support vector regression; JDA: joint distribution adaptation; MMD: maximum mean discrepancy; WF: witness function.

Error comparison histogram of four prediction models: (a) RMSE, (b) MAE, and (c) MAPE.
The main reason to explain good performance of our proposed method is the model based on distribution adaptability, which can solve the data interference due to the strong background noise of WT gearbox. Hence, it can further reduce incomplete adaptation caused by varying loads or speeds. However, there is nothing to do with it for JDA and MMD-JDA models. The conventional regression model only utilizes HI
Conclusion and future work
This article presents a novel transfer learning strategy and multi-sensor integrated learning mechanism for WT gearbox based on WF-MMD-JDA. Meanwhile, a useful statistical characteristic
However, owing to the insufficiency of data quantity, the proposed transfer learning method has only been applied to different vibration sensors on the same WT gearbox under one operating condition in this study, and therefore, future experiments would be done under multiple operating conditions between WTs, considering the potential vibration mechanism correlation. Meanwhile, experiments should be done on other rotating machinery in WT, such as main bearing, generator bearing, and pitch slewing bearing, to verify the generality of this method or to find the problems in generalizing.
Footnotes
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was funded by the Postgraduate Research & Practice Innovation Program of Jiangsu Province (KYCX17_0937) and the National Natural Science Foundation of China (51875273).
