Abstract
Mechanical fault diagnosis under time-varying conditions remains challenging due to speed fluctuations that induce signal nonstationarity. To address this issue, a multiscale correntropy matrix representation method is proposed. First, aligned Fourier decomposition is introduced to decompose multivariate sensor signals into structured multiscale components with consistent frequency alignment across sensors. Then, scale-aligned and scale-varied correntropy matrices are constructed to characterize cross-sensor coupling and within-sensor nonstationarity, respectively, capturing spatial coupling and nonstationary behavior. These matrices are mapped from a Riemannian manifold to Euclidean space via matrix logarithm operations and fused into a compact feature vector for joint signal representation. The effectiveness of the method is evaluated on bearing and gearbox datasets under time-varying and constant speed conditions. The results of comparative and ablation experiments demonstrate that the proposed approach achieves superior performance and practicality compared to existing methods.
Keywords
Highlights
A multiscale correntropy matrix representation is proposed for fault diagnosis under time-varying conditions.
Dual correntropy matrices model intersensor coupling and intrasensor nonstationarity.
Aligned Fourier decomposition enables structured multiscale representation of multivariate signals.
Introduction
Fault diagnosis of mechanical equipment is a critical technology for preventing production accidents.1–4 However, mechanical equipment frequently operates under time-varying conditions such as transitions between different speeds. These speed fluctuations modify the waveform and energy distribution of monitoring signals, making it difficult to establish stable relationships between signal characteristics and fault patterns.5,6 As a result, reliable fault diagnosis under time-varying conditions remains a significant challenge.
Existing approaches to address speed fluctuations can be broadly categorized into order tracking, deep learning-based (DL-based), and traditional machine learning-based (ML-based) methods. Order tracking transforms time-domain signals into the angular domain through equiangular resampling, aiming to produce stationary representations. 7 For example, Wang et al. 8 applied order tracking to stator current signals for eliminating the influence of turbine shaft speed variations. Schmidt et al. 9 combined short-time Fourier transform with a maximum likelihood frequency tracking procedure, where a probabilistic framework was used to estimate instantaneous frequencies. Su et al. 10 proposed a double estimation strategy for precise rotational speed reconstruction, in which Vold–Kalman order tracking is used to isolate the single-frequency component essential for phase demodulation. Zhang et al. 11 proposed an advanced self-tuning vibration modeling strategy, in which order tracking is used to convert nonstationary time-domain signals into stationary angular-domain signals for accurate extraction of instantaneous order components under speed fluctuations. Peng et al. 12 developed a kurtosis-based spectral index-guided polynomial approximation method, which adaptively optimizes speed estimation to suppress spectral smearing and accurately identify fault characteristic orders under fluctuating speed conditions. Despite their effectiveness, these methods rely heavily on accurate speed measurements, which are often unavailable in practical applications due to sensor limitations, installation constraints, or safety considerations.
With the rapid development of artificial intelligence, DL-based methods have gained prominence due to their end-to-end feature learning capabilities. Representative techniques include convolutional neural network (CNN), 13 generative adversarial network, 14 transformer, 15 and large language model. 16 For example, Peng et al. 17 designed a phoneme-inspired acoustic frame embedded lightweight transformer for bearing fault diagnosis, which effectively captures transient impulse patterns through a phoneme-like structural design while maintaining high computational efficiency. Beyond static scenarios, intelligent fault diagnosis under varying speed conditions has gained increasing attention over the last 4 years. In 2023, Liang et al. 18 presented a deep residual deformable convolution network with subdomain adaptation for fault recognition under time-varying speeds. In 2024, Chen et al. 19 developed a hybrid augmented network with a blind source window to diagnose bearing faults under abrupt speed variation. Dong et al. 20 proposed a multiscale supervised contrastive learning method that uses adaptive convolution and attention to diagnose bearing faults under variable speed conditions. In 2025, Shao et al. 21 proposed a multiscale graph attention model that achieves excellent performance on multiple fault datasets under variable-speed conditions. Qin et al. 22 established learnable wavelet-driven physically interpretable networks for bearing fault diagnosis under sharp speed transients. Li et al. 23 constructed a temporal cross-contrastive self-supervised learning framework for high-speed train bearing fault diagnosis under speed variability. In 2026, Kim et al. 24 proposed the multiscale signal transformer model to enable adaptive multiresolution feature learning for fault diagnosis under variable speed conditions. Wang et al. 25 propose a two-stage method that learns speed-invariant features from unlabeled varying speed data and then diagnoses faults accurately with very few labeled samples. Although these DL-based methods achieve promising performance, they face several limitations. First, they rely heavily on large amounts of labeled data and high computational resources, which are often impractical in industrial settings. Second, the internal mechanisms of deep networks are typically opaque, resulting in features that lack physical interpretability. This limits trust and acceptance among maintenance engineers who prefer transparent decision-making processes.
In contrast, less sophisticated ML-based methods require fewer annotated samples and offer greater transparency to facilitate their practical deployment in industrial settings. These methods typically follow a two-stage pipeline: feature extraction and fault classification. In the first stage, the time-domain, frequency-domain, and time–frequency domain features are usually extracted from the acquired sensor signals. 26 Subsequently, these features are fed into a classifier, such as support vector machines, logistic regression (LR), and Naive Bayes, to recognize different fault patterns. However, owing to nonstationary characteristics and operational variability of the acquired signals, the extracted features often lack stability across varying speed conditions, leading to degraded classification performance. According to a recent review, 27 the promising direction lies in developing advanced feature extraction techniques capable of capturing robust and speed-invariant representations from raw signals, thereby enhancing classifier reliability under time-varying conditions.
With the advancement of sensing technologies, monitoring sensors are now widely deployed across mechanical equipment, enabling simultaneous monitoring of diverse physical states and providing more comprehensive diagnostic information than single-sensor setups. 28 The resulting multivariate signals inherently exhibit two key characteristics: within-sensor nonstationarity and cross-sensor coupling. Within-sensor nonstationarity, arising from speed fluctuations and transient impacts, carries distinctive local fault signatures but poses significant challenges for reliable feature extraction. At the same time, cross-sensor coupling reflects the dynamic coherence among different sensing locations, offering valuable insights into fault propagation and system-level responses. Despite their complementary nature, existing methods often fail to simultaneously characterize both aspects within a unified framework. There remains a critical need for signal representations that can jointly model localized dynamics and spatial interactions under nonstationary conditions. To address this challenge, a multiscale correntropy matrix representation (MCMR) method is proposed. MCMR leverages a structured multiscale signal decomposition (MSD) and dual correntropy matrix modeling to jointly capture spatial coupling and nonstationary dynamics, thereby enabling accurate fault characterization under time-varying conditions. The main contributions of this article are summarized as follows:
A novel MCMR method is proposed for mechanical fault diagnosis under time-varying conditions. It effectively extracts and integrates discriminative multiscale features, thereby enhancing intraclass compactness and interclass separability in fault classification.
Scale-aligned and scale-varied correntropy matrices (SVCMs) are constructed to model cross-sensor coupling and within-sensor nonstationarity, enabling comprehensive fault characterization.
Aligned Fourier decomposition (AFD) is introduced to establish a structured multiscale representation of multivariate sensor signals, ensuring consistent frequency alignment across all sensors for reliable cross-sensor analysis.
The remainder of this article is structured as follows: The second section describes the proposed MCMR method, the required RR classifier, and the overall framework. The third section shows the comparative and ablation experiments. The fourth section summarizes this article.
Methodology
In this section, the proposed MCMR method for feature extraction is first described. Next, the excellent Ridge regression (RR) classifier for fault classification is introduced. Finally, the overall fault diagnosis framework is elaborated.
Proposed MCMR method
As illustrated in Figure 1, the MCMR method consists of two successive steps: MSD and dual correntropy matrix representation (DCMR). In the first stage, AFD is designed for multivariate sensor signals, decomposing them into scale-specific components with consistent frequency alignment across all sensors. In the second stage, scale-aligned correntropy matrices (SACMs) are constructed to capture cross-sensor coupling by computing correntropy across sensors at matched scales, while SVCMs are established to characterize within-sensor nonstationarity by computing correntropy across scales within each sensor. Being symmetric positive definite (SPD), these matrices reside on a Riemannian manifold and are mapped to Euclidean space via the matrix logarithm (logm) operation. After vectorization, they are concatenated into a compact feature vector that jointly represents the multivariate sensor signals.

Flowchart of the proposed MCMR method. MCMR: multiscale correntropy matrix representation.
Multiscale signal decomposition
To effectively characterize the complex, nonlinear, and nonstationary nature of multivariate sensor signals, we propose AFD, a framework grounded in Fourier theory. Within this framework, each sensor signal is decomposed into a set of Fourier intrinsic mode functions (FIMFs) through FD. These FIMFs are complete, locally adaptive, and can be efficiently obtained using a zero-phase filter bank. 29 Critically, a shared dyadic frequency partition is applied across all sensors, ensuring consistent spectral alignment and thereby enabling reliable cross-sensor correlation analysis.
The decomposition begins with the representation of each sensor signal
where
Theoretically, a finite-length signal can be expressed via Fourier series or the fast Fourier transform (FFT). Applying FFT yields the complex spectrum
The frequency axis
where
For each interval
where
Consequently, the
In discrete form, this becomes:
where
The AFD process is illustrated in Figure 2, where a shared, fixed frequency partition establishes a common spectral reference frame, ensuring that the
By treating the original signal as
where

The process of the AFD. AFD: aligned Fourier decomposition.
This aligned representation captures intrasensor multiscale dynamics and enables reliable intersensor correlations via a shared frequency basis.
Dual correntropy matrix representation
To reveal the underlying spatial relationships among sensors and the internal dynamic interactions within individual signals, the correntropy is adopted to quantify correlations across scale-specific signal components. Correntropy is an information-theoretic correlation measure that integrates kernel methods with statistical learning. Its key advantage lies in its locality: it effectively suppresses the influence of outliers distant from the data distribution center. Moreover, through nonlinear mapping into a reproducing kernel Hilbert space, correntropy can transform a nonlinear problem in the original space into a linear one, enabling robust analysis in the transformed domain. 30
For two random variables
where
where
In engineering applications, the correntropy is usually estimated from the finite-length signals
The correntropy with a Laplace kernel is symmetrical and admits the following Taylor series expansion:
In contrast to conventional second-order statistics (e.g., covariance), correntropy with the Laplace kernel is the summation of all order moments of the difference
Based on Equation (9), each scale-aligned column of the signal tensor can be denoted as
The principal diagonal of
Similarly, each sensor-specific row of the signal tensor is denoted as
SVCMs characterize intrasensor multiscale dependencies, revealing the internal nonstationary dynamics and self-similarity across frequency bands within a single-sensor signal.
SACMs and SVCMs are matrices with a SPD structure. They exist naturally in a nonflat Riemannian manifold space, which is a special differential manifold with a Riemannian metric. 31 The SACMs and SVCMs from the Riemannian manifold space can be mapped into the Euclidean metric space via the logm:
where
As illustrated in Figure 3, this mapping transforms the correntropy matrices such that their diagonal entries are no longer uniformly 1, thereby uncovering hidden global structural information and enriching the feature representation. Importantly, the resulting matrices remain symmetric. Consequently, the total number of unique elements across all mapped matrices is

The illustration of the logm operation. logm: matrix logarithm.
Next, the mapped SACMs and SVCMs are reshaped and fused into intersensor and intrasensor features, denoted as follows:
where
Finally, the complete feature representation is obtained by fusing inter- and intrasensor information:
Notably, every entry in the SACMs and SVCMs has a clear physical interpretation: it quantifies a generalized correlation between signal components. Incorporating the mathematically rigorous logm operation, the proposed method provides a transparent, interpretable, and physics-informed alternative approach to black-box DL models, making it particularly suitable for safety-critical engineering applications.
RR classifier-based fault classification
RR is a regularized variant of linear regression that introduces an L2 penalty term into the loss function to control model complexity and mitigate overfitting.
32
Owing to its simplicity and robustness, RR has been widely adopted in classification tasks. Given a training dataset
where
The optimal weight matrix
where
Notably, the RR formulation in Equation (20) provides a unified framework for both binary and multi-class classification. Moreover, its closed-form solution can be computed directly via linear algebra operations, avoiding the need for iterative quadratic programming as required in support vector machines. It is worth emphasizing that the regularization term is essential to ensure numerical stability, particularly when the number of training samples is smaller than the number of features, and the regularization parameter
Overall framework of the proposed method
To address the challenges faced in mechanical equipment fault diagnosis under speed fluctuations, this study proposes the MCMR method for extracting sensitive fault features to identify the health states of mechanical equipment. The overall framework of the proposed method is depicted in Figure 4, which includes the following procedures:
(1) Data acquisition: The data collection systems acquire monitoring signals from sensors installed on the mechanical equipment.
(2) Data segmentation: The acquired signals are segmented with an adequate length into a sample set by using a sliding window.
(3) Data preprocessing: To eliminate the influence of amplitudes on the correntropy calculations, the raw signals within the sample set are normalized by using Z-score normalization, scaling them to have a mean of 0 and a variance of 1. This process can be represented as follows:
where
(4) Feature extraction: The preprocessed sample set is input into the proposed MCMR method to generate a corresponding feature set.
(5) Classifier training: The feature set is randomly split into a train set and a test set. The train set is directly fed into the RR classifier for training.
(6) Diagnosis testing: Ultimately, the test set is input into the fully trained RR classifier to generate the diagnosis results.

Overall framework of the proposed method.
Experimental study
In this section, the hyperparameters of the MCMR are first analyzed and selected. Subsequently, comparative experiments are carried out to verify its effectiveness and superiority, and a series of ablation experiments are implemented to demonstrate the necessity and significance of its components.
Experimental setup and data description
To demonstrate the effectiveness of the MCMR, two datasets are introduced: a bearing dataset from the Korea Advanced Institute of Science and Technology and a gearbox dataset designed and established by Huazhong University of Science and Technology in collaboration with Weite Technologies Company.
Bearing dataset
The bearing fault simulation test rig comprises a brake, two bearings, a mass wheel, a gearbox, a torque meter, and a motor. 33 Electrical discharge machining was used to machine the bearings to generate the outer ring fault, inner ring fault, and ball fault. Including the normal state, there were four health states in total. Two accelerometers were mounted horizontally and vertically on the bearing housings to collect vibration signals at a sampling frequency of 25,600 Hz.
The test rig simulated time-varying conditions by adjusting the motor speed, which varied randomly in the range of 680–2460 revolutions per minute (rpm). The speeds for all health states in the first 50 s are shown in Figure 5, which demonstrates that the speed changes dramatically over time. In this article, the data from the first 50 s are considered to be representative and are used for subsequent experiments. In addition to time-varying conditions, the test rig also simulated constant speed conditions at 3010 rpm.

Speed variations on the bearing dataset: (a) inner ring fault, (b) outer ring fault, (c) normal state, and (d) ball fault.
Gearbox dataset
The gearbox fault simulation test rig was built to simulate time-varying conditions. As shown in Figure 6, the test rig primarily consists of a drive motor, a healthy gearbox, a fault simulation gearbox, and a loading motor. Three vibration sensors were mounted on the fault simulation gearbox, and an acoustic sensor was fixed nearby to collect vibration and sound signals from different directions, respectively. The sampling frequency was set to 20,000 Hz.

(a) The gearbox fault simulation test rig, (b) the locations of monitoring sensors, (c) the internal structure of the fault simulation gearbox, (d) pitted gear, (e) broken tooth, and (f) damaged bearing outer ring.
Time-varying conditions were simulated by adjusting the speed of the drive motor. The experiments were conducted over three periods, each lasting 40 s. Specifically, each period consisted of a constant speed condition at 200 rpm for the first 10 s, 1000 rpm from the 20th to 30th second, a time-varying condition with a linear speed increase from the 10th to 20th second, and a linear speed decrease from the 30th to 40th second. The acquired dataset encompasses four single health states (normal state (N), pitted gear (P), broken tooth (B), and damaged bearing outer ring (D)) and two compound health states (pitted gear and damaged bearing outer ring (PD), broken tooth and damaged bearing outer ring (BD)). All faults were artificially produced.
As shown in Figure 7, the time domain waveforms of the four collected sensor signals for the pitted gear vary dramatically with time. In the case of sensor 3, the amplitude of the signal changes dramatically from about 2 m/s2 to about 30 m/s2 when the speed increases from 200 to 1000 rpm. In addition to time-varying conditions, the experiment also simulated constant speed conditions at 400 rpm. To visually demonstrate the consistency across heterogeneous sensors, Figure 8 presents the time-domain waveforms and corresponding FFT spectra of a representative sample containing four monitoring sensors for the pitted gear. It can be observed that while the raw amplitudes differ significantly between the vibration sensors and the acoustic sensor, their frequency domain representations exhibit aligned characteristic peaks, particularly in the resonant bands. This spectral consistency suggests that, despite the different sensing mechanisms, all sensors capture consistent fault-related features across various transmission paths. Such similarity confirms that heterogeneous sensors exhibit highly correlated fault signatures, which provides an empirical foundation for MCMR to jointly model cross-sensor coupling and within-sensor nonstationarity.

Time domain waveforms for pitted gear on the gearbox dataset: (a) sensor 1, (b) acoustic sensor, (c) sensor 2, and (d) sensor 3.

Time-domain waveforms and corresponding FFT spectra: (a) sensor 1, (b) acoustic sensor, (c) sensor 2, and (d) sensor 3. FFT: fast Fourier transform.
Hyperparameter selection
This section aims to explore the impact of hyperparameters and select them. Hyperparameters include the K value, which represents the number of FIFMs generated by AFD, the L value, which indicates the signal length in the sample set, and the kernel size
Selection of the number K
The number K is crucial since the MCMR generates

Accuracies for different hyperparameters: (a) K, (b) L, and (c)
Selection of the length L
Although the periodic pattern of the signal waveform in time domain no longer intuitively appears due to speed fluctuations, the number of signal points for one revolution can still be calculated by using the following formula:
where
As mentioned above, the maximum speed is 2460 rpm, and the minimum speed is 680 rpm. After calculation, the maximum L ≈ 2259 and the minimum L ≈ 626. In general, a shorter signal fails to comprehensively capture the signal characteristics, while a longer signal fails to satisfy the practical engineering requirement for minimal training data. According to the previous section, K is set to 9. Set L from 300 to 4500 with an interval of 300. As shown in Figure 9(b), the average accuracy significantly increases as L is larger when L < 2259. As L continues to rise, there is only a slight increase in accuracy. It indicates that the L value should be at least one period larger. Therefore, we set L as 2700 to balance accuracy and data requirements.
Selection of the kernel size
Following the above two subsections, let L = 2700 and K = 9, and the kernel size
According to the above analysis, the hyperparameters for the bearing dataset under time-varying conditions are determined. For the gearbox dataset, the hyperparameters can also be set by using the same steps. To avoid repetitive discussions, L on the gearbox dataset is set to 6000 (the length for a signal period at the minimum speed), K is set to 5, and the
Comparative experiments
Comparative experiments under different operating conditions, with varying signal-to-noise ratios (SNRs), involving other methods, and using different classifiers, are implemented to verify the effectiveness and superiority of the MCMR. The experiments in this section are based on the gearbox dataset.
Comparative experiments under different operating conditions
The feature representation ability of the MCMR under both time-varying and constant speed conditions is verified by incrementally increasing the number of training samples per class from 1 to 40. Figure 10 illustrates the diagnostic accuracies under these conditions, along with the gap values between them. The MCMR achieves an average accuracy of 97.75% when trained with 20 samples per class, demonstrating its exceptional performance under time-varying conditions. This success is fundamentally attributed to the robustness of the fault features extracted by the MCMR, which remains unaffected by speed variations and effectively captures the unique characteristics of different fault types, ensuring excellent interclass separability. Notably, increasing the number of training samples per class enhances diagnostic performance under time-varying conditions and reduces the gap with the performance observed under constant speeds. Specifically, as the number of training samples per class increases from 1 to 40, the gap value decreases significantly from 28.41 to 1.13%. These gap values arise from speed variations, emphasizing the greater complexity of fault diagnosis under time-varying conditions compared to constant speed scenarios. It is worth highlighting that under constant speed conditions, the MCMR achieves an average accuracy of 99.48% with only seven training samples per class, underscoring its suitability for few-shot fault diagnosis in such settings.

Experimental results under different operating conditions with varying numbers of training samples.
Comparative experiments with varying SNRs
In real-world industrial scenarios, measured signals are often heavily contaminated by noise. To evaluate the immunity and robustness of MCMR under such conditions, Gaussian white noise is added to the original signal of each sensor at various SNRs. The SNR is defined as follows:
where
The performance of the proposed method is evaluated using 40 labeled samples per class under time-varying and constant speed conditions, with results summarized in Figure 11. As shown, MCMR achieves higher diagnostic accuracy on time-varying speed data than on constant-speed data under low-SNR conditions (≤12 dB), with a maximum improvement of 15.29% at SNR = 6 dB. This is likely because the time-varying condition acts as an implicit form of data augmentation, enriching fault representations and enhancing robustness to noise. In contrast, under high-SNR conditions (≥15 dB), the stable and repeatable signal patterns in constant speed operation enable slightly better performance.

Experimental results under different operating conditions with varying SNRs. SNR: signal-to-noise ratio.
Comparative experiments involving other methods
To further validate the superiority of the proposed MCMR method, we conduct comparative experiments against eight recent fault diagnosis approaches published within the last three years. These include two ML-based methods (continuous wavelet transform-K-nearest neighbor (CWT-KNN) 34 and short-time Fourier transform (STFT)-support vector machine (SVM) 35 ), four DL-based methods (enhanced wavelet shrinkage network (EWSN), 36 lightweight convolutional transformer (LiConvFormer), 37 wide-kernel convolution Mamba (WCamba), 38 and multi-level fork network (MLFork) 39 ), and two few-shot learning (FSL)-based methods (multiple local-global correlation network (MLGCN) 40 and selective channel Mamba-based few-shot learning (SC-MambaFew) 41 ). A detailed description of these methods is provided in Table 1.
Description of comparison methods.
ML: machine learning; CWL: continuous wavelet transform; KNN: K-nearest neighbor; FSL: few-shot learning; STFT: short-time Fourier transform; CNN: convolutional neural network.
For a fair comparison, ML-based methods concatenate the features extracted from each sensor in the multisensor data to form a unified input vector. DL and FSL models are configured to fuse multisensor signals along the channel dimension and trained under identical experimental settings: a maximum of 120 training epochs, a batch size of 16, and a learning rate of 0.001. Training continues until convergence. All models are implemented in Python 3.8 and Lingo 18.0, and executed on a workstation equipped with an AMD Ryzen 5 5600G CPU (3.9 GHz) and 16 GB RAM. In addition, the hyperparameters for all baseline models strictly adhere to the configurations recommended in their original papers. The performance of each method is evaluated using 20 labeled samples per class, with results summarized in Figure 12.

Experimental results for comparison methods: (a) time-varying conditions and (b) constant speed conditions.
Under time-varying conditions, the DL-based methods achieve a maximum accuracy of 88.89% and a minimum of 64.39%. In contrast, FSL-based methods perform significantly better, with accuracies ranging from 91.10 to 94.69%. The proposed MCMR achieves the highest performance, with a maximum accuracy of 98.91% and a minimum of 95.35%. This represents an improvement of at least 6.46% over DL-based methods and 0.66% over FSL-based counterparts. Under constant-speed conditions, DL-based methods show improved performance, with accuracies between 81.38 and 97.90%. FSL-based methods achieve near-perfect or perfect classification, with accuracies ranging from 98.78 to 100.0%. Notably, MLGCN achieves 100% accuracy in this scenario, matching the performance of MCMR. However, this equivalence is limited to stable operating conditions. As demonstrated in the time-varying case, MCMR maintains robust performance under more challenging and dynamic environments. To further assess MCMR’s feature extraction capability under extremely limited data, we compare it against the three best-performing baselines: LiConvFormer, MLGCN, and SC-MambaFew. The comparison is conducted across a range of sample sizes, specifically 1, 3, 5, 7, 10, 15, and 20 labeled samples per class. As shown in Figure 13, MCMR consistently outperforms all competitors across all sample sizes, further confirming its effectiveness in data-scarce scenarios.

Experimental results for different methods under smaller sample sizes: (a) time-varying conditions and (b) constant speed conditions.
The relatively poor performance of DL-based methods under time-varying conditions stems from their reliance on large amounts of labeled data to learn speed-invariant features, which are inherently scarce in few-shot settings. Although FSL-based methods such as MLGCN and SC-MambaFew alleviate this limitation to some extent, they still exhibit restricted generalization because their architectures do not explicitly account for the influence of speed variations. In contrast, MCMR explicitly models intersensor spatial coupling and intrasensor nonstationary dynamics, enabling it to adapt effectively to varying operating speeds and achieve consistently superior diagnostic accuracy.
In addition, the computational time required for each method to achieve its highest accuracy is further evaluated. For the comparative methods, this time includes both model training and testing, whereas for CWT-KNN, STFT-SVM, and MCMR, it comprises feature extraction and fault classification. Moreover, the computational time of MSD and DCMR is recorded to evaluate their efficiency.
As shown in Table 2, MCMR requires significantly less time than all other methods. Moreover, MCMR exhibits nearly consistent computational time across different operating conditions. In contrast, DL and FSL models require substantially more time to converge under time-varying conditions than under constant-speed conditions. These results demonstrate that MCMR not only achieves higher diagnostic accuracy but also provides superior computational efficiency and model stability. It can be observed that MSD requires only 0.62 s, which is significantly faster than DCMR. This difference arises because MSD relies on AFD, which involves only highly efficient FFT and IFFT operations with near-linear complexity. In contrast, DCMR first constructs correntropy matrices and then applies a logm operation to each of them. The computational cost of constructing a correntropy matrix scales quadratically with the matrix dimension, while the logm operation scales approximately cubically with the matrix dimension, making the overall procedure considerably more expensive.
Computational time required for comparison methods.
CWL: continuous wavelet transform; KNN: K-nearest neighbor; FSL: few-shot learning; STFT: short-time Fourier transform; MSD: multiscale signal decomposition; DCMR: dual correntropy matrix representation; MCMR: multiscale correntropy matrix representation.
Comparative experiments using different classifiers
To further validate the effectiveness of the MCMR, other classifiers are combined with it, including linear support vector machine (LSVM), naive Bayesian (NB), linear discriminant analysis (LDA), and LR. They are used to replace RR as the final classifier. Figure 14 shows the experimental results of all classifiers under time-varying conditions and constant speed conditions. Compared with RR, other classifiers perform terribly, which is attributed to the inherent characteristics of the classifiers themselves. For example, when using NB, the MCMR method achieves an average accuracy of 99.11 ± 0.32% under constant speed conditions, whereas under time-varying conditions, the accuracy drops to 81.68 ± 2.12%. The NB performs the worst because it treats each feature equally and assumes independence among them. It is often unreasonable since the extracted features originate from strongly nonstationary and nonlinear signals. Compared with LR and LDA, RR includes an L2 regularization term, making it more suitable for addressing issues such as model overfitting, high feature dimensionality, and small sample sizes. LSVM also has a similar ability to handle these problems and thus achieves second-best accuracies. Through the above analysis, RR is chosen as the classifier due to its superior performance.

Experimental results using different classifier: (a) time-varying conditions and (b) constant speed conditions.
Ablation experiments
Ablation experiments, including fusion strategy, multiscale decomposition method, and correlation measure method with matrix logarithmic operation, are carried out to validate the significance and necessity of corresponding components in the MCMR. The experiments in this section are based on the bearing dataset.
Ablation experiments for fusion strategy
MCMR constructs a high-dimensional feature vector by concatenating correntropy features extracted both within each sensor and across sensors. To investigate the impact of fusion strategy on diagnostic performance, this subsection compares feature concatenation with feature averaging under varying numbers of sensors. In the feature averaging, the feature matrices from all SVCMs and SACMs are first summed and then averaged. As a result, the dimensionality of the intrasensor features reduces to
As shown in Figure 15, feature averaging leads to a significant drop in diagnostic accuracy compared to feature concatenation. Notably, for intersensor features with S = 4, the accuracy of feature averaging is 31.87% lower than that achieved by feature concatenation. This clearly indicates that, despite resulting in higher-dimensional representations, feature concatenation preserves more discriminative information and is therefore a more effective fusion strategy in the MCMR.

Ablation experiments for fusion strategy: (a) feature concatenation and (b) feature averaging.
Furthermore, the first subfigure reveals that when S = 2, 3, or 4, the average accuracies obtained using either intersensor or intrasensor features alone are consistently lower than those of MCMR. Take S = 3 as an example, the average accuracy of intersensor features (83.35%) and intrasensor features (85.04%) are lower than that of MCMR (89.14%). These results confirm that jointly modeling intrasensor nonstationarity and intersensor coupling is essential for accurate fault diagnosis. Moreover, the performance of the MCMR using multisensor data outperforms that of using a single sensor, which proves that multiple sensors can provide more comprehensive equipment information than individual sensors alone.
Ablation experiments for signal decomposition method
As a signal decomposition method, AFD is vital as it produces adequate scale-specific signal components to facilitate the extraction of correlation information. This subsection evaluates its impact on MCMR performance. Besides AFD, other signal analysis methods that can control the number of decomposed components can also be utilized for signal decomposition. Therefore, empirical wavelet transform, 42 empirical Fourier decomposition, 43 time-varying filter-based empirical mode decomposition (TVF-EMD), 44 variational mode decomposition (VMD), 45 and random Fourier decomposition (RFD) are used as comparative methods. Notably, RFD replaces the dyadic cutoff frequencies in AFD with randomly generated frequency band partitions. These methods are substituted for the AFD, and the number of decomposed components is kept the same. Other parts of the MCMR remain unchanged. Additionally, AFD is removed to validate its necessity, where only the correntropy matrix of the original multisensor signals is computed. Similarly, Z-score normalization is removed to assess its impact on MCMR performance.
The average accuracy and feature extraction time for each sample are recorded in Figure 16. While adaptive decomposition methods such as TVF-EMD and VMD achieve competitive accuracies of 93.13 and 88.94%, respectively, they remain inferior to the 96.63% achieved by MCMR. These adaptive methods often suffer from scale misalignment, as components from different sensors may not correspond to the same frequency bands. In contrast, the structured dyadic partitioning of AFD ensures consistent spectral alignment. Quantitatively, MCMR also demonstrates superior efficiency by requiring only 0.07 s per sample. This is significantly faster than TVF-EMD at 19.12 s and VMD at 0.87 s. Although removing AFD reduces the computation time to 0.01 s, the resulting accuracy drops drastically to 51.61%. This result confirms that multiscale decomposition is indispensable for capturing discriminative fault signatures. Furthermore, the slight decrease in accuracy observed after removing Z-score normalization suggests that this process is beneficial for suppressing the influence of speed-induced amplitude fluctuations.

Ablation experiments for signal decomposition method.
Ablation experiments for correlation measure method with matrix logarithmic operation
Correntropy and the matrix logarithmic operation are indispensable due to their excellent performance. In addition to correntropy, many other methods can measure the correlation between two signals, such as covariance, Euclidean distance, Minkowski distance, Pearson’s correlation, and cosine distance. Pearson’s correlation is the standardized covariance, and both can be used to measure the trend consistency between signals. The larger the value, the more consistent the trend changes, indicating a stronger correlation. The distance between signals can be used to express correlation. The smaller the distance, the more correlated the signals are. These methods are grafted into MCMR and applied to replace the correntropy. To assess the impact of the kernel choice, the Laplace kernel is replaced by the standard Gaussian kernel with the kernel size held fixed, and the resulting variant is denoted as MCMR*.
Meanwhile, the matrix logarithmic operation is applied to all correlation matrices to evaluate their effectiveness. As shown in Figure 17, the I-value quantifies the improvement in diagnostic accuracy when mapping features from the Riemannian manifold space to the Euclidean metric space. It can be observed that, except for the covariance, all other correlation measures achieve good fault diagnosis performance, which validates the appropriateness of adopting correntropy as the correlation metric. Furthermore, classification accuracy in the Euclidean metric space is consistently higher than that in the Riemannian manifold space. To illustrate this visually, Figure 18 presents t-distributed stochastic neighbor embedding visualizations of the feature representations obtained by MCMR with and without the logm operation, where different colors correspond to different fault classes. The distinctness of the identified clusters is further validated through the Silhouette score, 46 providing a quantitative measure of how effectively each point aligns with its assigned cluster versus other categories. The embedded features of MCMR (with logm) exhibit markedly clearer class separation in the three-dimensional space compared to the version without logm. Quantitatively, this operation leads to a 6.64% boost in accuracy and an absolute increase of 0.1598 in the Silhouette score, underscoring the indispensability and effectiveness of the matrix logarithmic operation. Additionally, the blue and purple bars of MCMR* are consistently shorter than those of MCMR, indicating that the Laplace kernel yields superior performance compared to the standard Gaussian kernel.

Ablation experiments for correlation measure method with matrix logarithmic operation.

t-Distributed stochastic neighbor embedding feature visualization: (a) MCMR (without logm) and (b) MCMR (with logm). MCMR: multiscale correntropy matrix representation; logm: matrix logarithm.
Conclusion
This article introduces a MCMR method for mechanical fault diagnosis. The proposed method enables the extraction of discriminative multiscale features that effectively capture the unique characteristics of different fault types, even under the adverse effects of large speed fluctuations. The comparative experiments demonstrate the effectiveness and superiority of MCMR, while ablation experiments validate the significance and necessity of its components. The main conclusions are summarized as follows:
The proposed MCMR achieves superior fault diagnosis accuracy while requiring significantly fewer computational resources than DL-based methods under both time-varying and constant speed conditions. Furthermore, MCMR circumvents common issues associated with DL-based methods, such as the need for large amounts of annotated samples and poor interpretability.
The constructed scale-aligned and SVCMs enhance discriminative feature extraction and improve diagnostic accuracy by enabling joint modeling of cross-sensor coupling and within-sensor nonstationarity.
The introduced AFD has shown significant efficacy in enhancing model performance by establishing a structured multiscale representation of multivariate sensor signals.
In addition, various heterogeneous modal data, such as current signals and thermal images, are extensively used to monitor mechanical equipment in industrial scenarios. These modal data significantly differ from vibration and acoustic data in terms of signal characteristics, and thus, they may have weaker correlations with each other. Consequently, the extension of MCMR to current signals and thermal images will be explored in future work.
Footnotes
Funding
The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This research was supported in part by the National Natural Science Foundation of China under grant no. 52575113, the Interdisciplinary Research Program of Huazhong University of Science and Technology under grant no. 2024JCYJ028, and the Joint Training Fund Project of Hanjiang National Laboratory under the grant nos. LP2024057 and LP2024054.
Declaration of conflicting interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Data availability statement
Data is available on request from the authors.
