Abstract
Mental workload (MWL) refers to the cognitive effort required to perform tasks and plays a vital role in optimizing performance and maintaining well-being. Recent advances in machine learning (ML) and deep learning (DL), supported by various physiological sensors, have enabled promising methods for estimating MWL. This review examines 110 articles published between 2018 and April 2025, sourced from the Scopus and Web of Science (WoS) databases. These studies span diverse domains such as office environments, operational systems, and air traffic control. Techniques such as principal component analysis (PCA) and feature selection are commonly used to enhance data representation. Frequently utilized datasets include electroencephalogram (EEG), electrodermal activity (EDA), functional near-infrared spectroscopy (fNIRS), NASA Task Load Index (NASA-TLX), and Simultaneous Task EEG Workload (STEW). Results show the significance of EEG in real-world applications and the importance of optimized sensor placement to ensure data quality. Many studies also focus on tracking dynamic signal changes during human-machine interaction (HMI) tasks. Ensemble learning approaches that utilize DL models such as convolutional neural networks (CNNs) and long short-term memory (LSTM) networks are widely used, while newer ensemble architectures using recurrent neural networks (RNNs) and autoencoders are gaining attention for their potential to enhance accuracy. Future research may examine validation metrics and cross-domain comparisons for practical implementation.
Introduction
MWL – also referred to as cognitive workload – is an increasingly important topic in workplace research due to its significant influence on employee well-being, job satisfaction, and overall performance. It represents the cognitive load placed on an individual while completing a specific task, involving neurophysiological, perceptual, and motivational processes (Kesedzic et al., 2021; Teoh Yi Zhe & Keikhosrokiani, 2021). MWL is affected by various factors, including task complexity, individual abilities, environmental conditions, and social context (Kesedzic et al., 2021; Taori et al., 2022). MWL estimation has been widely applied across multiple domains, such as education, particularly in online learning and web browsing contexts (Mazher et al., 2017; Saeidi et al., 2021), transportation sectors involving vehicle driving and air traffic control (Bagheri & Power, 2022; Blanco et al., 2018; Meteier et al., 2022), and specialized fields requiring intense focus, like nuclear power plant engineering (Wu et al., 2020). Recently, its use has extended into healthcare and diagnostics, including applications related to cancer, depression, schizophrenia, and autism spectrum disorder (Watts et al., 2022; L. Zhang et al., 2017). Extensive research has shown that elevated MWL levels can lead to stress, burnout, and reduced productivity (Alberdi et al., 2016; Craik et al., 2019; Roy et al., 2019; Saeidi et al., 2021). Given the impact of MWL on employees, organizations need to understand and address these issues to create a supportive and healthy work environment. Effective strategies for managing MWL include reducing time pressure, simplifying tasks, and providing adequate training and resources. Overall, MWL is a complex and multifaceted issue that requires careful attention and management in the workplace. By taking steps to understand and address these issues, organizations can improve employee well-being, job satisfaction, and performance, thereby increasing productivity and success.
The most common techniques used to measure MWL are subjective techniques based on a subject’s perception and objective measures based on physiological responses (Gupta et al., 2021; Hoang et al., 2020). Subjective measures for MWL rely on self-report assessments, such as rating scales and questionnaires. This typically entails asking employees to self-report their perceived MWL levels for a given task or time interval. Examples of commonly used subjective measures include the NASA-TLX (Hart & Staveland, 1988) and Subjective Workload Assessment Technique (Reid & Nygren, 1988). An advantage of subjective measures is that they directly assess how employees perceive their workload. However, subjective measures are limited by self-report biases and inter-individual variability in workload perception and reporting tendencies (Asgher et al., 2020b; Becerra-Sánchez et al., 2020; DellrAgnola et al., 2022). Meanwhile, objective measures for MWL include employing physiological and behavioral data to deduce the MWL level experienced by subjects. Such measures encompass various techniques, including but not limited to the use of ECG for monitoring heart electrical activity, EMG for assessing skeletal muscle electrical activity, EEG for detecting brain electrical activity (Gonçales et al., 2021; Taori et al., 2022; Zammouri et al., 2018), Photoplethysmography (PPG) for tracking volumetric changes in blood flow and respiration rate sensors (Beh et al., 2023; Ding et al., 2020; Meteier et al., 2022), EDA for eye movement trackers, oxygen density in the brain, and skin surface temperature reading (Rim et al., 2020). These techniques enable a more objective and reliable assessment of MWL than subjective techniques because they measure physiological dynamic changes that cannot be controlled consciously, making them increasingly popular among researchers in recent years. However, they can be invasive and may not provide a complete picture of the MWL of a task (DellrAgnola et al., 2022; Ding et al., 2020; Taori et al., 2022).
Recent advances in Artificial Intelligence (AI) have opened new possibilities for MWL assessment. In particular, Machine Learning (ML) and Deep Learning (DL) approaches have demonstrated strong potential in analyzing complex physiological and behavioral signals, enabling more accurate and real-time estimation of MWL. These methods go beyond traditional subjective and objective techniques by automating feature extraction, enhancing classification, and allowing cross-domain applications. The use of AI techniques has emerged as a promising approach for detecting and predicting MWL in the workplace (Meteier et al., 2022; Yin et al., 2019). Compared with conventional subjective measures, the integration of AI techniques, such as ML and DL, with physiological, behavioral, and environmental data can provide organizations with a more objective and precise MWL estimation (Gonçales et al., 2021). Furthermore, accuracy in evaluating MWL levels presents new opportunities for understanding and addressing crucial performance and human well-being issues. By replacing conventional subjective techniques, such as self-report questionnaires, ML and DL algorithms can offer a more reliable MWL estimation to organizations.
This article presents a systematic review of AI algorithms, including ML and DL, applied to MWL estimation. It covers key research directions, techniques, and experimental findings. The paper is structured as follows: the Related Work section highlights the limitations of prior surveys and the need for a centralized MWL resource; the Methodology section outlines the inclusion criteria; the Results and Discussion section analyzes ML and DL techniques used in MWL estimation; and the Conclusion provides final insights and suggests directions for further research.
Related Work
The MWL topic has gained considerable attention recently. Furthermore, only a few surveys have covered a wide range of applications, strategies, and techniques, thereby remarkably influencing ongoing research in the MWL estimation. In this section, we review existing literature with a special focus on review articles that have examined ML and DL techniques to understand, evaluate, or estimate MWL.
The study in Alberdi et al. (2016) investigated the MWL of office workers and examined the use of multimodal measurements for developing automatic MWL estimation systems in office environments. They explored different modalities used in detection systems, such as physiological measurement, speech analysis, and facial expression recognition, and evaluated the performance of various algorithms used for MWL estimation. Moreover, they discussed the potential applications of automatic MWL estimation systems in office environments and ethical considerations associated with their use. Their review covered literature from 2004 to 2015 and incorporated data from three databases, namely Compendex, Inspec, and PubMed.
Another specific survey paper delved into the applications of ML and DL techniques for estimating MWL (Zhou et al., 2022). In this case, the authors focused on comprehending MWL, which assesses the mental effort individuals invest in completing various work settings. They employed various conventional ML and DL techniques that use EEG to capture brain signals. Their review encompassed the examination of various approaches used for estimating MWL, highlighting remarkable advancements, and explored techniques for feature selection, categorization, and evaluation. Although specific databases were not mentioned for the reviewed works, this review covered research conducted up to 2020.
Recently, Khan et al. (2024c) examined the use of ML and DL techniques for classifying MWL using fNIRS signals in research published between 2011 and 2023. The authors reviewed literature sourced from databases including ACM, WoS, PubMed, IEEE Xplore, Scopus, Google Scholar, and EuropePMC. They summarized existing ML and DL methods applied to fNIRS data, explaining the foundational concepts and processing pipelines involved in these approaches. Additionally, the review discussed the performance metrics employed across studies and offered recommendations for future research aimed at enhancing model robustness and generalizability.
In a more specific context, (P. Wang et al., 2024) conducted a review examining the relationship between physiological signals, specifically heart rate variability (HRV), and MWL in pilots. The study highlighted the use of ML methods to classify and estimate MWL based on HRV data, using literature from 2000 to 2023 retrieved from databases: PubMed, Scopus, and WoS. The review provides insights into the efficacy of physiological measures in aviation contexts, highlighting the predictive value of HRV features for monitoring cognitive states in real time. Additionally, the paper discusses key performance metrics employed across studies.
In a similar systematic review, (Filipa Ferreira et al., 2024) explored ML and DL approaches for assessing various dimensions of MWL including attention load, fatigue, and alertness, using physiological signals derived from eye tracking. The study examined literature published in 2024, sourced from databases namely: WoS, ScienceDirect, Scopus, and PubMed. The review provided a comprehensive analysis of technical and methodological advances in gaze-based workload inference but omitted discussion of evaluation metrics.
In Vishnu and Gupta (2024), the authors presented a review of ML and DL techniques for EEG-based MWL estimation, with a specific focus on attention, engagement, and working memory. The review aimed to identify experimental paradigms, feature extraction methods, and classification models commonly used in EEG signal analysis for workload inference. It covered research published up to 2023, focusing on how EEG features can be harnessed to build robust approaches. However, the study did not specify or compare metrics used in the reviewed works.
In addition, Demirezen et al. (2024) focused on establishing best practices for reproducible ML studies in MWL estimation using EEG signals. The review gathered literature from 2024, sourced from multiple databases, including Scopus, WoS, ACM Digital Library, and PubMed. By outlining key reproducibility guidelines, the study supports the advancement of reliable ML-based MWL assessment tools that can be validated and compared across studies.
In a more recent study, Kingphai and Moshfeghi (2025) investigated the use of DL techniques to analyze EEG signals for MWL assessment. The review explored both the potential and challenges of DL methods in interpreting brain activity data, synthesizing findings from multiple databases, including ACM, IEEE Xplore, ScienceDirect, Scopus, SpringerLink, and Wiley Online Library. While the paper highlights relevant performance metrics, it does not specify the timeframe of the reviewed literature.
Table 1 compares our review with previous studies on ML and DL techniques for MWL estimation. The comparison covers key aspects such as targeted MWL areas, detection methods, associated challenges, datasets used, and future work suggestions. Furthermore, the table specifies whether each study explicitly focused on ML, DL, or both, thereby clarifying the methodological scope of prior works. The analysis reveals that while prior reviews addressed some of these aspects, none provided a comprehensive overview. Blank cells denote missing information, while checkmarks (✓) indicate adequately covered areas.
Comparison Between This Review and Existing Literature Reviews.
The main objective of this paper is to extend previous studies in the MWL estimation to fill gaps in the existing literature. Therefore, the important contributions of this review are as follows.
We highlight several novel studies that have not been covered in previous reviews. Additionally, we discuss both ML and DL techniques together and their application in MWL estimation.
We address the limitation of prior works that focused solely on either contextual data or EEG by discussing a range of dataset sources and highlighting commonly used datasets in the literature.
We review widely used ML and DL techniques for MWL estimation and assessment, emphasizing their potential to enhance estimation accuracy.
Finally, we summarize open challenges and propose future research directions for ML and DL techniques in MWL estimation, outlining areas for further investigation.
Methodology
In this study, we adhered to the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines as the standard for conducting systematic reviews and meta-analyses (Moher et al., 2009). In addition, our research method followed the Cochrane Collaboration definitions (Islam et al., 2021) to minimize the risk of bias.
Search Strategy
This systematic literature review aims to analyze previous studies that have focused on assessing MWL using DL and ML algorithms. This endeavor contributes to the existing body of knowledge in the domain of MWL estimation. To maintain high standards, our review concentrates on empirical research articles published in esteemed, peer-reviewed journals with rankings of Q1 or Q2, according to their respective disciplines. Key conference papers that met our criteria were also included in this review. Specifically, key conferences were defined based on the CORE Conference Ranking, with only conferences ranked A or A* considered. These venues are typically high-impact and highly selective in computer science and related fields. In addition, only conferences indexed in Scopus or WoS and relevant to artificial intelligence, biomedical engineering, or human-computer interaction were included. Meanwhile, we exclude book chapters, preprints, and papers ranked Q3 and Q4. The selection criteria are based on the understanding that leading journals play a pivotal role in advancing academia (Pereira et al., 2023). Their rigorous peer review process ensures the dissemination of high-quality research. Furthermore, our criteria are influenced by previous studies that emphasized the importance of focusing on publications in top-tier journals. For this study, two distinguished databases, Scopus and WoS, are used because of their extensive collection of relevant articles in healthcare, computer, management, business, economics, and related fields (Kranzbühler et al., 2018; Stumbitz et al., 2018). A combination of these two databases provides a comprehensive overview of journal articles in MWL estimation.
Article Selection
In this study, we utilized the PRISMA method, a systematic review and meta-analysis proposed by Moher et al. (2009), to identify and narrow down studies for inclusion in this review. The search strategy employed a systematic approach that encompassed two databases (Scopus and WoS) and all sections of each article, including titles, abstracts, and keywords along with predefined search criteria to identify relevant articles. Keywords, including the following, were selected to ensure a comprehensive and targeted search:
“Machine Learning” OR “Artificial Intelligence” OR “AI” OR “Image Processing”
“Prediction” OR “Detection” OR “Classification” OR “Identification”
“Mental workload” OR “Mental Work Stress.”
We initially identified 431 articles across the selected databases. To refine the collection for analysis and eliminate redundancy or irrelevant content, a careful screening process was implemented, incorporating specific inclusion and exclusion criteria. The process was as follows:
Removal of duplicate articles.
Identification and removal of irrelevant articles following the inclusion and exclusion criteria.
Inclusion of high-quality papers that have undergone a rigorous evaluation process.
The inclusion criteria are as follows:
Only published articles from 2018 to April 2025 are included.
Articles focusing on the use of ML and DL techniques in the field of MWL are considered.
Articles published in Q1 and Q2 peer-reviewed journals are included.
Articles from conferences ranked A* or A according to the CORE Conference Ranking and indexed in Scopus or WoS are included.
The exclusion criteria are as follows:
Articles related to MWL but not using ML or DL techniques are excluded.
Articles that use ML or DL techniques but do not focus on MWL are excluded.
Book chapters, preprints, and Q3- and Q4-ranked articles are excluded.
Articles from conferences ranked below A (e.g., B or C) or not indexed in Scopus/WoS are excluded.
Through a collaborative effort, we established the inclusion and exclusion criteria for articles. The final step involved a comprehensive assessment of the full texts of the retained articles against the established criteria. This stringent screening protocol ensured the inclusion of only high-quality studies that closely matched the goals of our systematic literature review, thereby contributing to the precision and validity of our analysis. We identified 110 published studies from 2018 to April 2025, and any concerns about specific studies were resolved through collaboration. A flowchart is shown in Figure 1 to illustrate the process of study inclusion and exclusion.

PRISMA study selection diagram.
Results and Discussion
This section synthesizes the reviewed literature on ML and DL techniques for MWL estimation by examining how the field has developed across datasets, feature engineering strategies, learning models, validation practices, and application settings. The discussion highlights the main directions of development, the shifts in methodological emphasis over time, and the remaining challenges that continue to shape this area of research.
MWL Estimation in the Literature
The reviewed literature shows a clear development in MWL estimation research. Early studies mainly focused on improving classification using handcrafted features and conventional ML models applied to single-modal data, especially EEG. More recent work has shifted toward DL, multimodal integration, and application-oriented systems that aim to support real-time and cross-context workload assessment. In this review, the literature is synthesized according to the main directions through which this progression can be observed.
Investigating Feature Engineering Techniques for MWL Estimation
Several studies have aimed to enhance feature engineering by selecting relevant features and applying innovative signal processing techniques. Research has evaluated the effectiveness of feature selection on EEG data (Becerra-Sánchez et al., 2020; Plechawska-Wojcik et al., 2018; Saadati et al., 2020; T. Wang et al., 2024) and leveraged various physiological sensors to capture stimulus-related responses over short time spans (Jimenez-Molina et al., 2018; Safari et al., 2024; G. Wang et al., 2022). Additionally, novel methods such as the fixed-value modified Beer-Lambert law (FV-MBLL) and U-Net architectures have been proposed to improve workload classification accuracy (Asgher et al., 2019; Xie et al., 2019).
Furthermore, P. Zhang et al. (2019a) and Cao et al. (2022) explored the fusion of different EEG and ECG features to detect MWL variations, examining various neural networks, such as two-stream neural networks (TSNNs) and CNNs for robust classification of MWL. Cao et al. (2022) proposed a framework for integrating hybrid EEG-fNIRS features with ML techniques to enable multilevel workload classification. Finally, the impact of specific band ratios and connectivity features on creating discriminative models for workload perception has been investigated (Hassan et al., 2024; Khan et al., 2024b; Park et al., 2024; Raufi & Longo, 2022).
Enhancing Current Approaches to MWL Estimation
To improve existing techniques for estimating MWL, studies across diverse domains focus on improving the classification accuracy of existing MWL estimation approaches. Research efforts are concentrated on improving supervised models, such as classification algorithms and diagnostic models, by leveraging ML with diverse inputs, such as EEG data, simulated cognitive behavior, and physiological indicators (Aksu & Çakıt, 2023; Appriou et al., 2018a; Elkin & Devabhaktuni, 2020; Lira et al., 2018; Pandey et al., 2020; Parent et al., 2019; Z. Zhang et al., 2023). Moreover, few studies have investigated the differentiation between task and rest periods during experiments, such as the Stroop task (Hassan et al., 2024; Ho et al., 2019), to improve MWL estimation during multitasking activities and identify hazardous conditions via MWL estimation (Das Chakladar et al., 2020; He et al., 2022; Siddiquee et al., 2020; Yiu et al., 2022). Furthermore, researchers have explored eye-tracking data analysis to enhance MWL estimation during memory tasks and driver-state assessment (Aksu et al., 2024; Aksu & Çakıt, 2023; Kothe et al., 2024; Oppelt et al., 2022). Finally, the potential of graph convolution networks and other advanced techniques has been explored to improve the classification of MWL and emotional statuses of workers (Chen et al., 2023; Z. Zhang et al., 2023).
Introducing and Investigating MWL Datasets
Recent studies have significantly advanced MWL estimation by both introducing and evaluating diverse datasets. Notably, Becerra-Sánchez et al. (2020) proposed GALoRIS, an EEG-based feature selection model, while Ktistakis et al. (2022) introduced COLET, a dedicated eye-tracking dataset. More recently, Chakraborty et al. (2024) and Hemakom et al. (2024) presented datasets leveraging physiological signals for improved multi-level stress and workload classification.
In parallel, numerous investigations have explored existing datasets across varied contexts using techniques like Near-infrared spectroscopy (NIRS), EEG, ECG, and eye tracking (Asgher et al., 2020a, 2020b; Le et al., 2018; Saikia et al., 2021; Wei et al., 2023). These datasets have been applied in driving simulations, ADAS systems, VR training, and outdoor deployments to assess MWL and provide real-time feedback (Aksu & Çakıt, 2023; Cardone et al., 2022; Hajra et al., 2020). Additionally, several studies have evaluated the effectiveness of physiological indices, self-assessments, and neural biomarkers in diverse simulated environments (Fawwaz et al., 2024; Kothe et al., 2024; Saadati et al., 2019; Zemla et al., 2023), emphasizing the need for multimodal data and cross-context validation (Figure 2).

Overview of an EEG-based ADAS system.
Proposing Novel MWL Estimation Approaches
Recent advances in MWL estimation have introduced various neural network architectures for analyzing EEG and fNIRS signals. These include recurrent three-dimensional CNN (R3DCNN), temporal segment networks (TSNN; P. Zhang et al., 2019b), extreme learning machines (ELM; Tao et al., 2019), transfer dynamical autoencoders (TDAE; Yin et al., 2019), three-dimensional CNN (3DCNN; Kwak et al., 2020), hybrid convolutional autoencoders with SVM classifiers (Islam et al., 2021), CNN-RNN combinations (Saadati et al., 2021), and deep CNN architectures (Dhulipalla & Al Hafiz Khan, 2022; Ding et al., 2020; Hernández-Sabaté et al., 2022). Additionally, novel frameworks incorporating ensemble learning and feature fusion techniques have been developed for multimodal EEG-fNIRS analysis (Ma et al., 2024; Moustafa & Longo, 2019; Saadati et al., 2021; Tao et al., 2019).
Investigating Automated MWL Estimation and Validation Techniques
Research on automated real-time MWL estimation has expanded beyond driving and brain-computer interfaces to various operational contexts. Approaches leveraging EEG signals have been proposed for MWL detection during tasks (Cheema et al., 2018), alongside cognitive human-machine interfaces supporting adaptive automation (Lim et al., 2021). ML algorithms enable real-time monitoring to inform operator replacement or resource allocation in scenarios such as search and rescue, while passive BCI systems allow simultaneous subject-specific and cross-subject classification of MWL and stress (DellrAgnola et al., 2022; Zhu et al., 2023). Concurrently, validation techniques have been refined to better accommodate physiological data characteristics; for example, Chihara et al. (2020) adapted cross-validation methods to align with the time-series nature of EEG signals, improving the evaluation of MWL estimation models.
Taken together, these research directions reflect a gradual shift in the field from improving isolated components of MWL estimation toward developing more integrated and adaptive systems. Earlier work placed greater emphasis on feature design and classifier selection, whereas later studies increasingly combined multimodal data, deep architectures, and validation-aware frameworks to improve robustness and practical applicability.
Datasets Used for MWL Estimation
Researchers have employed various datasets to investigate and quantify MWL in diverse contexts. The datasets used vary according to their collection method, purpose of use, preprocessing techniques, and the computational power required for any model to comprehend the dataset. Below are the widely used datasets in the covered literature.
EEG
Researchers have extensively used EEG data to study MWL. Sciaraffa et al. (2019) applied EEG-based ML and DL algorithms to classify workload in VR environments. To address imbalanced classes, feature oversampling techniques were explored (Ved & Yildirim, 2021; G. Wang et al., 2022). Hybrid frameworks combining EEG with fNIRS have been proposed for multilevel MWL estimation (Cao et al., 2022; S. Shao et al., 2024). Studies also evaluated various feature selection methods and classification accuracy using EEG (Dong et al., 2025; Plechawska-Wojcik et al., 2018). Novel neural network architectures, such as deep R3DCNN, TSNN, and CNN hybrids, have been developed to classify EEG features across tasks (Hernández-Sabaté et al., 2022; P. Zhang et al., 2019a, 2019b). Lastly, efforts to automate MWL estimation using EEG during operations have been pursued (Cheema et al., 2018).
EDA
EDA has been a key physiological signal for workload classification among various modalities. Jimenez-Molina et al. (2018) combined EDA with EEG, ECG, and PPG to effectively estimate MWL. Meteier et al. (2021) used EDA alongside ECG and respiration in a simulated automated driving context, complementing subjective NASA-TLX assessments. EDA is often part of comprehensive physiological measurements, including ECG, EMG, PPG, respiration, skin temperature, eye tracking, facial action units, reaction time, and questionnaire feedback for holistic workload evaluation (Oppelt et al., 2022). Wei et al. (2023) combined EDA and ECG with NASA-TLX annotations to understand driver workload. Mirzaeian and Ghaderyan (2023) proposed a novel method for detecting time-frequency changes in EDA using enhanced textural features and GLCM texture descriptors to improve MWL estimation.
ECG
ECG data are widely used to estimate MWL across contexts. Sakib et al. (2021) studied wearable devices collecting ECG and fNIRS simultaneously to design predictive models for general and specific populations, highlighting ECG’s role in combined physiological assessments (Kesedzic et al., 2021). ECG is also incorporated in datasets with EDA and respiration to investigate automated driving scenarios (Meteier et al., 2021) and forms part of broader assessments including EMG, PPG, skin temperature, eye tracking, facial action units, reaction time, and subjective feedback (Oppelt et al., 2022). Additionally, ECG alongside EEG has been evaluated for detecting pilot MWL and deriving heart rate and HRV measures, underscoring its importance in high-stress workload monitoring (Hajra et al., 2020).
EMG
In the literature, EMG has been used in a few applications. Recently, Park et al. (2024) focused on EMG-based prosthetic devices, incorporating varied input features such as eye-tracking measures, task performance, and cognitive performance model outcomes to estimate MWL. In addition, EMG was considered a part of a comprehensive array of physiological measurements, including ECG, EDA, PPG, respiration rate, skin temperature, eye tracker data, and behavioral metrics such as facial video-derived action units, reaction time, and subjective feedback via questionnaires (Oppelt et al., 2022).
fNIRS
fNIRS data have been used to capture brain hemodynamic activity during various task assessments (Ho et al., 2019). The related studies encompassed multiple subjects, employing fNIRS devices to record changes in the subjects’ hemoglobin concentration in the prefrontal cortex using multiple channels and wavelengths. The experiments involved tasks such as Stroop tasks, mental logic, or working memory assessments to elicit stress levels or evaluate cognitive workload. fNIRS signals, which measure oxyhemoglobin and deoxyhemoglobin, were analyzed using SVM, CNN, and DCNN for MWL estimation and brain signal processing (Asgher et al., 2020b). Hybrid EEG-fNIRS frameworks integrate ML features and bivariate functional brain connectivity for multilevel MWL estimation, while other studies have proposed the application of LSTM on fNIRS data to achieve the optimum accuracy in the MWL estimation level (Cao et al., 2022). fNIRS has been used as a valuable substitute for EEG, offering improved spatial resolution and requiring fewer protocols for brain signal acquisition compared to EEG (Asgher et al., 2020a; Khan et al., 2024a).
NASA-TLX
In literature, NASA-TLX has been used as a subjective MWL estimation tool to evaluate the perceived workload in various scenarios. Meteier et al. (2021) considered 90 subjects in the context of conditionally automated driving in a simulator. The NASA-TLX questionnaire was employed to subjectively assess the workload experienced during driving tasks while physiological signals such as ECG, EDA, and respiration were concurrently collected. Another investigation by Ktistakis et al. (2022) monitored participants’ eye movements and ratings based on the NASA RTLX workload index while they engaged in puzzles of varying complexity and duration, each differing in terms of time constraints and a secondary task. Further, in a driver-oriented context, sensors gathered physiological data such as ECG and EDA while drivers annotated the data using the NASA-TLX scale to assess their perceived workload during driving tasks (Gogna et al., 2024; Wei et al., 2023).
STEW Dataset
STEW has been used for MWL estimation and temporal dynamics modeling of EEG signals (Afzal et al., 2024). In addition, it was employed for offline EEG data analysis focusing on binary and multiclass classification of MWL. By proposing a pipeline interface specifically designed for estimating MWL in a continuous engagement or attention environment, the temporal dynamics within EEG signals were modeled (Afzal et al., 2024; Das Chakladar et al., 2020). By leveraging STEW’s diverse experimental tasks, encompassing scenarios without specific tasks and those involving multitasking activities graded by workload levels, scholars have sought to investigate and estimate MWL dynamics through EEG signals, enhancing insights into cognitive processes under varying task demands and multitasking complexities.
SWELL-KW Dataset
Questionnaire data and physiological sensor data have been investigated as effective indicators of MWL, with a focus on maintaining a manageable scope and optimizing relevant features for the extreme learning adaptive neuro-fuzzy inference system (ELANFIS; Teoh Yi Zhe & Keikhosrokiani, 2021; Z. Wang et al., 2024; Yin et al., 2024). To enhance estimation accuracy, the authors integrated particle swarm optimization into a microgenetic algorithm for predicting the MWL of knowledge workers using data from the SWELL-KW dataset.
Overall, the evolution of dataset usage indicates that the field has moved beyond reliance on single-source measurements toward broader multimodal and context-aware representations of workload. Although EEG remains the dominant modality, the increasing use of EDA, ECG, eye tracking, fNIRS, and subjective scales reflects an effort to capture MWL as a multidimensional phenomenon rather than a single physiological response.
Feature Engineering Techniques Used for MWL Datasets
Extensive research has been devoted to feature engineering techniques, highlighting their importance in the field. The following categories summarize the main directions of investigation.
PCA
Owing to the nature of the MWL dataset, PCA has been a useful technique because it helps in reducing dataset dimensionality and extracting new features. Ho et al. (2019) employed PCA as a preprocessing step to obtain a clean dataset termed “non-PCA inputs.” Following this, PCA was applied to generate inputs to enhance the robustness of the MWL estimation scheme. The primary objective of the study was to reduce generalization error and computational time and prevent overfitting, particularly when handling numerous features, thereby improving the efficiency of subsequent classifiers. Similarly, Yin et al. (2019) used PCA specifically for reducing feature dimensionality, thereby creating six hybrid classifiers: PCA-LSSVM, PCA-ELM, PCA-NB, PCA-LR, PCA-KNN, and PCA-ANN. Furthermore, principal components were determined based on a threshold of 0.9 for the total variance as it retains much information about the data, effectively reducing the feature space while preserving essential information for classification tasks. Another study proposed to use PCA as a feature selection technique alongside statistical analysis (Mohanavelu et al., 2022). Moreover, features were normalized using min-max normalization to ensure a consistent range of values for different features in the dataset. Similarly, the authors in Z. Zhang et al. (2023) used PCA for feature selection, emphasizing its role in reducing feature space while retaining the most informative components for subsequent modeling tasks. Furthermore, Chen et al. (2023) indirectly integrated PCA with RFE-SVM for feature selection, indicating that PCA can be included as part of a feature selection process.
Feature Selection
A variety of feature selection techniques have been explored in the literature to improve MWL estimation. In Plechawska-Wojcik et al. (2018), PCA was used for artifact rejection, followed by techniques such as K-Fisher, eigenvector centrality, and Mitinff’s mutual information-based method to identify the most informative features. These techniques were evaluated using two ML models, bagged decision trees (DT) and cubic SVM to determine the optimal feature set for modeling. In Lira et al. (2018), RFE-SVM was used as a baseline, and a new model was proposed that used Random Forest (RF) for feature selection. This method considered the cost of each feature, giving lower selection probability to features with a higher cost, thereby optimizing the process.
Random Forest combined with Recursive Feature Elimination (RF-RFE) is commonly employed, as demonstrated in Jimenez-Molina et al. (2018), which applied it after time-window-based feature extraction. Sequential Forward Selection (SFS) and the Relief algorithm were also investigated in Siddiquee et al. (2020), with both methods evaluated using a linear SVM and the feature set yielding the highest MWL estimation accuracy was selected. Additionally, Mohanavelu et al. (2022) introduced a hybrid approach combining PCA with statistical analysis, using min-max normalization to ensure feature uniformity. Similarly, Z. Zhang et al. (2023) applied PCA alongside feature selection to reduce dimensionality while retaining relevant information. In Chen et al. (2023), RFE-SVM was tested with varying training sample sizes to assess its effectiveness in identifying optimal features. Overall, techniques such as RFE, RF, SFS, Relief, and hybrid statistical methods are widely used to refine feature sets for MWL estimation models.
Feature Extraction
The study in Jimenez-Molina et al. (2018) used a feature extraction-based approach on time windows after standardizing signals for comparability. However, the data needed preparation before extracting important characteristics using RF-RFE as feature selection. Similarly, the preprocessing steps involved converting raw signals into hemoglobin values, filtering to remove noise, and finally applying PCA to extract new features before being fed to four classifiers in an experiment (Ho et al., 2019). Another comprehensive approach proposed by Shafiei et al. (2020) includes band-pass filtering, artifact removal, wavelet transform, and DFA to process EEG data and highlights the importance of preprocessing before feature extraction. Meanwhile, Islam et al. (2021) presented a unique approach using a CNN encoder and power spectral density for feature extraction, highlighting the advanced DL technique’s role in the context.
Feature Scale
Scaling the dataset features is another aspect of feature engineering that is considered an important preprocessing technique. In Cheema et al. (2018), the authors conducted feature normalization and optimization after aiming to scale features into a similar range to enhance the model performance and stability. Another similar approach proposed by Jimenez-Molina et al. (2018) highlighted the necessity of standardizing signals from different scales. Moreover, normalization techniques were employed by dividing signals either by their mean values or by the maximum value of window samples plus an average, to ensure consistency within the dataset and improve the convergence of ML and DL models (Asgher et al., 2020a; He et al., 2022; Saadati et al., 2020). Furthermore, the authors in Ktistakis et al. (2022) and Meteier et al. (2021) used MinMaxScaler and RobustScaler techniques for feature normalization, highlighting the robustness of these techniques in scaling data to optimize model performance and reduce computational complexity.
Across these studies, feature engineering evolved from manual preprocessing and dimensionality reduction toward more flexible combinations of feature selection, extraction, and normalization tailored to specific modalities and tasks. This trend also helps explain the later transition toward DL, where part of the representation learning process is handled automatically by the model itself.
ML and DL Algorithms Used for MWL Estimation
Numerous ML and DL algorithms have been explored in the literature; the most commonly used algorithms for MWL estimation are summarized below.
DT
Considering classic ML algorithms, Plechawska-Wojcik et al. (2018) used DT to evaluate feature selection techniques and improve EEG classification accuracy. Similarly, Cheema et al. (2018) proposed an approach to estimate MWL using EEG by employing DT as a supervised ML techniques among ensemble learners. Other approaches introduced by Ho et al. (2019), Asgher et al. (2020b), Le et al. (2018), and Salimi et al. (2019) aimed to predict MWL and employed the Stroop task on pilots and normal participants using NIRS and EEG datasets.
KNN
The authors of Plechawska-Wojcik et al. (2018) employed KNN as a classifier to evaluate and discuss the efficiency of feature selection techniques and the accuracy of classification techniques using EEG data. Moreover, KNN was adopted by Moustafa and Longo (2019) to present a novel ML approach that develops MWL models from EEG data without any theoretical assumptions. Interestingly, KNN is used for automatic MWL estimation with EEG during medical operation and real multitasking activities such as air traffic management (Cheema et al., 2018; Sciaraffa et al., 2019). An additional unique utilization of KNN was presented by Le et al. (2018) aiming to estimate increased MWL in drivers using hemodynamic data from a NIRS device. KNN was also applied to classify MWL levels in pilots using EEG brainwaves during VR-based tasks (Mohanavelu et al., 2022; Ved & Yildirim, 2021). Finally, Zemla et al. (2023) explored differences in brain cortical activity during relaxation or MWL estimation using dense array EEG and subsequently modeled and classified these differences using KNN and generalized linear models.
SVM
SVM has been extensively employed to address various objectives related to MWL estimation, brain activity analysis, and cognitive performance modeling using EEG, fNIRS, EMG, and other neurophysiological data (Asgher et al., 2020a; Li et al., 2024; Lim et al., 2021; Siddiquee et al., 2020). Several studies have evaluated the efficiency, accuracy, and reliability of SVM as a classification or regression tool in the context of MWL estimation using neurological signals. These studies explored diverse applications from aviation to driver’s state detection, healthcare, air traffic, HMI, and neurocognitive research (Hussain et al., 2021; Massaeli et al., 2023; Meteier et al., 2021; Şahin Sadık et al., 2022; Q. Shao et al., 2024).
Naive Bayes (NB)
NB was incorporated into models that aimed to recognize binary MWL, propose TDAE, and combine EEG features for estimating MWL (Tao et al., 2019; Yin et al., 2019; P. Zhang et al., 2019a). Moreover, MWL estimation was investigated using techniques such as the Gray-Wolf Optimizer and ELANFIS integrated with NB as part of their respective frameworks (Das Chakladar et al., 2020; Siddiquee et al., 2020; Teoh Yi Zhe & Keikhosrokiani, 2021). In addition, novel approaches for cognitive performance, such as feature extraction from EEG signals using Gaussian-SMOTE-based feature ensembles involving NB, have been introduced by researchers (Islam et al., 2021; Sharma et al., 2021; G. Wang et al., 2022). Furthermore, Cao et al. (2022) presented a MWL estimation framework by relying on hybrid EEG-fNIRS features and ML, where NB was compared with other classifiers for multilevel MWL estimation. Finally, a unique study investigated NB for real-time MWL estimation among pilots, comparing its performance with KNN and SVM classifiers (Zhu et al., 2023). In another distinctive application, NB was combined with CNN to estimate drivers’ MWL for developing safe HMIs and adaptive driving systems (Caber et al., 2024).
RF
RF was employed by Nittala et al. (2018) to predict the skill level and MWL of aircraft pilots. RF is also used to enhance the accuracy of MWL estimation using FV-MBLL and during multitasking activities, such as air traffic management (Asgher et al., 2019; Sciaraffa et al., 2019), and with law enforcement officers (Wozniak & Zahabi, 2024). For hybridization with DL, Kwak et al. (2020) proposed a 3DCNN with RF for MWL estimation. Further, Meteier et al. (2021) investigated driver conditions with RF among other classifiers, while G. Wang et al. (2022) employed RF within an oversampling framework for imbalanced class workload recognition. In addition, Raufi and Longo (2022) used RF to analyze the impact of band ratios on MWL estimation. Finally, RF was used by Aksu and Çakıt (2023) to estimate MWL based on eye-tracking data, while Park et al. (2024) employed it to estimate MWL in prosthetic devices.
Logistic Regression (LR)
LR is often used in the literature alongside other ML and DL models for estimating MWL from EEG and fNIRS signals. For example, Le et al. (2018) included LR among several classifiers to assess MWL using fNIRS data collected from drivers. In Dhulipalla and Al Hafiz Khan (2022), a DCNN was combined with LR to estimate MWL based on fNIRS signals. Additionally, LR has been evaluated in studies focusing on feature selection and model performance comparison (Sciaraffa et al., 2019).
XGBoost
A novel study conducted by Shafiei et al. (2025) evaluated MWL during surgical tasks by applying XGBoost to integrated EEG and eye-tracking data from 26 participants. The tasks included Matchboard, Ring Walk, Pattern Cut, and Suturing exercises. The analysis incorporated multidimensional features like pupil diameter and temporal lobe connectivity, achieving high predictive accuracy (R2 = .81–.83). Although the data fusion significantly enhanced overall performance, the Pattern Cut task showed weaker results. Similarly, another study proposed a multimodal fusion approach using EEG, ECG, and EDA to assess seafarers’ workload, where XGBoost achieved 85.72% accuracy, outperforming unimodal methods by 9.49%, thus validating the benefits of integrated physiological analysis (Yin et al., 2024).
Linear Discriminant Analysis (LDA)
Several studies have used LDA alongside classifiers such as SVM, KNN, CNN, RF, and multilayer perceptron (MLP) for MWL estimation (Asgher et al., 2020a; Cheema et al., 2018; P. Zhang et al., 2019b). Many of these works proposed novel frameworks that leverage different physiological signals to estimate MWL levels, while others integrated DL techniques like CNN and LSTM to improve estimation accuracy. Additionally, some studies utilized LDA for tasks such as feature selection, data modality fusion, or as part of ensemble learning strategies to enhance the performance of MWL classification models (Huang et al., 2022; Islam et al., 2021; Z. Zhang et al., 2023).
CNN
Extensive research has adopted CNN for MWL estimation using different physiological signals predominantly EEG and occasionally fNIRS psychophysiological sensors, and other biometric data. CNN models have been frequently used with other ML algorithms for MWL estimation. For instance, Appriou et al. (2018b) and Salimi et al. (2019) focused on CNN to estimate MWL from EEG signals, while P. Zhang et al. (2019b) proposed a concatenated structure of deep R3DCNN to learn EEG features across tasks. In addition, CNN models have been used or combined with other models to address MWL estimation from various physiological data sources (Chaturvedi & Ahirwal, 2024; Ho et al., 2019; Massaeli et al., 2023; Saadati et al., 2019, 2020), typically following a common structure as illustrated in Figure 3 below.

A commonly employed architecture in the reviewed studies (Kwak et al., 2020; P. Zhang et al., 2019b).
Meanwhile, some studies have employed various models for MWL estimation, including SVM, KNN, RF, ANN, RNN, LSTM, LR, and ensemble classifiers, often in conjunction with CNN or independently, using EEG, fNIRS, EMG, HR, and eye-tracking data (Abdurrahman et al., 2021; Aksu et al., 2024; Jimenez-Molina et al., 2018; Pandey et al., 2020; Sciaraffa et al., 2019). These models were designed to estimate different MWL levels across a range of cognitive tasks in domains such as aviation, driving, virtual reality VR, robotics, and healthcare.
Autoencoders
Autoencoders have been widely used to enhance feature extraction and improve MWL estimation. For instance, Islam et al. (2021) combined autoencoders with SVM to interpret EEG features extracted by convolutional autoencoders. Similarly, Dhulipalla and Al Hafiz Khan (2022) proposed a DCNN that incorporates autoencoders to estimate MWL from fNIRS signals. In other cases, autoencoders were used for feature learning or as components in hybrid models. For example, Shafiei et al. (2020) integrated autoencoders with SVM, KNN, and RF to estimate MWL during robot-assisted surgery using functional brain network measurements. Likewise, Salimi et al. (2019) introduced an ensemble model that automatically extracts features from EEG channel data, highlighting the potential of autoencoders for effective feature extraction in EEG-based MWL estimation.
MLP
MLPs have been used to predict and classify different levels of MWL based on EEG, fNIRS, heart rate, and eye-tracking data. For example, Cheema et al. (2018) proposed a real-time MWL estimation approach using EEG signals, employing MLP alongside other supervised ML methods. In Asgher et al. (2019), classification accuracy was improved using FV-MBLL in combination with SVM and MLP. Studies such as Elkin and Devabhaktuni (2020) and Ved and Yildirim (2021) conducted comprehensive analyses involving MLPs alongside models like SVM, DT, and RF. Similarly, Pandey et al. (2020) compared several ML algorithms for EEG-based MWL estimation, including MLP. A hybrid approach combining DCNN and MLP was used in Shafiei et al. (2020) to estimate MWL during robot-assisted surgery, taking advantage of both architectures’ strengths. Additionally, Chihara et al. (2020) explored the use of one-class SVM for MWL estimation in driving scenarios.
ANN
ANNs have been extensively utilized in numerous studies for MWL estimation. For example, P. Zhang et al. (2019b) proposed a concatenated structure of R3DCNN alongside traditional ANN models to learn EEG features across different tasks without prior knowledge. Similarly, Asgher et al. (2020a) analyzed and estimated MWL states using a mix of SVM, CNN, and ANN algorithms on EEG and fNIRS data. Furthermore, Salimi et al. (2019) and Xie et al. (2019) specifically mentioned ANNs as part of their models for estimating MWL levels from EEG signals. Similarly, Elkin and Devabhaktuni (2020) presented a comprehensive analysis of DL techniques, including ANNs, used for estimating MWL. Finally, ANNs have been integrated into hybrid models or combined with other DL techniques. For instance, Saadati et al. (2020) investigated the use of CNN and ANN to estimate MWL, bypassing challenges in feature selection. Shafiei et al. (2020) proposed an innovative framework for estimating MWL in robot-assisted surgery using a combination of DCNN and ANN.
LSTM
LSTM is frequently applied to process time-series data such as EEG signals. For instance, Yin et al. (2019) proposed a novel TDAE using LSTMs to capture the dynamical properties of EEG features and individual differences. Asgher et al. (2020b) implemented LSTMs on four-level MWL-fNIRS data, achieving optimal classification accuracies for MWL. Furthermore, Afolabi et al. (2025), Ktistakis et al. (2022), and Kwak et al. (2020) introduced innovative approaches for MWL estimation involving 3DCNN by employing multilevel feature fusion algorithms and LSTMs. Islam et al. (2021) used mutual information to explain EEG features extracted via convolutional autoencoders and incorporated LSTMs into their hybrid autoencoder-SVM classifier. Furthermore, LSTMs have been compared with other ML models. For example, Ktistakis et al. (2022) introduced a dataset for MWL estimation based on eye tracking data and assessed various classifiers, including LSTMs, to evaluate their effectiveness in MWL prediction.
ELM
ELMs have been used to develop individual-specific ensemble classifiers for binary MWL estimation (Tao et al., 2019). In the same study, a heterogeneous ensemble ELM was compared with traditional models such as KNN and ANN. Additionally, Teoh Yi Zhe and Keikhosrokiani (2021) enhanced the ELANFIS model by integrating particle swarm optimization within a microgenetic algorithm to estimate MWL in knowledge workers. The effectiveness of ELM in handling physiological data was further demonstrated by G. Wang et al. (2022), where a novel EEG feature oversampling method, Gaussian-SMOTE-based Feature Ensemble (GSMOTE-FE), was proposed to address class imbalance in MWL estimation. Finally, Yu et al. (2025) applied ELM to estimate MWL in air traffic controllers by combining EEG and eye-tracking data, achieving better performance than RF, SVR, and FFNN models.
In summary, the literature reveals a clear methodological progression from conventional classifiers such as DT, KNN, SVM, and RF toward deep and hybrid architectures capable of learning more complex spatial, temporal, and multimodal patterns. This shift reflects both the growing availability of richer physiological data and the increasing need for models that perform well beyond tightly controlled experimental settings.
MWL Estimation Outcomes and Limitations
The outcomes of the reviewed studies can be interpreted not only in terms of performance, but also in terms of how the field has matured. Across the literature, progress is visible in three broad ways: the expansion of MWL estimation into new domains, the improvement of model performance through better feature and model design, and the gradual move toward more realistic, multimodal, and real-time applications.
Exploring New Areas of MWL Estimation
In terms of exploring new areas, the authors of Hajra et al. (2020), Nittala et al. (2018), and Sciaraffa et al. (2019) presented approaches to predict skill and MWL of pilots using various physiological signals and ML algorithms. Overall, they reported good results for three MWL levels (low, medium, and high) with an accuracy of 84% for single data sources (flight, ECG, and EEG), and the results improved with the introduction of multimodal approaches and ensemble techniques. However, limited data were used in this specific context (aviation) in addition to using a single set of SVM regressors. Due to a similar lack of sufficient data, approaches using combined modalities (e.g., PPG, thermography, and fNIRS) reported an accuracy of more than 78% (Asgher et al., 2021; Cho et al., 2019). Approaches focusing on real-time assessment (Chihara et al., 2020; Lim et al., 2021; Massaeli et al., 2023; Sakib et al., 2021) have reported a moderate to high accuracy of >66%; however, further investigation into accuracy is needed in complex environments, as accuracy varied across participants performing specific tasks.
Alternative techniques and modalities were proposed by Luo et al. (2025), Mirzaeian and Ghaderyan (2023), Şahin Sadık et al. (2022), Teoh Yi Zhe and Keikhosrokiani (2021), Ved and Yildirim (2021), and Yiu et al. (2022) to develop and investigate novel approaches (e.g., ELANFIS optimization, new textural features for EDA) using various data sources (e.g., EEG, eye-tracking, odors). Although with limited data, a promising result with an accuracy of up to 91% for MWL estimation was reported; however, this finding requires validation on larger datasets and in real-world scenarios.
Investigating Feature Engineering Techniques for MWL Estimation
The authors employed advanced DL algorithms, such as CNN and TSNN, for feature engineering on multiple modalities, including EEG, fNIRS, ECG, and EDA (Cao et al., 2022; Ding et al., 2020; Jimenez-Molina et al., 2018; Saadati et al., 2020; P. Zhang et al., 2019a). Selecting the appropriate features achieved accuracies of 91.24%, 91.9%, 94%, and 96.4%, respectively. The authors reported a lack of data and limitations on specific tasks (i.e., n-back and web browsing). Similarly, Asgher et al. (2019), Becerra-Sánchez et al. (2020), S. Shao et al. (2024), and G. Wang et al. (2022) reported high accuracies; their objectives were to apply novel feature selection and optimization techniques to improve MWL estimation accuracy. Asgher et al. (2019) used the FV-MBLL algorithm, which improved classification accuracy to 94%, and Becerra-Sánchez et al. (2020) used GALoRIS for feature selection and achieved 90% on the precision metric. Moreover, G. Wang et al. (2022) adopted the GSMOTE-FE oversampling framework and showed performance comparable to state-of-the-art approaches. Despite the reported performance, overall, the authors focused on specific tasks or data types with limited data sizes in some studies.
Enhancing and Proposing Novel Approaches to MWL Estimation
Studies have employed advanced models such as CNN, LSTM, and RF on multimodal physiological data, including EEG, ECG, galvanic skin response, and eye-tracking, achieving baseline accuracies around 75%, which improved to 97.8% with optimized implementations (Aksu & Çakıt, 2023; Asgher et al., 2019; Das Chakladar et al., 2020; He et al., 2022; Z. Zhang et al., 2023). Despite these advancements, limitations persist, including small datasets, context-specific models, and real-world data imbalance. In parallel, several novel approaches have explored multimodal and DL techniques for MWL estimation. Notable examples include R3DCNN (88.9%) and TSNN (91.9%) using EEG data (P. Zhang et al., 2019a, 2019b), and DCNNs combined with dynamic brain networks and fNIRS for robot-assisted surgery (>91%; Dhulipalla & Al Hafiz Khan, 2022; Shafiei et al., 2020). A multimodal ECG-EDA approach reached 96.4% accuracy (Ding et al., 2020). While these models show promise, they often incur high computational costs. To mitigate this, traditional ML methods have also been explored, for example, empirical mode decomposition on ECG and PPG signals using RF yielded 88.64% accuracy and 94.9% recall with lower complexity (Dayal et al., 2025).
Exploring Novel and Existing MWL Datasets
Novel datasets such as GALoRIS (Becerra-Sánchez et al., 2020), derived from EEG signals, reduced the original data size by over 50% and achieved over 90% precision in classifying high and low MWL. Similarly, the COLET dataset (Ktistakis et al., 2022), based on eye-tracking data, reached up to 88% accuracy in binary and multiclass MWL classification using the GNB classifier. In addition to these, multiple studies have evaluated existing datasets, achieving state-of-the-art performance with accuracy exceeding 80% across various modalities for example, 81.30% to >95.40% (NIRS), 78.33% (multimodal), 80% to 87% (CNN EEG), 94.00% (EEG MLA), 93.70% (EEG LSA), and 97.8% (combined CNN + LSTM and RF multifactor using NIRS, PPG, EEG, ECG, EDA, traffic flow, and environment data). However, researchers emphasize the need for validation on broader datasets to ensure generalizability (Cho et al., 2019; Huang et al., 2022; Hussain et al., 2021; Le et al., 2018; Sharma et al., 2021; Taori et al., 2022; Wei et al., 2023).
Investigating Automated MWL Estimation
Real-time MWL monitoring and adaptive systems have been used for operational tasks. Lim et al. (2021) demonstrated the feasibility of real-time workload assessment inference and HMI adaptation with variable accuracy (RMSE 0.2–0.6). Similarly, DellrAgnola et al. (2022) proposed a subject-specific ML model for search and rescue operators, achieving 87.3% to 91.2% accuracy in MWL estimation for real-time operation. Multimodalities were introduced in real time to explore multimodal and cross-subject approaches for simultaneous workload and stress classifications, achieving 77.5% and 84.1%, respectively, using cross-subject transfer learning with EEG data (Bagheri & Power, 2022). Another study demonstrated effective pilot workload identification using multimodal features and an RF classifier, achieving 90.5% average accuracy (Zhu et al., 2023).
Investigating Validation Techniques for MWL Assessment
Chihara et al. (2020) and Kingphai and Moshfeghi (2023) introduced OCSVM for anomaly detection to drive MWL evaluation and successfully identified 95% of high workload states (3-back task) during automobile driving based on eye and head movement features. However, its generalizability was limited due to the specific driving environment and the small sample of participants. Moreover, a cross-validation method was proposed for time-series EEG data used in LSTM models for workload classification and achieved high accuracy in two tasks 87.44% (Task 1) and 88.79% (Task 2) with optimized training and test window sizes. The authors highlighted the trade-off in selecting the time window size, balancing noise reduction with real-time applicability (Chihara et al., 2020; Kingphai & Moshfeghi, 2023).
Overall, the literature suggests that MWL estimation research has evolved from proof-of-concept classification studies toward more ambitious systems designed for multimodal fusion, online adaptation, and practical deployment. At the same time, the persistence of small datasets, task-specific designs, and limited external validation shows that methodological progress has not yet fully translated into broad real-world generalizability.
Future Research Directions for MWL Estimation
Recent advances in ML and DL have substantially improved the accuracy and effectiveness of MWL estimation. The reviewed studies suggest several future research directions, which can be categorized as follows.
Future Works on Datasets
Several studies have highlighted the need to improve datasets for MWL estimation. Cho et al. (2019) and Nittala et al. (2018) recommended incorporating EDA in real-world scenarios, along with a broader range of physiological features such as EEG, cardiac features, and respiration signals (Hussain et al., 2021; Meteier et al., 2021; Parent et al., 2019). Researchers also emphasized the importance of using large and highly diverse samples to develop more generalizable and scalable models (Chen et al., 2023; He et al., 2022; Raufi & Longo, 2022). Other suggestions include optimizing sensor placement (Siddiquee et al., 2020), exploring additional modalities (Islam et al., 2021; Kesedzic et al., 2021), and refining signal processing techniques (Cao et al., 2022; Sharma et al., 2021). Additionally, the dataset integration process is recommended to include EEG, HRV, and brainwave signals to enhance data quality (Chen et al., 2023; Yiu et al., 2022).
Future Work on Feature Engineering
Studies have highlighted the need for a comprehensive evaluation of features most relevant for predicting overall MWL perception (Moustafa & Longo, 2019). This includes creating simpler MWL models by identifying and incorporating the most critical features, thereby developing a highly generalizable model applicable across various domains and experimental contexts. Similar studies need to be conducted, particularly in fNIRS data analysis (Kesedzic et al., 2021), with mean activation underscored as a crucial feature, driving the need for meticulous feature selection to enhance the MWL estimation model performance. Moreover, a study highlighted the use of EEG and HRV features to determine cardiac-cognitive correlations for perceived MWL (Mohanavelu et al., 2022). This approach suggests leveraging physiological signals to capture workload variations and correlations. A similar suggestion was made by Mirzaeian and Ghaderyan (2023), focusing on optimizing computational complexity while ensuring high performance in feature engineering; for instance, using fewer textural features can offer a better tradeoff between complexity and performance. Furthermore, the incorporation of time-series strategies for physiological signal analysis was proposed to better capture dynamic behaviors, especially in driving contexts (Wei et al., 2023). This involves exploring diverse metrics beyond mean and variance to capture fluctuating physiological responses during high workload. These considerations can extend to analyzing external factors (e.g., in-vehicle passengers) and their dynamic influence on MWL estimation, suggesting the need for a comprehensive approach encompassing various factors.
Future Work on MWL Estimation
The future directions for MWL estimation continue to be shaped by a wide range of applications and increasing demands for accuracy and adaptability. One notable area is web browsing tasks, where Jimenez-Molina et al. (2018) recommended enhancing MWL estimation through psychophysiological signals and refining experimental designs by incorporating tasks of varying difficulty. In operational environments, Hajra et al. (2020) emphasized broader adoption of objective workload assessments, particularly in safety-critical contexts. Human-machine task allocation also presents optimization opportunities, as highlighted by Pandey et al. (2020), who proposed classification models to enhance task distribution. Sakib et al. (2021) further emphasized addressing the gap between VR and real-world settings to improve calibration.
In medical and therapeutic domains, Asgher et al. (2021) advocated shorter task windows to facilitate real-time classification for individuals with motor disabilities, while Şahin Sadık et al. (2022) highlighted the importance of understanding MWL differences in individuals with cognitive impairments. Training and simulation environments have been identified as promising for MWL estimation. Sakib et al. (2021) suggested leveraging VR-based models to simulate stress and cognitive load where real-world exposure is limited or hazardous. Z. Zhang et al. (2023) highlighted the value of real-time MWL detection in complex human-machine systems such as subway or aircraft operation for enhancing system performance, operator safety, and overall well-being.
In aviation and navigation, Yiu et al. (2022) proposed using unmanned aerial systems as decision support tools to improve pilot situational awareness. Similarly, Sobrie et al. (2024) introduced a real-time ML framework based on LightGBM to predict MWL in digital railway control rooms using granular 15-min interval data from Infrabel, demonstrating superior performance compared to existing models. Lastly, Das Chakladar et al. (2020) emphasized expanding MWL classification models to cover more complex tasks and identify associated brain region patterns through connectivity analysis. Complementarily, Ding et al. (2020) recommended defining workload thresholds for various task types to facilitate timely interventions.
Future Work on ML and DL Techniques
From an algorithmic perspective, ML and DL algorithms have been suggested for MWL estimation in multiple approaches and applications. For ML models, studies recommended adopting ensemble learning and structures for EEG classification tasks based on different brain regions and investigating generic MW classifier ensembles for individual-specific online use (Salimi et al., 2019; Tao et al., 2019). Another suggestion is adopting approaches that dynamically select efficient ML techniques based on MWL data for early estimation (Elkin & Devabhaktuni, 2020). Similarly, G. Wang et al. (2022) and Yiu et al. (2022) highlighted the need for dynamic and adaptive regression models to mitigate issues and improve adaptability.
To improve DL model performance, studies mainly focused on enhancing architectures with minor refinements. Appriou et al. (2018b), Asgher et al. (2020b), Ho et al. (2019), Saadati et al. (2020), and Yin et al. (2019) emphasized exploring CNN, LSTM, and attention-based models using EEG and fNIRS data, considering computational efficiency and knowledge transferability. New DL models employing RNNs, autoencoders, and augmentation techniques have been proposed to boost performance (Islam et al., 2021; Ktistakis et al., 2022; Meteier et al., 2022). Simplifying DL models for broader MWL applications and real-world conditions is a promising direction (P. Zhang et al., 2019b). Continued investigation into novel architectures such as convolutional/LSTM hybrids and Lambda Nets is ongoing. Addressing computational costs and optimizing training parameters is essential to balance accuracy and practical relevance, particularly in automated driving and decision-making applications (Ktistakis et al., 2022; Tao et al., 2019).
Overall, the reviewed literature shows an increasingly coherent methodological pattern in MWL estimation research. Although earlier studies differed considerably in modality choice, feature design, and model selection, more recent work shows growing convergence around a general pipeline that includes signal acquisition, preprocessing, representation or feature learning, model development, and evaluation. As illustrated in Figure 4, this convergence reflects the maturation of the field and suggests that future advances are likely to depend less on isolated performance gains and more on reproducibility, generalizability, and deployment in realistic settings.

Steps for MWL estimation: EEG use, preprocessing, ML/DL methods, and evaluation (Zhou et al., 2022).
Conclusion
The overall trajectory of the literature indicates a clear movement from conventional, feature-driven ML approaches toward deep, hybrid, and multimodal frameworks aimed at improving robustness, interpretability, and real-world usability. This systematic review examined 110 peer-reviewed articles published between 2018 and April 2025, highlighting the widespread adoption of ML and DL techniques for MWL estimation across a range of domains and contexts. The reviewed literature explores novel application areas, including office environments, operations, digital railway systems, and air traffic control, where MWL monitoring is critical for performance and safety.
The analyzed studies utilize various techniques such as PCA, feature selection, feature extraction, and feature scaling to identify informative patterns within physiological data. In addition, widely used datasets include EEG, EDA, ECG, EMG, fNIRS, NASA-TLX, STEW, and SWELL-KW, as well as newly introduced datasets like COLET and GALoRIS. Researchers have implemented a range of traditional ML algorithms (e.g., DT, KNN, SVM, NB, RF, LR, and LDA) and advanced DL architectures (e.g., CNN, Autoencoders, MLP, ANN, LSTM, and ELM) to enhance MWL estimation accuracy and performance. Furthermore, new datasets have been introduced, such as COLET and GALoRIS, as sub-set features for EEG data in addition to other explored datasets, such as EDA, ECG, EMG, fNIRS, NASA-TLX, STEW, and SWELL-KW.
For future studies, the reviewed articles have pointed toward essential areas for improving MWL estimation, such as focusing on better data collection, incorporating EDA in real-life situations, and optimizing the placement of sensors used for effective data gathering. In addition, there is a need to thoroughly understand the features that are most important for estimating MWL and to use strategies that can track dynamic changes in physiological signals during tasks such as driving. Additionally, current research focuses on refining MWL assessment during web browsing, applying MWL estimation in diverse work environments, and optimizing task allocation between humans and machines. In terms of technique, researchers suggest using ensemble learning by integrating different DL techniques, such as CNN and LSTM, and investigating newer ensemble architectures that utilize RNN and autoencoder to improve MWL estimation accuracy and efficiency.
This review underlines the increasing interest in using ML and DL to improve MWL estimation, an area with substantial potential for real-world impact. Future studies may benefit from focusing on underexplored aspects, such as comparative analysis of methodologies, robust validation strategies, and applications in domain-specific settings like healthcare, transportation, and human-machine collaboration.
Footnotes
Ethical Considerations
This study is a systematic literature review based exclusively on previously published literature. It did not involve the recruitment of human participants, the collection of identifiable personal data, or any intervention involving human subjects. Therefore, ethical approval was not required.
Consent to Participate
As this study did not involve human participants or the collection of personal data, consent to participate was not required.
Funding
The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by a Universiti Sains Malaysia, Bridging Grant with Project No: R501-LR-RND003-0000002090-0000.
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Data Availability Statement
Data sharing is not applicable to this article, as no datasets were generated during the current study.
