Abstract
Big data classification has become popular for classifying important data such as healthcare, stock market, and other fields’ data for effective handling of many classical databases. In recent years, deep learning and machine learning techniques have become more reliable for such classification, which has resulted in effective results with better accuracy. In this survey, articles from around 10 years are taken for analyzing various techniques recently for big data classification, and this study helps to improve the future evolution of fine, adequate, and highly acceptable models for classification-based approaches. This survey contains the advantages and disadvantages of several methods, which aid in overcoming the challenges of building new ensemble models for classification. Moreover, this research provides complete knowledge about the existing methods and emphasizes the ideas for creating high-potential models as well as paving the way for inventing a real-world application to classify big data. The performance metrics to measure the classification techniques are accuracy, sensitivity, and specificity, and also provide improvement for future classification processes.
Introduction
In recent years, for the development of digital technology, enormous data has been required from various fields such as healthcare, Internet of Things platforms, social media platforms, and so on. These widespread usages of digital technology have demanded the growth of big data. Big data refers to the integration of vast and complex datasets that existing database management tools struggle to handle effectively. A classical database is a traditional relational database that is designed to store and manage structured data with a fixed scheme. Big data involves large, expanding datasets that are complex and sourced from multiple autonomous origins. The key difference lies in architecture and capability: classical databases rely on vertical scaling, fixed schemas, and batch processing, making them rigid and slow for modern, diverse data. In contrast, big data leverages distributed, horizontally scaled architectures to handle massive amounts of unstructured, semistructured, and structured data, enabling near real-time processing and dynamic analysis. Essentially, classical systems are for organized, static data, while big data platforms are for fluid, vast datasets generated at high speed from sources such as social media and sensors. The primary difference between big data and classical data lies in the methods, procedures, objectives, and strategies utilized to extract valuable insights from the information. Previous technologies were unable to manage the storage and processing of such enormous data, leading to the development of the big data concept. Big data describes large datasets that cannot be effectively managed or handled by classical database systems. (Banchhor & Srinivasu, 2021; Benabderrahmane et al., 2017) and quite challenging due to the expeditious gathering of data, storage, and other related techniques (Raghav et al., 2017; Thanekar et al., 2016). Big data has five key characteristics: volume, velocity, variety, value, and veracity. The big data generally utilizes two traditional approaches, namely classification (Manogaran et al., 2018) and clustering (Yildirim, 2019), which are relevant to data mining techniques. Here, the classification approach belongs to the supervised learning approach, whereas clustering is an unsupervised approach that requires no pretraining strategy for cluster performance. The classification is defined as grouping the data according to the respective similarities and differences from the gathered data, and providing data tagging. Clustering is a technique that groups data points based on their similarities, forming clusters without any prior labels or predefined knowledge of the groupings. This makes it possible for better identification and segregation of data, and also makes it easy to find and locate the related data. Big data classification is strongly admired in real-time applications such as healthcare management, epidemic outbreak prediction, and so on (Khanna et al., 2021).
Specifically, big data analytics, particularly in sectors such as healthcare and the stock market, were quite complex due to the large volumes of data and the high computational costs required. The challenges faced in the field of big data can be solved efficiently using a classification approach. The classification approach proceeded with analysis strategy, processing techniques, and storage (Suthaharan, 2012). There were multiple possible techniques used for classification, which initially worked by extracting huge amounts of information from a large set of data (Koturwar et al., 2015). However, two major tools, such as Apache Spark and MapReduce techniques, are used for the classification. Apache Spark refers to a kind of tool used for security analysis in big data (Brahmane & Krishna, 2021; Lighari & Hussain, 2017). It is also known for speedy cluster computing that is efficiently managed for high accuracy. On the other hand, MapReduce is also a technique, embedded with a mapper and reducer to break down the large or complex data into a simple form for classification (Bhukya & Gyani, 2015). This technique utilizes a Hadoop file system that reads and writes to the disk, and also has shortcomings due to inefficiency and redundancy (Elsebakhi et al., 2015). These two efficient methods are the key tools to support the traditional particle swarm optimization technique and genetic algorithm (Ditzler et al., 2017).
Machine learning (ML) techniques have been used in big data classification (Wang & Alexander, 2016). There were several classification techniques based on ML to classify the big data, namely, K-nearest neighbors (KNNs; Saadatfar et al., 2020), Naïve Bayes (Shah et al., 2020) algorithm, decision tree (DT; Shaik et al., 2023), Hadoop MapReduce (Nagesh & Prabhu, 2017), and random forest (RF; Manikandan, 2018). It also required supervised and unsupervised techniques such as a DT and a support vector machine (SVM; Huang et al., 2014). In addition, several optimization-based approaches were developed to solve big data classification. The DT is a hierarchical model adopted with nodes and leaves, which highly divides the input space into class regions. However, KNN is another technique that was adopted for finding the nearest samples and selecting them with the respective traditional methods. Moreover, the hybrid model of SVM–KNN is utilized to compute the query distances, which are comparatively slower than other traditional approaches. In big data, the computation process was much more complex, so fuzzy classification was developed. The high usage of the fuzzy technique results from weak identification and increases computational expenses, which is the major drawback of this method (Manikandan, 2018). The development of neural networks has led to the rapid growth of deep learning (DL) techniques such as convolutional neural networks (CNNs), Deep CNN (DCNN), long short-term memory (LSTM), and so on. DL offers the capability to tackle data analysis and learning challenges posed by a large amount of input data. Notably, it excels at automatically extracting complex data representations from extensive amounts of unsupervised data. As data volumes increase and advancements in graphics processors and processing power continue, DL is becoming more crucial for delivering big data predictive analytics solutions, necessitating further investigation.
The main objective of this study is not to provide a comprehensive overview of all the relevant work in ML and DL, but rather to highlight current research endeavors, the difficulties associated with big data, and key issues about learning from vast amounts of data, as well as future trends. Specifically, this review analyzes 55 research articles relevant to big data classification utilizing advanced strategies and algorithms, exploring their advantages and disadvantages to formulate advanced proposals for effective big data classification in the future. This systematic review contributes to exploring the potentials and difficulties of big data classification to formulate an efficient contribution. Furthermore, the review analyzes the advanced methods of big data classification, limitations, datasets used, and performance metrics contemplated for big data classification. This review is initiated by offering the background of multiple big data classification models and provides an overview of the technical details of several mechanisms involved in the models, which paves the way for improving an effective model for future research. Moreover, the review underscores that the DL models provide exceptional performance compared to other ML classifiers.
The research article is categorized in the following structure: The classification of multiple methods of big data is categorized in Section 2, analysis and discussion of various big data classification techniques are presented in Section 3, Section 4 exposes research gaps and future works, and lastly, the conclusion is in Section 5.
Research Questions
In this review, the research questions are meticulously developed to overcome the inherent challenges associated with big data classification techniques. Furthermore, the selection of relevant articles and the analysis are made in alignment with the research questions. Moreover, the review offers a valuable source to foster informed decision-making and innovation in the context of big data classification.
Which specific metaheuristic techniques have been applied for feature selection? In the context of big data classification, what are the datasets commonly utilized for assessing the effectiveness of the classification models? What are the major challenges prevalent in the recent methods of big data classification? Which classifiers are highly effective for big data classification? What are the future works that can be applied to enhance the big data classification?
Research Objectives
The majority of state-of-the-art reviews focus on describing and explaining ML or DL techniques without evaluating, and presenting the results as indicators taken from other recent works, or presenting the results without a methodical and repeatable approach to the collection of articles and results, and statistical analysis. The main objectives of this proposed review for big data classification are listed as follows:
To provide a systematic analysis and a comprehensive description of recent methods in the field of big data classification. To analyze the recent and cutting-edge techniques applied for feature selection and preprocessing in big data classification. To identify and investigate the datasets commonly utilized for assessing the effectiveness of the classification models. To explore the challenges prevalent in the recent methods of big data classification. To investigate the results of highly effective classifiers and advanced mechanisms prevalent for big data classification. To explore the future works and advancements for addressing the challenges and improving the big data classification.
Overview of Big Data Classification
Categorizing the big data problems based on the type makes it easier to see the characteristics of each kind of data. These characteristics provide insight into the data's acquisition process, processing the data into the appropriate format, and the frequency with which new data becomes available. Figure 1 depicts the architecture of big data classification.

Architecture of big data classification.
This section will detail the methodology, emphasizing the systematic approach used in article selection. It will also include an extensive literature review that provides an overview of big data classification. Figure 2 outlines the stages followed in the PRISMA-based review.

Flow diagram of a systematic approach.
The proposed work conducts a systematic literature review adhering to PRISMA guidelines to provide a comprehensive review. Initially, a detailed search filters articles from data sources. To include the most recent works in the field of big data classification, only articles from 2020 to 2024 were considered. The literature search was conducted across several electronic databases, including IEEE, MDPI, Springer, Elsevier, Research Gate, Wiley, and others. The research strategy was refined iteratively to ensure broad coverage of relevant studies while excluding irrelevant or duplicate records. The search includes keywords such as “Big data,” “Classification of big data,” “big data medical,” and “big data analytics” for finding the relevant research articles. After selecting the articles, each one was evaluated for relevance to the research topic; papers with irrelevant titles, abstracts, contributions, papers of poor quality, and insufficient analysis were excluded based on eligibility criteria. Finally, 55 papers were reviewed for an in-depth analysis. The PRISMA flow diagram visually represents the number of articles selected, identified, and excluded, along with the reasons for exclusion. This systematic approach enhances the validity and reliability of the proposed review.
Figure 3 illustrates the preprocessing techniques for big data classification, which is also an initial step for all DL and ML techniques, where the unwanted background from the image datasets is removed to enhance the quality of an effective classification process. Several preprocessing techniques are handled by various approaches, such as feed-forward neural network (Ashraf et al., 2020), K-means algorithm (Manikandan, 2018), MapReduce (Banchhor & Srinivasu, 2021; Game et al., 2022; Lakshmanaprabu et al., 2019; Narayana et al., 2022; Sujitha & Seenivasagam, 2021), linguistic hedges neuro-fuzzy classifier with selected features (LHNFCSF; Azar & Hassanien, 2015), maximum Likelihood (Boulila et al., 2021), bottom-hat filter (Naronglerdrit & Mporas, 2021), manifold analysis and nearest neighbor propagation (Jiang & Li, 2019), and locality sensitive hashing–synthetic minority oversampling technique (Hassib et al., 2020). Most reviewed papers utilized MapReduce techniques to reduce the backlogs from the datasets. Furthermore, the above-mentioned techniques are highly reliable and more viable, which reduces the time duration for computation and also earns a high classification performance of the model.

Preprocessing techniques for big data classification.
Feature extraction means extracting the key features from the given data and works by reducing the high-dimensional parameters. Due to this feature extraction, the model can be highly trained for better classification and improve the computational speed effectively. In other words, raw data is transformed into numerical formats that are highly compatible with both ML and DL methods. This approach is an important and mostly used method for effective classification, some of the methods are collected from recent articles such as the GoogleNet method (Ashraf et al., 2020), Spark framework (Tchito Tchapga et al., 2021), gray level difference matrix (GLDM; Almutairi et al., 2024), SVM (Ashraf et al., 2020; Zhu & Chen, 2020), Spark-Standalone mode (Xing & Li, 2019), principal component analysis (PCA; Wu, 2022), and DT (Rawal & Agarwal, 2019) are shown in Figure 4. These techniques are deployed for accurate classification and provide superior performance.

Feature extraction and selection techniques for big data classification.
The GoogleNet method, a robust DL architecture specifically a CNN, is employed for feature extraction in big data classification. It excels in efficiently extracting complex and highly discriminative features from large datasets.
Spark Framework
The Spark framework prepares data for iteration, allows for repeated querying, and loads it into memory. The main program manages multiple workers and collects their results. Workers read data partitions from a distributed file system, perform computations, and save the results to disk.
GLDM
A GLDM is employed in feature extraction for big data classification, especially for image data. It captures textural information within an image by analyzing the frequency of different gray level combinations between neighboring pixels, offering valuable features for distinguishing between various data classes.
SVM
An SVM aids in feature extraction for big data classification by pinpointing the most relevant features in high-dimensional datasets. This reduces the feature space while ensuring high classification accuracy, particularly in handling complex, nonlinear relationships between features and class labels.
Spark-Standalone Mode
The Spark-Standalone mode allows for cluster launching either manually or via a launch script. In its application master/worker management mode, the master primarily handles resource management and scheduling. However, in Spark-Standalone, this master/worker mode is vulnerable to single-point failure issues.
PCA
The PCA produces new variables that are linear combinations of the original variables. It can extract the most informative features from large datasets while retaining the most relevant information from the original dataset.
Feature Selection Techniques
This crucial technique for ML and DL approaches involves selecting extracted data by reducing irrelevant parameters, thereby enhancing the model's learning capability while minimizing computational expenses. The analysis of new articles utilized a library for SVMs (Lin et al., 2016), Whale optimization algorithm (WOA; Hassib et al., 2020), Pearson correlation-based Black hole entropy (BHE) fuzzy clustering (Narayana et al., 2022), and improved Dragonfly algorithm (IDA; Lakshmanaprabu et al., 2019), which are shown in Figure 5. The feature selection methods reduce the complexity of the model for effective classification.

Optimization techniques for big data classification.
WOA is employed for feature selection by treating each feature subset as a whale's position. Each subset contains a random number of features, up to the total number of features in the original dataset. The whale representing the subset with the fewest features and highest classification accuracy is deemed the best solution.
Pearson Correlation-Based BHE Fuzzy Clustering
The Pearson correlation-based BHE fuzzy clustering is used as the feature selection technique in the existing method. BHE fuzzy clustering generates samples using the fuzzy membership function and performs feature selection by integrating three key aspects: the Bayesian inference model, fuzzy clustering, and BHE-based information. This method is particularly effective for large datasets due to its integration of fuzzy clustering with BHE. Additionally, the Pearson correlation distance is used to characterize clustering behavior.
IDA
In the conventional method, the IDA is used for the feature selection process. The IDA depends on both static and dynamic swarming capabilities, such as partitioning, formation, cohesion, attraction to food sources, and evasion of predators. These swarming behaviors closely mirror the two main phases of optimization using metaheuristics.
Review of Big Data Classification Models
Big data classification techniques are one of the efficient strategies for analyzing, processing, storing, and classifying the data accordingly for easy locating and aggregation without complexity. The reviewed papers have huge techniques that effectively solve the problem. Moreover, the ML and the DL are the prominent techniques used in many classification tasks as well as other complex problems to make simple and unique, and these techniques are considered for high accuracy performance. In this paper, the techniques, methods for processing, and other strategies for classifying big data are emphasized to forge multiple models in the future. The taxonomy of different classification techniques is shown in Figure 6. Table 1 presents the abbreviations of the model included in Figure 6.

Technique used in big data classification.
Abbreviations.
In big data classification, ML techniques are the most prominently used methods inspected in recent papers, which are highly required for classification, and the survey is provided in the context below.
DT Method
Game et al. (2022) introduced the DT model for big data classification, which was ideally done using a divergence-based grey wolf optimization (DGWO) algorithm. The DGWO algorithm selects the major attributes for classification, minimizes the class mixture, and makes the determination of class easier, hence providing better performance. However, this model does not provide high accuracy and also requires maximum computational expenses for the complete procedure. Chern et al. (2021) developed a hybrid method to solve the big data problems in credit assessment and achieve better results compared to the artificial neural network (ANN) model, also lowering the computational complexity of classifying the big data. However, the model requires more time and drives complications during the classification process, which includes the model's intended purpose to minimize the complexity of the attributes. Moreover, the attributes are not enough for the data mining model, in addition suffer from low sensitivity. Rawal and Agarwal (2019) deployed C4.5 DT-based classification for big data, which involves calculating the weight parameters between several nodes. This mechanism was also represented as a classifier of statistics that brought efficient classification with high accuracy. Moreover, the minority range of this model was specifically designed for measuring some kind of impurities.
Gradient Boosting Method
Kadkhodaei et al. (2021) utilized an ensemble method, a distributed heterogeneous boosting-inspired ensemble classifier (DHBoost), where the boosting approach enhanced the automatic selection for proper learning procedures and provided increased classification accuracy. Also, this results in maximum computational expenses and is also sensitive to unwanted noises. Ait Hammou et al. (2019) also initialized with the boosting model, extreme gradient boosting (XGBoost), to classify the early and late stages of cancers. This method has scalable benefits to handle large datasets and provides efficient computation. Despite the method being slow for the training process and vulnerable to overfitting problems. Hussein and Abdulkader (2022) developed a Light Gradient Boosting Machine (Light GBM; Abbasniya et al., 2022; Asif et al., 2023) classification model deployed with VGG 19 DCNN. The Light GBM model has several benefits of scalability and reliability, and provides significant results compared to DT. Despite that, the model has overfitting issues and provides computational complexity. In addition, the selection of leaves during tree generation led to the majority loss of important data.
SVM Method
Tchito Tchapga et al. (2021) utilized the ML model, SVM, for finding the optimal hyperplane parameter, which aids in error minimization during classification, and also utilized the CNN approach for classification that highly extracts the high-level representations belonging to the big data and provides approximate results. Despite that, the SVM model only provides better performance using medium or small datasets, but not large datasets. Sujitha and Seenivasagam (2021) also developed an SVM model with different algorithms, binary classification, and multiclass classification with threshold technique (T-BMSVM) for classification. The combination of SVM and T-BMSVM outperforms, with promising performance, high scalability, and provides maximum convergence for classification. Despite this, the model was not implemented using a large number of datasets, which consumes more time and cost.
Complement Naïve Bayes Method
Banchhor and Srinivasu (2021) utilized a correlative Naïve Bayes (CNB) classifier with cuckoo search and grey wolf optimization technique for big data classification, which has achieved comparatively high results with other CNB-based classifiers. The model provides quite high accuracy as well as reliability for better outcomes. However, this approach was highly expensive during computation. Banchhor and Srinivasu (2020) deployed the Cuckoo–Grey wolf-based Correlative Naive Bayes classifier with MapReduce model (CGCNB-MRM), which highly selects an important parameter with the help of an optimizer and affords classification with high speed and fewer computation costs. Despite the performance rate being improved, the output was always affected based on the quality of the datasets. This limitation can provide low accuracy and decreased efficiency of the model.
Other ML Techniques
Xing and Bei (2019) introduced a class-based weighted KNN model to develop a classification model; this method significantly reduced the computational time and was also proficiently performed using a large number of datasets. The improved KNN model only focused on a single-class classification, which decreases the sample sparsity of the category description. Manikandan (2018) utilized an RF (Lakshmanaprabu et al., 2019) classifier that generates test error in the forest building process, as well as an internal unbiased estimation, which was highly capable of dealing with the classification process. Using this RF model, high computation costs and lower convergence speed for classification were required. Liang et al. (2018) utilized the extreme inception and LSTM (Xception-LSTM) model, which tends to extract spatial–temporal features that are intended for segmentation and classification. However, the Xception-LSTM model required high computational expenses, and the problem can be solved using computer hardware. Azar and Hassanien (2015) developed an LHNFCSF for preprocessing, which deliberately reduced the dimensionality, discarded redundant noise, and achieved high-speed classification. Somehow, the results of the LHNFCSF model are not completely interpretable and are not implemented in complex problems. Hernández et al. (2020) developed linear-morphological neural network (LMNN) and morphological-linear neural network (MLNN). Especially, the morphological neurons in the model possibly generate great responses that provide great initial capacity for big data classification. Moreover, the deployed model was not able to solve the overfitting problems due to the two high-dimensional datasets. Xing and Li (2019) developed a hybrid model of k-means spectral clustering (KSC) for big data classification that achieved clustering by merging the merits of KSC, which decreased the computational complexity as well as minimized the hyperspectral dimensionality of the data. Sometimes, the k-means became sensitive to the initial condition, and however hard to determine the data for classification. Jiang and Li (2019) deployed a radial basis function (RBF) neural network that performs global as well as local approximation performances and also increases the accuracy for classification. However, the RBF neural network consumes more time and is limited to complex data.
DL Techniques for Big Data Classification
In this section, the DL methods are described for big data classification with the upper hand techniques and their limitations from recent articles.
CNN Methods
Boulila et al. (2021) distributed CNNs (D-CNNs) as a type of CNN model, which utilizes an effective maximum likelihood supervised technique to efficiently classify big data and provide high accuracy. The huge amount of data led to the development of the curse of dimensionality problem, resulting in complex classification. Naronglerdrit and Mporas (2021) employed the most common DL method, a CNN (Almutairi et al., 2024; Wu, 2022), in which the model is highly preprocessed using a bottom-hat approach that enhances the image dataset to ensure effective classification. Also, the segmented process was capable of segmentation and also achieving high accuracy for classification performance. Sometimes, the CNN model was not reliable due to large memory footprints.
DCNN Methods
Ashraf et al. (2020) used the DCNN technique that participates in the big data classification problem. In this model, image-based datasets are taken; in addition, two phases, namely the feature extraction phase and the classification phase, were occupied in this process, which results in accurate performance with the help of reliable tools as well as techniques, such as benefits for high classification. Even though the model provides accurate classification, it was not explored in a large-scale image-based dataset, which was the only limitation. Jaya Sudha and Sneha (2022) also introduced a DCNN-based model; here, the neurons are considered as the basic building block to perform subsequent classification. In this approach, the classification task was effectively done using interclass and intraclass, which automatically allocates magnetic resonance imaging, ultrasound, X-rays, and other major components for effective classification. However, the privacy exploitation activity was quite improper and not a reliable model.
CNN With LSTM
Salehin et al. (2023) established a hybrid model, combining CNN with LSTM (CNN–LSTM; Rai & Chatterjee, 2022; Shahzadi et al., 2018), using real-time datasets. This technique is efficient in handling multiple sequential data and, more importantly, works in time series data and classifies effectively. However, the CNN–LSTM model was not efficient enough to handle highly complex problems, and hence, effective classification was not determined. Rai et al. (2020) also developed a CNN–LSTM model for automatic prediction problems and cardiac arrhythmias using big datasets. This model was quite efficient in predicting the disease as well as producing better accuracy performance by solving gradient issues. However, the hybrid CNN–LSTM model was inefficient in solving imbalance issues and required high computational expense during prediction.
Recurrent Neural Network (RNN) Methods
Hassib et al. (2020) developed a model for big data classification, bidirectional RNNs (BRNNs) that showed many processes for classification and utilized an optimization algorithm, WOA, for accurate selection of features that aid better results. Despite that, the WOA–BRNN model requires more time for computation while using large datasets. Narayana et al. (2022) effectively handle the classification process using a deep RNN (Deep RNN) model. The Ant Cat Swarm Optimization (ACSO) algorithm was implemented in the utilized model, which highly identifies the weight parameters and effectively tunes the model for enhanced classification, also achieving better results. In addition, this model was quite reliable and scalable because of adopting the advantages of the ACSO algorithm, which includes tolerance of faults during classification and deals well with complex patterns. However, the model provides less computational speed at a maximum cost.
Other DL Techniques for Big Data Classification
Li et al. (2017) deployed a deep computation model (DCM) that highly classified the big data for uncomplicated processing and storing the data. The utilized DCM model achieved increased performance with the help of a tensor representation, which makes it easy for classification. Moreover, the DCM took more time than the traditional CNN model due to more parameters in the tensor space. Zhu and Chen (2020) introduced the structure-informed locally distributed deep nonlinear embedding (SILDDNE) method, in which the structural and attribute characteristics of each node were taken for improved classification. The SILDDNE method directly utilizes a deep neural network to reduce several parameters and improve computational efficiency. The nodes in the network were connected only with a minimum of edges, which made it quite difficult to learn a vector representation. Brahmane and Krishna (2021) utilized a deep-stacked auto-encoder for classification, in which the model was trained using the Rider Chaotic Biography streamlining (RCBO) algorithm that effectively trained the model and increased the computational speed for classification. Although the model does not deal with large datasets, the computational expense was very high. Zhai et al. (2019) developed a restricted Boltzmann machine with a Hadoop MapReduce and fuzzy integral (MR-RBM-FI) model that involved an eligible MapReduce procedure for effective classification and achieved improved accuracy. However, the RBM model was quite challenging for the energy gradient function and increased computation time.
Federated Learning and Transfer Learning Methods
Liu et al. (2023) proposed a transfer learning-based classifier for big data classification, which deals with both static and real-time data effectively, thus increasing the overall classification. However, the model is computationally expensive while handling larger datasets. Kaleem et al. (2023) introduced a federated averaging strategy for the classification of big data, which effectively addresses challenges in real-world environment settings and data integration, thus demonstrating an improved performance in classification. However, the model requires massive computational resources and time to train the model. Dhiman et al. (2022) presented a federated learning approach to supervise the privacy preservation of large medical and health data. Finally, due to the general privacy protection problems that occur throughout the medical big data life cycle throughout the industry, appropriate options were developed at the management level.
Ensemble Methods
Demidova et al. (2016) proposed a two-level SVM classifier for big data classification, which incorporates particle swarm optimization, which reduces the processing time of the SVM classifier, which plays a major role in addressing the challenges in handling larger data. However, these models can be computationally intensive and complex to tune, and potentially suffer from less interpretability and sensitivity to initial conditions. Ponmalar and Dhanakoti (2022) proposed an ensemble SVM model for intrusion detection in a big data platform. It integrates the Chaos Game optimization algorithm and the ensemble SVM, thus identifying different types and achieving higher accuracy. However, it high computational cost and memory requirements that impact scalability for large-scale real-time data processing. Ramachandran and Manikandan (2021) proposed an ensemble classification algorithm for medical big data processing. It includes SVM and RNN classifiers for the medical data classification, precisely processes the big data, and results in a significant performance than a single classifier. However, the model may face issues such as overfitting and generalizability issues when handling new datasets. Figure 7 displays the performance ranking of the different classification models in terms of metric accuracy for big data classification.

Performance ranking chart.
In general, optimizers are employed to tune parameters for model training and enhance the computational efficiency for specific tasks. Recent articles indicate that various optimization algorithms are utilized for multiple processes, such as the selection and extraction processes, and fine-tuning the model for improved classification. The utilized optimizers from the review of recent papers such as ACSO (Narayana et al., 2022), Cuckoo–Grey wolf-based Optimization (CGWO; Banchhor & Srinivasu, 2020, 2021), WOA (Hassib et al., 2020), RCBO optimization, improved cat swarm optimization (Lin et al., 2016), genetic algorithm and gradient approximation (Almutairi et al., 2024), DGWO (Game et al., 2022), vulture optimization algorithm (VOA) and Haris Hawks optimization (HHO) are shown in Figure 5.
ACSO
The ACSO is the integration of Ant Lion Optimizer (ALO) and Cat Swarm Optimization (CSO). The ALO algorithm simulates the foraging behavior of ant lion larvae, mimicking the interactions between ant lions and ants. The CSO is a swarm intelligence-based optimization method inspired by the behavior of cats, and it is employed to address various optimization problems.
CGWO
The CGWO algorithm is an enhancement of the Grey wolf-based optimization (GWO), which incorporates a population-based algorithm, CS. The GWO harnesses the hunting behavior of grey wolves, including chasing, encircling, and attacking prey characteristics. In the CGWO algorithm, the position update mechanism of the GWO is modified by integrating the update equation of the CS, resulting in faster convergence of the CGWO algorithm.
WOA
Whales are regarded as some of the smartest animals due to the presence of spindle cells in their brains. These cells facilitate judgment, emotions, and social behaviors similar to those found in humans.
DGWO
The GWO algorithm is instrumental in solving various engineering problems. However, it has limitations, such as its performance in finding the global optimal solution, which affects its convergence rate. To overcome the efficiency, exploration, and exploitation limitations of the conventional GWO algorithm, a novel modified approach known as DGWO has been developed.
In this section, Table 2 provides the comparison of ML techniques for Big Data classification. Table 3 presents the comparison of DL techniques for Big Data classification. Table 4 depicts the comparison of federated learning and transfer learning techniques for Big Data classification. Table 5 displays the comparison of ensemble learning techniques for Big Data classification.
Comparison of Machine Learning Techniques for Big Data Classification.
Comparison of Machine Learning Techniques for Big Data Classification.
Note. DT = decision tree; DTCAA = DT credit assessment approach; ANN = artificial neural network; DHBoost = distributed heterogeneous boosting-inspired ensemble classifier; XGBoost = extreme gradient boosting; Light GBM = light gradient boosting machine; SVM = support vector machine; CNN = convolutional neural network; DL = deep learning; CNB = correlative Naïve Bayes; CG-CNB = Cuckoo–Grey wolf-based CNB; CGCNB-MRM = Cuckoo–Grey wolf-based Correlative Naive Bayes classifier with MapReduce model; KNN = K-nearest neighbor; RF = random forest; Xception-LSTM = extreme inception and long short-term memory; LHNFCSF = linguistic hedges neuro-fuzzy classifier with selected features; RMSE = root mean square error; MLNN = morphological-linear neural network; LMNN = linear-morphological neural network; KSC = K-means spectral clustering; RBF = radial basis function.
Comparison of Deep Learning Techniques for Big Data Classification.
Note. D-CNN = distributed convolutional neural network; CNN = convolutional neural network; BDA-BMIC = big data architecture-based biomedical image classification; ECG = electrocardiogram; BRNN = bidirectional recurrent neural network; WOA = Whale optimization algorithm; Deep RNN = Deep recurrent neural network; ACSO = Ant Cat Swarm Optimization; DCM = deep computation model; MR-RBM-FI = restricted Boltzmann machine with Hadoop MapReduce and fuzzy integral; SVM = support vector machine; SMOTE = synthetic minority oversampling technique.
Comparison of Federated Learning and Transfer Learning Techniques for Big Data Classification.
Note. TLC=thin layer chromatography; FedAvg=federated averaging; cPDS = cluster primal-dual splitting.
Comparison of Ensemble Learning Techniques for Big Data Classification.
Note. SVM = support vector machine.
This section elaborates on the analysis and discussion of big data classification techniques utilizing different research works based on classification models, datasets, and performance metrics.
Analysis Based on Metrics
This section presents a quantitative analysis of the selected research articles using various metrics to provide insights into big data classification techniques, as illustrated in Table 6. The graphical illustration is depicted in Figure 8.

Analysis based on metrics.
Analysis Based on Metrics.
The analysis presented in this section demonstrates the performance metrics of each established method with accuracy, specificity, and sensitivity, which are described in the context below.
Table 7 clearly illustrates the dataset analysis, highlighting the number of papers utilizing the specific database. Furthermore, the real-time datasets are utilized for most of the big data classification models.
Analysis Based on the Dataset.
Analysis Based on the Dataset.
The experimental analysis is performed on PyCharm software version 3.9.0, using the Python programming language for the implementations. Furthermore, the Windows 11 operating system, equipped with 16 GB of RAM, is utilized for the implementation of the classification models evaluation.
Analysis of Different Classification Models
This section discusses the classification models used in different big data classification works. Several established methods, such as DT, CNN–LSTM, DCNN, Light GBM with CNN (LGBCNN), VOA–LGBCNN, and HHO–LGBCNN, are analyzed with various training percentages (TPs).
Performance Analysis of DT
The performance analysis of the DT method using TP of 40, 50, 60, 70, 80, and 90 for big data classification in terms of metric accuracy, sensitivity, specificity, and F1-score is revealed in Tables 8–11, respectively. Here, the accuracy of DT with maximum TP-90, the log loss criterion is 77.65%, entropy is 78.78%, and 87.54%. The accuracy performances at each criterion differ and provide low to high values. However, the sensitivity of log loss, entropy, and gini criteria is 77.87%, 79.25%, and 93.47%, whereas the specificity is 77.43%, 78.31%, and 81.62%, respectively. The F1-score at epochs 300, 400, and 500 is 75.27%, 75.64%, and 87.82%, respectively. This analysis shows the achieved results of DT for big data classification.
Accuracy Analysis of DT.
Accuracy Analysis of DT.
Note. DT = decision tree; TP = training percentage.
Sensitivity Analysis of DT.
Note. DT = decision tree; TP = training percentage.
Specificity Analysis of DT.
Note. DT = decision tree; TP = training percentage.
F1-Score Analysis of DT.
Note. DT = decision tree; TP = training percentage; CNN–LSTM = convolutional neural network with long short-term memory.
The performance analysis for the CNN–LSTM method with TP of 40, 50, 60, 70, 80, and 90, along with varying epochs 300, 400, and 500 in terms of metric accuracy, sensitivity, specificity, and F1-score, is revealed in Tables 12–15, respectively. At maximum TP-90 with epoch 300, the performance of the CNN–LSTM model is 78.06%, 77.63%, and 78.49%, whereas the accuracy, sensitivity, and specificity at epoch 400 are 78.97%, 79.20%, and 78.73%. At epoch 500, the accuracy, sensitivity, and specificity are reported as 83.19%, 82.20%, and 84.17%, respectively. The F1-score of the CNN–LSTM model concerning epochs 300,400, and 500 is 74.21%, 74.63%, and 76.19%, respectively. This analysis shows the achieved results of CNN–LSTM for big data classification.
Accuracy Analysis of CNN–LSTM.
Accuracy Analysis of CNN–LSTM.
Note. CNN–LSTM = convolutional neural network with long short-term memory; TP = training percentage.
Sensitivity Analysis of CNN–LSTM.
Note. CNN–LSTM = convolutional neural network with long short-term memory; TP = training percentage.
Specificity Analysis of CNN–LSTM.
Note. CNN–LSTM = convolutional neural network with long short-term memory; TP = training percentage.
F1-Score Analysis of CNN–LSTM.
Note. CNN–LSTM = convolutional neural network with long short-term memory; TP = training percentage.
The performance analysis for the DCNN method with TP of 40, 50, 60, 70, 80, and 90, along with varying epochs of 300, 400, and 500 in terms of metric accuracy, sensitivity, specificity, and F1-score is revealed in Tables 16–19. At maximum TP-90 with epoch 300, the accuracy, sensitivity, and specificity of the established model are 79.67%, 80.48%, and 78.85%, whereas the performance at epoch 400 is 79.85%, 80.69%, and 79.01%, and at epoch 500, the performance percentage is 90.02%, 87.17%, and 92.87%, respectively. At epochs 300, 400, and 500, the DCNN model achieved an F1-score of 78.63%, 79.03%, and 83.21%, respectively. This analysis shows the achieved results of DCNN for big data classification, which is effectively higher than the previous model.
Accuracy Analysis of Deep CNN.
Accuracy Analysis of Deep CNN.
Note. CNN = convolutional neural network; TP = training percentage.
Sensitivity Analysis of Deep CNN.
Note. CNN = convolutional neural network; TP = training percentage.
Specificity Analysis of Deep CNN.
Note. CNN = convolutional neural network; TP = training percentage.
F1-Score of Deep CNN.
Note. CNN = convolutional neural network; TP = training percentage; CNN–LSTM = CNN with long short-term memory.
The performance analysis of the LGBCNN method with TP of 40, 50, 60, 70, 80, and 90, along with varying epochs of 300, 400, and 500 for big data classification in terms of metric accuracy, sensitivity, specificity, and F1-score is revealed in Tables 20–23, respectively. At maximum TP-90 with epoch 300, the accuracy, sensitivity, and specificity of the established model are 82.96%, 80.10%, and 85.82%, whereas the performance at epoch 400 is 83.41%, 80.62%, and 86.20%, and at epoch 500, the performance percentage is 89.85%, 83.38%, and 96.31%, respectively. The F1-score of the LGBCNN method at epoch 300 is 75.63%, at epoch 400 is 76.10%, and at epoch 500 is 87.67%. This analysis shows the achieved results of LGBCNN for big data classification.
Accuracy Analysis of LGBCNN.
Accuracy Analysis of LGBCNN.
Note. LGBCNN=light gradient boosting machine with convolutional neural network; TP=training network.
Sensitivity Analysis of LGBCNN.
Note. LGBCNN=light gradient boosting machine with convolutional neural network; TP=training network.
Specificity Analysis of LGBCNN.
Note. LGBCNN = light gradient boosting machine with convolutional neural network; TP = training network.
F1-Score of LGBCNN.
Note. LGBCNN = light gradient boosting machine with convolutional neural network; TP = training network; CNN–LSTM = convolutional neural network with long short-term memory.
The performance analysis for the VOA–LGBCNN with TP of 40, 50, 60, 70, 80, and 90, along with varying epochs 300, 400, and 500 for big data classification, in terms of metric accuracy, sensitivity, specificity, and F1-score, is revealed in Tables 24–27, respectively. At maximum TP-90 with epoch 300, the accuracy, sensitivity, and specificity of the established model are 78.82%, 79.58%, and 88.38%, whereas the performance at epoch 400 is 79.37%, 79.72%, and 79.02%, and at epoch 500, the performance percentage is 88.38%, 94.46%, and 82.30%, respectively. This analysis represents the achieved results of VOA–LGBCNN for big data classification.
Accuracy Analysis of VOA–LGBCNN.
Accuracy Analysis of VOA–LGBCNN.
Note. VOA–LGBCNN = vulture optimization algorithm–light gradient boosting machine with convolutional neural network; TP = training network.
Sensitivity Analysis of VOA–LGBCNN.
Note. VOA–LGBCNN = vulture optimization algorithm–light gradient boosting machine with convolutional neural network; TP = training network.
Specificity Analysis of VOA–LGBCNN.
Note. VOA–LGBCNN = vulture optimization algorithm–light gradient boosting machine with convolutional neural network; TP = training network.
F1-Score of VOA–LGBCNN.
Note. VOA–LGBCNN = vulture optimization algorithm–light gradient boosting machine with convolutional neural network; TP = training network.
The performance analysis for the HHO–LGBCNN method with TP of 40, 50, 60, 70, 80, and 90, along with varying epochs 300, 400, and 500 for big data classification in terms of metric accuracy, sensitivity, specificity, and F1-score is revealed in Tables 28–31, respectively. At maximum TP-90 with epoch 300, the accuracy, sensitivity, and specificity of the established model are 77.71%, 77.69%, and 77.72%, whereas the performance at epoch 400 is 78.17%, 78.38%, and 77.96%, and at epoch 500, the performance percentage is 86.38%, 81.43%, and 91.33%, respectively. The HHO–LGBCNN method achieved an F1-score of 74.25%, 75.10%, and 85.31% at epochs 300,400, and 500. This analysis represents the achieved results of HHO–LGBCNN for big data classification.
Accuracy Analysis of HHO–LGBCNN.
Accuracy Analysis of HHO–LGBCNN.
Note. HHO–LGBCNN = Haris Hawks optimization–light gradient boosting machine with convolutional neural network; TP = training percentage.
Sensitivity Analysis of HHO–LGBCNN.
Note. HHO–LGBCNN = Haris Hawks optimization–light gradient boosting machine with convolutional neural network; TP = training percentage.
Specificity Analysis of HHO–LGBCNN.
Note. HHO–LGBCNN = Haris Hawks optimization–light gradient boosting machine with convolutional neural network; TP = training percentage.
F1-Score of HHO–LGBCNN.
Note. HHO–LGBCNN = Haris Hawks optimization–light gradient boosting machine with convolutional neural network; TP = training percentage; CNN–LSTM = convolutional neural network with long short-term memory.
The complexity analysis of several ML and DL models is represented in Figure 9. At epoch 50, the classification models such as DT, CNN–LSTM, DCNN, LGBCNN, VOA–LGBCNN, and HHO–LGBCNN use time of 48.9, 42.9, 43.6, 44.4, 44.6, and 42.8 s, respectively. Similarly, at epoch 70, the CNN–LSTM, HHO–LGBCNN, DCNN, and LGBCNN models use a shorter time of 56.4, 56.4, and 57.6 s, respectively. Finally, at epoch 99, the HHO–LGBCNN model consistently uses a lower time of 75.9 s, compared to all other classification models. The inclusion of HHO with the LGBCNN model effectively reduces the computational time, thus increasing the model performance.

Time complexity analysis.
The classification models’ memory usage analysis of several ML and DL models, including DT, CNN–LSTM, DCNN, LGBCNN, VOA–LGBCNN, and HHO–LGBCNN models, is depicted in Figure 10. At epoch 30, the DT, CNN–LSTM, DCNN, and LGBCNN models utilize memories of 188.5 kb, 169.8 kb, and 175.90 kb, while the HHO–LGBCNN model uses less memory of 168.48 kb. Similarly, at epoch 70, LGBCNN, VOA–LGBCNN, and HHO–LGBCNN models utilize memory of 321.0 kb, 323.8 kb, and 292.6 kb, respectively. Finally, at epoch 99, the HHO–LGBCNN model utilizes less memory of 410.3 kb, highlighting its scalability in classifying big data in resource-constrained environments.

Memory analysis.
In this review, several papers with multiple DL and ML techniques for big data classification are examined and utilized. The achievements, merits, and demerits of several existing methods are determined and helpful for future implementation works. Moreover, the significant analysis of the reviewed papers shows that DL methods are very effective in the classification of big data. Specifically, the hybrid DL methods, such as CNN–LSTM and LGBCNN, are more efficient than the other methods. The DL techniques are efficient in handling multiple sequential data and, more importantly, work in time series data and classify effectively. However, these DL methods have an imbalance dataset issue and also high computational complexity during the prediction. Moreover, the existing ML method faces issues in high computational time due to the utilization of large datasets and also faces scalability issues, which affect the training and lead to inefficient classification and overfitting. Model parallelism requires frequent communication between devices to exchange intermediate outputs and pass gradients for backpropagation. This can introduce significant latency and create performance bottlenecks, especially in models with many interdependent layers, as model parallelism's scalability is constrained by the number of layers in the model. Simply adding more graphics processing units does not improve performance beyond a certain point, limiting its effectiveness for horizontal scaling. In cloud-based big data processing techniques, data from diverse, heterogeneous sources can be difficult to integrate and unify for classification. Compiling this variety of data types and formats is a complex task. Furthermore, in cloud-based preprocessing, data is stored and processed by a third-party provider, increasing its vulnerability to cyber threats, unauthorized access, and data breaches. Storing and processing large volumes of sensitive big data, such as personal or health information, is a major concern for regulatory compliance. In DL for Big Data classification, determining the optimal number of model parameters and enhancing their computational practicality present significant challenges. The large volume of big data often reflects and amplifies past societal biases. As a result, these models can learn and reinforce unfair patterns, leading to biased and discriminatory results. Additionally, the black box nature of deep neural networks makes it hard to understand how they make classification decisions. This lack of clarity complicates accountability and debugging. Without transparency, trust in critical applications, such as medical diagnoses or loan decisions, breaks down, making it very hard to justify or audit a model's possibly biased outcomes. The hybrid DL model has specific challenges in big data classification, which include streaming data, high-dimensional data, and scalability issues.
Research Gaps and Future Works
This section explains the limitations and future work of existing methods in big data classification using the recently reviewed papers, which is thus helpful for the development of innovative models in future classification.
Limitations and Future Work of DT
Limitations
The DT model with optimization technique for big data classification is occupied with certain limitations, such as finding the global optima impacting the convergence rate, providing low accuracy, and requiring maximum computational expenses for the classification procedure (Game et al., 2022).
The DT credit assessment approach (DTCAA) model used in credit assessment problems using big data components drives complications during the process of classification, including the model's intention to minimize the complexity of the attributes. However, the attributes are not enough for the data mining model, in addition to suffering from low sensitivity (Chern et al., 2021).
In the C4.5 DT model, the minority range of this model is specifically delayed for measuring some kind of impurities, which results in low accuracy for classification (Rawal & Agarwal, 2019).
Future work
The DT model with the DGWO algorithm will be deployed in a large-scale medical platform and will be utilized in some other efficient optimization techniques to improve the performance rate for classification (Game et al., 2022). The limitation of the DTCAA model will be improved in future work (Rawal & Agarwal, 2019).
Limitations of Gradient Boosting Methods
Limitations
The DHBoost model has automatic selection, which results in maximum computational expenses and is also sensitive to noise (Kadkhodaei et al., 2021).
The XGBoost model was quite slow for the training process to classify the big data, and also more vulnerable to overfitting problems (Ait Hammou et al., 2019).
The Light GBM model has overfitting issues and provides computational complexity. In addition, the selection of leaves during tree generation led to the majority loss of important data, reducing the efficiency of the model (Hussein & Abdulkader, 2022).
Limitations of SVM Methods
The SVM model with the DL model, CNN used for big data classification, only provides better performance using medium or small datasets, but not large datasets (Tchito Tchapga et al., 2021)
The T-BMSVM model consumes more computational time and cost due to the large number of datasets, and also reduces the scalability for classification (Sujitha & Seenivasagam, 2021)
Future work
The spark algorithm in the SVM model will be implemented in real-world applications as well as evaluated with other performance metric parameters to identify the improved performance rate of the developed model (Tchito Tchapga et al., 2021). In the T-BMSVM model, multiple datasets will be deployed for the big data classification procedure for future work (Sujitha & Seenivasagam, 2021).
Limitations and Future Work of CNB Models
Limitations
The CNB classifier was highly expensive during computation due to the developed optimizer, search, and grey wolf optimizer, which effectively tunes the model with multiple tuning parameters, leading to increases in the computational cost and time (Banchhor & Srinivasu, 2021).
The output of the CGCNB-MRM was always affected due to the quality of the datasets; if the quality was poor, the outcome of the model was also poor. This limitation can provide low accuracy and decreased efficiency of the classification model (Banchhor & Srinivasu, 2020).
Future work
In the future, the performance of the CNB model will be examined using log loss and training loss functions to exaggerate the model performance (Banchhor & Srinivasu, 2021). Also, the optimized model, CGCNB-MRM, will be implemented in DL to reduce the limitations of traditional ML techniques, to boot, improved classification (Banchhor & Srinivasu, 2020).
Limitations and Future Work of ML Techniques
Limitations
The improved KNN model only focused on a single-class classification, which decreases the sample sparsity of the category description and also results in decreased performance for classification (Xing & Bei, 2019).
Due to the complexity of the RF classifier, the prediction of each node required more computational costs and was estimated with minimum convergence speed for classification (Manikandan, 2018).
The Xception-LSTM model required high computational expenses for the classification process, and the problem can be solved using only computer hardware (Liang et al., 2018).
The outcome of the LHNFCSF model was not completely interpretable for classification and cannot be implemented in complex problems (Azar & Hassanien, 2015).
The designed MLNN with the LMNN model did not solve the overfitting problems due to the two high-dimensional datasets (Hernández et al., 2020).
In the KSC model, sometimes the k-means algorithm becomes more sensitive to the initial condition; however, hard to determine the data for classification and provides high computational expenses (Xing & Li, 2019).
RBF neural network consumes more time due to multiple hidden layers for input and output information, and is also limited for complex data (Jiang & Li, 2019).
Future work
The enhanced KNN model will be concentrated on building a multiclass classification for improving the sample sparsity for effective classification (Xing & Bei, 2019). The Xception-LSTM model will be improved with a task-specific classifier to solve the complex problems (Liang et al., 2018).
Limitations and Future Work of CNN–LSTM
Limitations
The CNN–LSTM model was not efficient in handling highly complex problems, and hence, effective classification was not determined (Salehin et al., 2023).
The hybrid CNN–LSTM model was inefficient in solving imbalance issues in several problems and required high computational expense during prediction (Rai et al., 2020).
Future work
The resampling methods and faster intensive learning model will be deployed to handle imbalance problems in data, as well as reduce computational time (Rai et al., 2020).
Limitations and Future Work of DCNN
Limitations
Even though the DCNN model provides accurate classification, it was not explored in large-scale image-based datasets, which was the only limitation (Ashraf et al., 2020).
In big data classification, the DCNN model provides improper privacy exploitation, which achieves unreliable classification in big data, and also provides a minimum accuracy rate (Jaya Sudha & Sneha, 2022).
Future work
This work will be focused on improving the model to perform with large-scale datasets and will also be developed for detection-based problems (Ashraf et al., 2020). For big data classification, effective privacy-preserving cloud-based crowd-sourcing will be deployed in the future to ensure high privacy during classification (Jaya Sudha & Sneha, 2022).
Limitations and Future Work of RNN Methods
Limitations
Ø The WOA–BRNN model required more time for computation while using large datasets. It even reduces the number of attributes for an effective process and reduces the optimal accuracy of the developed model (Hassib et al., 2020).
Ø The Deep RNN model with the ACSO algorithm required a high privacy strategy for defense or healthcare-based datasets for effective classification (Narayana et al., 2022).
Future Work
In the future, the distribution version of the WOA–BRNN model will be implemented to reduce the time consumption during classification (Hassib et al., 2020). The cryptography algorithm in the Deep RNN will be deployed for privacy concerns, and also, performance with different datasets will be analyzed.
Limitations and Future Work of Other DL Methods
Limitations
The DCM model took more time to compute than the traditional CNN model because of the high number of parameters in the tensor space (Li et al., 2017).
In SILDDNE, the nodes in the network were connected only with a small number of edges, which intensified the difficult learning process with the vector representation (Zhu & Chen, 2020).
The deep-stacked auto-encoder does not deal with large datasets due to the limited memory and requires high computational expense for classification (Brahmane & Krishna, 2021).
In the MR-RBM-FI model, the RBM was quite challenging for the energy gradient function and required maximum computation time for effective classification. This represents that the model was inefficient (Zhai et al., 2019).
Future Work
In the SILDDNE model, difficult learning processes will be solved for better accuracy performances and provide high efficiency to the model (Zhu & Chen, 2020). The incremental learning will be developed in the deep-stacked auto-encoder model by sorting the data of the existing model, which effectively processes the new incoming data (Brahmane & Krishna, 2021).
Limitations and Future Work of Transfer Learning and Federated Learning Methods
Limitations
Transfer learning model reliance on active sampling might be computationally expensive, especially when handling larger datasets or if high-frequency real-time analysis is needed (Liu et al., 2023).
Federated learning models require massive computational resources and time to train a model on a larger dataset (Kaleem et al., 2023).
Future Work
The computational overhead and computational cost challenges in the federated learning (Kaleem et al., 2023) and transfer learning model (Liu et al., 2023) are addressed by including the active sampling techniques and integrating hybrid optimization techniques to minimize the model overhead and make it scalable in resource-constrained environments.
Limitations and Future Work of Ensemble Methods
Limitations
Ensemble models may be susceptible to local optima and face computational complexity challenges due to SVM's inherent costs, as well as they will struggle with highly imbalanced data and risk of overfitting (Demidova et al., 2016).
An ensemble SVM model requires high computational resources and memory requirements (Ponmalar & Dhanakoti, 2022).
Ensemble classification algorithms face overfitting and generalizability challenges while handling new or unseen data.
Future Work
The local optima and computational complexity challenges in the ensemble methods can be solved by hybridizing with other metaheuristics, adopting parallel and distributed SVM approaches, using data reduction techniques, integrating cost-sensitive learning, or employing explainable artificial intelligence (AI) methods (Demidova et al., 2016). The high computational resources, memory requirements, and generalizability challenges (Ponmalar & Dhanakoti, 2022) are addressed through the integration of an advanced optimization algorithm with a lightweight architecture model (Ramachandran & Manikandan, 2021).
Applications
Healthcare
Big data changes healthcare by allowing predictive analysis for early disease detection, looking at patient vitals from wearables, and improving clinical operations. By gathering large amounts of data from Electronic Health Records, genomic sequencing, and medical imaging, big data enables personalized treatment plans for individual patients. This analysis also helps hospitals manage staff and resources better by predicting patient admission trends. Real-time patient monitoring through big data can alert medical professionals, allowing for immediate intervention for at-risk patients and reducing medical errors. Additionally, it speeds up medical research and drug discovery by examining large datasets of trial results and genetic information, which cuts down development time and costs.
Security
In the security field, big data analytics improves threat detection and prevention by processing vast amounts of organized and unorganized data in real time. By examining network traffic, user behavior, and system events, algorithms can spot unusual patterns that suggest cyberattacks or insider threats. This method shifts cybersecurity from reactive to proactive, enabling organizations to foresee and reduce risks before they cause serious damage. Big data also drives advanced fraud detection systems, which can analyze millions of transactions each second to identify irregular activities instantly, significantly lowering financial risk.
Social media
Big data from social media serves many purposes, such as market research, personalized marketing, and sentiment analysis. Companies analyze large amounts of social media interactions, including posts, comments, and likes, to understand what consumers want, their preferences, and their behavior. For example, Netflix uses viewing habits to create personalized content recommendations, which improves user engagement and satisfaction. In the same way, brands monitor public opinion through sentiment analysis. This helps them respond quickly to negative feedback and manage their brand reputation. Social media data also helps identify market trends and predict future customer needs, guiding product development and marketing strategies.
Conclusion
The research survey aims to develop a highly efficient model for big data classification to classify important data in numerous fields, such as healthcare, marketing, and other fields, with improved security. In this research, several reviewed papers from recent years with multiple DL and ML techniques, as well as the established methods, are examined. In this paper, the achievements, merits, and demerits of several existing methods are determined, which are helpful for future implementation works. These analyses will make future research quite easy, simple, and straightforward approaches to effectively handle classification with more adapted strategies, such as preprocessing, feature extraction, and feature selection. Moreover, the significant analysis of the reviewed papers shows that DL methods are very effective in the classification of big data. Specifically, hybrid
Future Work
Future work will focus on developing dynamic and heterogeneous ensemble techniques that adapt in real-time to evolving data streams, combining diverse base learners and incorporating advanced optimization strategies, such as hybrid cloud-edge training and parallel processing, to enhance the scalability and performance of the model big data classification and to address the problems associated with Big Data classification, including high dimensionality, streaming data analysis, and scalability of DL models. The feature research will be considered, integrating self-supervised learning and automated feature engineering to enable models to automatically discover optimal features and learn from massive unlabeled datasets, which is particularly vital for overcoming labeling bottlenecks and handling complex, high-dimensional data efficiently within distributed computing frameworks. This integrated approach, combining intelligent automation, adaptive learning, and robust optimization, will pave the way for a new generation of big data classification models that are more autonomous, resilient, and effective.
Supplemental Material
sj-docx-1-web-10.1177_24056456261462613 - Supplemental material for An Empirical Study of Big Data Classification Methods and Challenges
Supplemental material, sj-docx-1-web-10.1177_24056456261462613 for An Empirical Study of Big Data Classification Methods and Challenges by Pradnya Bhangale, Pradheep Manisekaran and RP Sharma in Web Intelligence
Footnotes
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Supplemental Material
Supplemental material for this article is available online.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
