Abstract
In the rapidly evolving digital landscape, virtual reality (VR) has emerged as a transformative tool in education, particularly in language acquisition. Traditional language learning methods often struggle to provide immersive, interactive, and engaging experiences that foster real-world communication skills. VR environments are increasingly transforming English language acquisition by providing immersive and interactive learning experiences that enhance communication skills development. This paper explores the impact of immersive VR technologies on English language acquisition, focusing on communication skills development through dynamic and interactive learning experiences. A novel Dynamic Monarch Butterfly Optimized Bidirectional Gated Recurrent Unit (DMBO-BiGRU) model is presented to analyze learner interactions within VR environments, assessing communication patterns, language complexity, and learning progression. The dataset, collected from student interactions in VR settings, was preprocessed using noise filtering and normalization to ensure data quality. t-Distributed Stochastic Neighbor Embedding (t-SNE) is employed to extract key features, such as fluency, vocabulary usage, and contextual appropriateness. The DMBO-BiGRU model leverages adaptive learning strategies to provide tailored feedback and optimizes language acquisition based on learners’ communication skills. The effectiveness of the suggested method is demonstrated through experimental results, which indicate good accuracy and positive VR learning scores. The results underscore how immersive, customized, and interactive experiences provided by VR-integrated deep learning models hold the potential to transform language acquisition, showing accuracy of 97.2%, VR score of 97.52%, training gain of 98.49%, detection score of 92.6%, and a retention rate of 96% at iteration 13. This work contributes to the growing field of technology-enhanced education, emphasizing the role of VR and AI-driven methods in enhancing English communication skills.
Keywords
Introduction
Technological advancements have initiated a significant shift in language learning, with virtual reality (VR) emerging as a revolutionary educational tool. As the global demand for English language competence increases, educators and scholars are exploring new methods to enhance language learning.
1
VR presents an interactive and immersive learning atmosphere, providing students with a unique opportunity to engage with the language in a tangible and meaningful manner, as illustrated in Figure 1. The core component of VR is the role of technology in its environment.
2
VR refers to a computer-generated simulation that provides a synthetic, three-dimensional environment where users can engage with a digital landscape that appears realistic and immersive. Sensory components such as visual and, in some cases, haptic input enhance the user’s sense of presence.
3
By offering interactive scenarios, VR allows for deeper cognitive and experiential learning compared to traditional digital learning approaches that rely on textbooks or screen-based materials. VR enables language learners to practice their skills through stable encounters and meaningful dialogues that replicate real-world experiences.
4
VR in English language acquisition.
The context in which communication is provided by the environment significantly influences language development. Traditional classroom-based language instruction often lacks real-world engagement, which can hinder the development of communication capabilities. 5 However, through VR environments, language learners can apply their skills in everyday situations, such as online marketplaces, workplaces, social activities, or travel scenarios. These immersive environments encourage the use of natural language by giving students the opportunity to utilize their knowledge in real-world contexts. 6 The adaptive nature of VR settings accommodates a variety of learning styles and levels. English language acquisition refers to the acquisition of skills in speaking, listening, reading, and writing. 7 Traditional methods often focus on grammar practice and rote memorization, which may not foster natural language use. In contrast, VR-based learning promotes engagement and active participation. 8 VR facilitates improvements in comprehension, fluency, and pronunciation through simulations of authentic conversations. Additionally, VR encourages self-observation and mitigates the language anxiety that can arise in traditional learning environments. 9
These components collectively establish VR as a valuable tool for enhancing English language learning. VR fosters learning experiences that transcend academic knowledge by immersing students in an interactive, realistic environment. 10 Such direct interactions allow language learners to develop their skills in conditions that reflect real-world situations. Furthermore, the personalized and adaptable nature of VR platforms ensures that learners receive tailored experiences according to their skill levels and learning paces. 11 As VR technology continues to advance and become more accessible, it can revolutionize English language instruction. This paper examines the advantages, challenges, and potential applications of VR environments in the context of learning the English language. 12 However, varying personal learning styles and technological accessibility may limit generalizations. Additionally, issues such as motion sickness and cognitive stress could negatively affect the user experience. A unique DMBO-BiGRU model is introduced to evaluate communication patterns, language complexity, and learning development through learner interactions within VR environments.
Related works
Recent advancements in immersive technologies and artificial intelligence (AI) have significantly influenced language education research. This section synthesizes key studies exploring the integration of virtual reality (VR), augmented reality (AR), and AI in language learning contexts. Shi et al. 13 developed a VR-based English-speaking training and assessment system to examine the impact of cognitive style (field-dependent vs field-independent) and test settings (virtual vs real environments) on learning outcomes. Their findings revealed that field-independent learners exhibited superior performance in virtual testing environments, whereas field-dependent learners demonstrated better outcomes in traditional settings. The study highlighted the critical interplay between cognitive styles, environmental contexts, and technological interventions in shaping language acquisition efficacy. In a comparative analysis of vocabulary acquisition, Tai et al. 14 investigated the effectiveness of mobile-rendered head-mounted displays (HMDs) for seventh-grade EFL learners. The experimental group using VR applications outperformed the control group (video-based learning) in both immediate retention and long-term recall of vocabulary. Despite the VR group’s success, the video group reported mixed perceptions of multimodal support, underscoring the need for further exploration of HMDs’ pedagogical implications.
Lu 15 proposed a machine learning (ML) approach, termed the Re-defined Harris Hawks Optimized Intelligent Support Vector Machine (RHH-ISVM), to enhance English language learning in VR environments. By analyzing learners’ interaction patterns, language proficiency metrics, and communication behaviors, the model demonstrated high accuracy in skill classification and progress evaluation. This work emphasized the synergistic potential of immersive VR and adaptive ML algorithms for personalized language instruction. Chen et al. 16 integrated AI and VR technologies to create a robot-assisted language learning (RALL) system for training English-speaking tour guides. During a 10-week intervention, learners engaged with a 3D VR environment and AI-driven robot interactions, resulting in measurable improvements in speaking fluency and learner autonomy. Participants acknowledged the system’s capacity to foster active participation and flexible learning pathways, validating RALL’s role in complementing traditional pedagogies.
Divekar et al. 17 introduced the Cognitive Immersive Language Learning Environment (CILLE), combining Extended Reality (XR) and AI to facilitate non-dyadic multimodal interactions for Chinese as foreign language (CFL) learners. A 7-week mixed-methods study with university students revealed significant gains in vocabulary, comprehension, and conversational competence, demonstrating XR’s ability to deliver immersive experiences without reliance on intrusive hardware. Qiao and Zhao 18 explored AI-driven interventions for enhancing self-regulation and L2 speaking skills among Chinese EFL learners. Their results indicated that AI-based training not only improved oral proficiency but also strengthened metacognitive strategies and learner autonomy, suggesting AI’s transformative potential in fostering independent language acquisition.
Lee 19 compared the efficacy of printed materials and mobile AR games for narrative-driven language learning with South Korean undergraduates. While both media elicited positive engagement, print-based resources were perceived as more conducive to structured language practice. The study advocated for location-based AR applications to enrich contextualized learning experiences. Smutny 20 conducted a market analysis of VR educational applications (2019–2021), identifying history, art, medicine, and natural sciences as dominant domains. Notably, over 50% of sampled applications were freely accessible and English-centric, highlighting both opportunities and linguistic biases in VR’s educational adoption.
Ironsi 21 investigated VR’s role in developing speaking skills through teacher and student perspectives. While VR enhanced oral communication outcomes, the study cautioned against technological overreliance and emphasized balancing immersive tools with pedagogical fundamentals. Beyond language-specific applications, Li et al. 22 designed a distributed VR framework for entrepreneurship education, employing 3D point cloud processing and edge contour recognition to optimize collaborative virtual spaces. Their system reduced time overhead and information saturation, aligning with evolving demands for scalable, immersive learning infrastructures. Hu et al. 23 proposed a multimodal feature integration model to assess learner concentration in VR environments. By analyzing visual and interaction data, their method achieved superior accuracy in concentration recognition, correlating high immersion levels with improved learning outcomes.
Collectively, these studies underscore VR, AR, and AI’s transformative potential in language education while highlighting contextual challenges such as cognitive diversity, technological accessibility, and pedagogical integration.
Methodology
The virtual reality environments were specifically designed to mimic real-world scenarios that language learners are likely to encounter, such as social gatherings, professional meetings, and everyday interactions in marketplaces. The environments were selected based on their ability to stimulate authentic language use and encourage engagement. The classification of communication patterns was based on several established linguistic features, including but not limited to fluency rates, vocabulary richness, and contextual appropriateness. We utilized a framework to categorize these patterns into low, medium, and high communication levels, which enabled us to assess the effectiveness of the VR interactions more accurately. Noise filtering and normalization are two steps in data preparation that improve data quality for precise learner assessment. t-SNE is used to extract important language properties, such as word richness, contextual appropriateness, and fluency. In immersive VR environments, the suggested DMBO-BiGRU model analyzes learner communication, offers feedback that is adjusted, and improves individualized language acquisition by utilizing adaptive learning techniques. Figure 2 shows the overall flow for VR in English language acquisition. Overall flow for VR in English language acquisition.
Dataset
The collection of learner interactions, linguistic characteristics, engagement metrics, and AI-generated feedback, the VR-LearnEng dataset facilitates English language learning in VR environments and analyzes the evolution of communication abilities. A learning success categorization (low, medium, high) is used to monitor progress, and it also contains learner demographics, VR session facts, language complexity, interaction metrics, and individualized feedback. A great resource for assessing and improving immersive language learning experiences, which consists of 1000 structured samples in CSV format, offers insightful information on vocabulary variety, fluency, contextual comprehension, and learning development.
Source: https://www.kaggle.com/datasets/ziya07/vr-learneng-dataset/data.
Data preprocessing
Data preprocessing is necessary for accuracy and reliability when measuring the impact of VR environments on English language learning. Z-score normalization is conducted on linguistic performance measures to normalize data and make comparisons more meaningful by bringing the data in view around zero with unit variance. The Wiener filter also enhances the measurement of response time and pronunciation quality by reducing noise in speech recordings and sensor inputs. By enhancing the robustness of the dataset, this preprocessing facilitates easier analysis of the outcomes of VR-based language learning with increased precision.
Z-score normalization
Z-score normalization is a statistical technique that standardizes the linguistic performance measures by transforming them into a function of the mean and standard deviation of the dataset. This method ensures that the data possesses a mean of zero and a standard deviation of one, which allows for meaningful comparison between various linguistic indicators. The formula for this transformation is given by:
Weiner filter
The Wiener filter is employed to enhance the quality of the data collected by reducing background noise in the recorded interactions. This filter functions by estimating the desired signal from the noisy input based on statistical estimations of the signal and noise present. The Wiener filter optimally minimizes the mean square error between the estimated and the actual signal, thereby improving the clarity of spoken interactions in the VR environment. The filter model is represented in the frequency domain by the expression:
It should be noted that the Wiener filter in equation (2) reduces to the inverse filter in equation (3), provided that detector noise is absent
Here,
The Wiener filter was implemented using a 512-sample Hamming window with 50% overlap to mitigate ambient noise from VR headset microphones, particularly addressing the 60 Hz electrical hum prevalent in laboratory environments. The filter’s frequency response was calibrated using noise profiles captured during system idle states, ensuring adaptive suppression of device-specific artifacts.
Feature extraction using t-SNE
A useful dimensionality reduction method for displaying high-dimensional data is t-SNE. t-SNE can extract useful characteristics from learner interactions, speech patterns, and engagement metrics while exploring how VR environments affect English language learning. t-SNE assists in identifying student groups according to competency, immersion levels, or engagement patterns by translating intricate linguistic and behavioral data into a lower-dimensional environment. Groups of neighboring data points are called neighbors, and the t-SNE approach models their distribution. The initial high-dimensional space is treated as a Gaussian distribution, while the low-dimensional output space is represented as a t-distribution. This four-step process intends to find the transformation that helps distribute data points more uniformly in low-dimensional conditions by reducing the gaps between all points and transferring high-dimensional space to low-dimensional space.
The highly dimensional Euclidean distance between each point of data in equation (4) is converted into a conditional probability of resemblance or the conditional probability
The formula for calculating
Here,
Using the gradient descent approach, SNE finds the smallest sum of the conditional probability differences by minimizing the
Finally, the Gaussian variance
DMBO-BiGRU
BiGRU
BiGRU is used to investigate how VR environments affect learning English. Through the utilization of VR immersive learning experiences, the BiGRU model analyzes advances in vocabulary retention, pronunciation, and understanding by processing sequential language learning data. Recurrent neural networks, or RNNs, use sequence data as input and process it repeatedly along the sequence’s development path. Each node is connected to the others by a chain. Bidirectional and long short-term memory networks (LSTM) are two common varieties of recurrent neural networks. Nevertheless, it has issues with gradient expansion and disappearance. The issues of disappearing gradients and explosion of gradients can be resolved using LSTM, an enhanced RNN-based variant. It performs well in the domains of machine translation, language modeling, and speech recognition. In addition, it is applied to a variety of time series attribute issues. Understanding bidirectional GRU requires knowledge of gated recurrent units (GRUs). A simpler variant of the LSTM network is the GRU. The input, forget, and output gates constitute an LSTM. The update gate and the reset gate are the only two gates in GRU. The input and forget gates of an LSTM function are similar to the update gate. It decides which data should be maintained as current and which should be left behind. When determining the current time, the reset gate decides if a certain amount of the previous data is unnecessary. Due to having fewer gating points, GRU executes computations faster than LSTM. The following is the formula for calculating the GRU equations (9)–(14).
Here,
The sigmoid activation function, denoted as BiGRU structure.
DMBO
In VR environments, English language acquisition is enhanced through immersive learning experiences. When it comes to language learning, VR exhibits traits that set it apart from traditional methods. Learners engage with dynamic, interactive scenarios that mimic real-world communication. The structured yet flexible learning process in VR has an impact on the development of effective language acquisition strategies. Exploration and exploitation are two crucial techniques that guide the learning process, much like in other adaptive learning systems. VR-based learning utilizes these two techniques: exploring new linguistic contexts and adjusting learned skills. Equations (15) and (16) are used to initialize the two phases of learning, referred to as
There are several solutions for the entire population, represented by
The symbols
The rand is a number extracted from a distribution, and the variable peri determines the period of migration to regulate MBO functioning. The peri value in the original MBO was
The
The randomly selected solution, represented by the symbol
The adjustment rate is represented by
The weighting factor
The maximum value for the step is represented by
Hyperparameter for DMBO-BiGRU.
Experimental findings
Experimental setup for the suggested DMBO-BiGRU method.
Three classes, designated as 0, 1, and 2, constitute the confusion matrix in Figure 4, which illustrates how well a model predicts various categories. Class 1 is incorrectly categorized as class 2 three times; however, classes 0 and 2 have excellent accuracy (80 and 117 correct predictions, respectively), indicating that the model successfully separates these groups. Such a matrix can display how successfully a machine learning model categorizes language competence levels or learning outcomes in the context of exploring how VR environments affect English language acquisition. Although there could be places for improvement at intermediate competence levels (class 1), the low misclassification indicates that VR-based language instruction could enhance specific language abilities. Confusion matrix for VR environments affect English language acquisition.
The distribution of language skill levels among students is depicted in Figure 5, where the majority (611 learners) are in the intermediate group, followed by the advanced (369 learners), and the beginners (20 learners) category is the smallest. This distribution implies that learners at intermediate and advanced levels could discover VR-based learning more popular or more successful when it comes to VR environments and English language acquisition. The low percentage of beginners could indicate that beginning language learners who need more fundamental assistance can benefit from immersive VR experiences for language learning, or it could have restricted access to VR resources. Distribution of language proficiency levels performance.
A model’s ability to differentiate between various degrees of language skill in a VR-assisted English language learning environment is demonstrated by the Precision-Recall (PR) curve in Figure 6. The model appears to correctly identify intermediate and advanced learners with minimal misclassification, as indicated by the high AUC values (1.00 for these learners). However, the beginner class’s lower AUC (0.85) suggests that it could be challenging to differentiate starting learners, either as a result of overlapping linguistic traits or fewer data points. This implies that even while VR environments help intermediate and advanced learners very well, additional effort could be required to improve language learning in the early stages. Graphical representation of Precision-Recall curve.
The average VR learning scores for each level of English competence are displayed in Figure 7, with no difference between beginners (85.0), intermediate learners (84.5), and advanced learners (84.7). It indicates that regardless of skill level, virtual reality (VR) environments offer a consistently successful learning experience. The minimal variations imply that the advantages of VR-based language learning are comparable for all students, confirming its potential as an inclusive language learning assistant. However, an additional investigation could determine if certain VR elements have a greater impact on various skill levels. Average VR learning score by proficiency.
Accuracy
The accuracy of grammar, vocabulary, pronunciation, and sentence structure in VR-based English language learning is referred to as accuracy. It assesses how effectively students adhere to accepted language use norms. Through immersive voice recognition, writing exercises, and interactive learning, virtual reality improves accuracy. Accuracy performance is shown in Figure 8 and Table 3. The ANN performed well in pattern recognition and classification, demonstrated by its 96.37% accuracy rate. The F-D2QLN’s slightly lower accuracy of 95.6% demonstrated its adaptability to dynamic learning conditions with some minor drawbacks. With an accuracy of 97.2%, the suggested DMBO-BiGRU performed better than both, proving its greater capacity to capture temporal dependencies in language acquisition. Therefore, it is a very useful tool for evaluating learning outcomes in virtual reality environments. Accuracy performance for VR in English language acquisition. Overall values for the proposed and the existing methods.
VR score
The VR score is a numerical indicator that evaluates how VR environments affect learning English. Vocabulary retention, pronunciation accuracy, understanding, engagement levels, and participation in the VR environment are some of the criteria utilized to assess students’ success.
Training gain
The improvement in English language proficiency that students attain following exposure to VR-based instruction in comparison to their starting skill level is referred to as training gain.
Detection score
The detection score indicates how well the system can evaluate and categorize students’ replies or interactions inside the VR environment.
The outcomes indicate that in determining how VR environments impact English language learning, the DMBO-BiGRU model is superior to the ANN, as illustrated in Figure 9 and Table 3. The ability of the model to determine the effectiveness of VR-based learning is expressed in the VR score, which is slightly greater for DMBO-BiGRU (97.52) than for ANN (97.37). Moreover, DMBO-BiGRU (98.49) performs better than ANN (98.34) in training gain, which signifies learning progressions over time. Moreover, DMBO-BiGRU has a significantly superior detection score (92.6) compared to ANN (85.47), which signifies accuracy in identifying key language learning patterns. This means that DMBO-BiGRU is a better choice for assessing learning development in immersive settings since it provides a more accurate and reliable assessment of VR-based language learning. Performance comparison for ANN and DMBO-BiGRU.
Retention rate
In the context of VR environments’ effects on English language learning, retention rate is the proportion of acquired vocabulary, language skills, or understanding that students can successfully retain and use after a certain amount of time. Throughout several iterations, the suggested DMBO-BiGRU model outperforms the F-D2QLN model in terms of retention rates when used in VR environments for English language learning, as shown in Figure 10 and Table 4. At iteration 1, DMBO-BiGRU outperforms F-D2QLN by showing more engagement from the beginning with a 73% retention rate compared to 70% for F-D2QLN. With 79% at iteration 2, 80% at iteration 4, and 82% at iteration 9, DMBO-BiGRU consistently performs better than F-D2QLN, which lags behind with 75%, 74%, and 75%, respectively. With DMBO-BiGRU achieving 93% at iteration 12 and 96% at iteration 13 compared to F-D2QLN’s 84% and 94%, an evident difference manifests in subsequent iterations, highlighting the model’s higher ability to sustain learner interest. This suggests the longer-term retention and enhanced language learning through the integration of VR-facilitated learning and BiGRU’s ability to recognize long-term patterns in sequential data. Retention rate performance for VR in English language acquisition. Retention rate versus number of iterations.
Discussion
English language learning is greatly impacted by VR environments that offer immersive, dynamic, and engaging educational opportunities. By practicing real-life communication in simulated environments, VR helps students improve their vocabulary, pronunciation, and fluency. The existing methods are F-D2QLN 25 and ANN. 24 The F-D2QLN’s high computational complexity is a drawback when it comes to examining how VR environments affect English language acquisition. Real-time adaption in immersive VR environments is difficult due to the increased processing cost caused by the fuzzy logic integration. Furthermore, D2QLN and other deep reinforcement learning models need a lot of training data and adjustment, which can make them less applicable to a wide range of language learners. 25 It can be difficult to obtain the vast volumes of high-quality labeled data that ANNs need for efficient training in instructional virtual reality environments. It’s also challenging to understand how certain VR components affect language learning results because of their black-box nature. 24 The DMBO-BiGRU improves sequence learning efficiency, captures bidirectional contextual dependencies, and dynamically adjusts model parameters to enhance the analysis of how virtual reality settings affect learning English. While BiGRU efficiently stores past and future contextual information, DMBO guarantees quicker convergence and better feature selection, which improves pronunciation recognition, sentiment analysis, and engagement monitoring in VR-based language learning.
Conclusion
VR has become an essential tool in education, especially in language learning, in the quickly changing digital world. It investigates how immersive VR technologies affect learning English, emphasizing the improvement of communication skills through engaging and dynamic learning environments. Noise filtering and normalization were used as preprocessing techniques to assure the quality of the dataset, which was gathered via student interactions in virtual reality environments. Fluency, language use, and contextual appropriateness are among the important variables that are extracted using t-SNE. The DMBO-BiGRU model evaluates language complexity, communication patterns, and learning development through learner interactions in VR environments. The suggested method is compared to the existing method in terms of accuracy (97.2%), VR score (97.52%), training gain (98.49%), detection score (92.6%), and the retention rate at iteration 13 (96%). VR-based English language instruction provides immersive and interactive learning; its drawbacks include expensive implementation costs, the requirement for specialist technology, and the possibility of motion sickness in users. For future enhancement of AI-driven adaptive learning in VR, increasing accessibility with affordable solutions, and incorporating real-time feedback mechanisms should be the primary targets.
Footnotes
Funding
The author received no financial support for the research, authorship, and/or publication of this article.
Conflicting interests
The author declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Data Availability Statement
The authors declare that the data supporting the findings of this study are available within the article. The raw/derived data supporting the findings of this study are available from the corresponding author at request.
