Abstract
Forum discussions in Massive Open Online Courses (MOOCs) play a crucial role in promoting learning engagement and academic achievement. In particular, discussion topics significantly influence learners’ emotional and cognitive engagement. However, the complex interrelationships among these factors remain underexplored. This study introduces an innovative two-step methodological approach to investigate the relationships between topic complexity, emotional engagement, cognitive engagement, and academic achievement in MOOC discussions. Using BERT for engagement detection and developing a Joint Emotion and Cognition Topic Model (JECTM) based on Bayesian networks, we analyzed 27,428 discussion posts from 2857 learners in a psychology MOOC. Our findings reveal three key insights: (1) The proposed two-step approach efficiently detects topics and analyzes their patterns of emotional and cognitive engagement. (2) As topic complexity increases, learners demonstrate higher-order cognitive engagement while experiencing reduced positive emotions along with increased confused and negative emotions. (3) In high-complexity topics, learners who maintain both positive emotions and higher-order cognitive engagement are more likely to achieve academic success than those who have negative emotions or lower-order cognitive engagement. These fine-grained analyses provide valuable insights for optimizing discussion design and interventions. This study also provides implications for the analysis of classroom dialogues and AI tutor-based conversations.
Keywords
Introduction
Massive Open Online Courses (MOOCs) have revolutionized education by offering unprecedented global learning opportunities (Xing & Du, 2023). The scale of this transformation is remarkable - as of July 2023, Udemy alone serves 662 million learners across 202,000 courses (Dhawal & Archisha, 2023). However, MOOCs continue to face significant challenges in sustaining learner engagement, primarily due to limited face-to-face interaction and potential feelings of isolation in online learning environments (Orji & Vassileva, 2023; Tang et al., 2019). To address these challenges, online discussion forums have emerged as a vital solution, providing learners with opportunities to connect, share ideas, and support each other in ways that effectively mirror traditional classroom interactions (Gašević et al., 2018).
Research has consistently demonstrated that well-designed forum discussions play a crucial role in enhancing both learner engagement and academic achievement (Liu et al., 2024; Ouyang, 2021; Ramirez-Arellano et al., 2019). The effectiveness of these discussions is largely determined by topic design (Chen et al., 2021), particularly topic complexity - defined as the level of cognitive effort required for meaningful participation (Bloom, 1956; Nistor et al., 2020). This concept is firmly grounded in established frameworks such as Bloom’s taxonomy, which posits that analysis-level topics require more higher-order knowledge transfer compared to recall-level topics (Krathwohl, 2002). We randomly selected two discussion threads from our dataset to illustrate possible emotional and cognitive responses under different topic complexity levels. As illustrated in Figure 1, the discussion threads show distinct patterns of emotional and cognitive engagement. Thread (a) represents a low-complexity topic focused on understanding, which tends to elicit positive emotion and lower-order cognitive engagement. In contrast, thread (b) represents a high-complexity topic focused on creation, which tends to trigger confusion and higher-order cognitive engagement. These examples demonstrate hypothetical relationships between topic complexity and engagement patterns that require further empirical investigation to validate. Two hypothetical examples illustrating the impact of topic complexity on emotional and cognitive engagement in forum discussions.
Learning engagement in online education encompasses behavioral, emotional, and cognitive dimensions (D’Mello, Dieterle, & Duckworth, 2017; Sun & Rueda, 2012). While researchers have extensively examined how topic complexity affects behavioral engagement (Chen et al., 2020; Lee & Recker, 2021; Peng & Xu, 2020), we know less about its simultaneous impact on emotional and cognitive engagement. This knowledge gap warrants attention for several key reasons.
First, emotional and cognitive engagement provide deeper insights into learning processes compared to surface-level behavioral engagement like participation rates and time spent on tasks (Liu et al., 2024). Specifically, cognitive engagement directly reflects learners’ mental investment by revealing how they analyze, synthesize, and construct knowledge (Chi & Wylie, 2014). Emotional engagement reflects learners’ motivation, persistence, and willingness to engage with course content (Dubovi, 2022). By understanding how topic complexity influences these engagement patterns, we can uncover crucial mechanisms for optimizing learning outcomes in MOOCs.
Second, topic complexity influences emotional and cognitive engagement in mixed ways. Complex topics may promote higher-order thinking while inducing negative emotions such as confusion or frustration (Wei et al., 2021). In contrast, simpler topics may elicit positive emotional responses but provide insufficient cognitive challenge. Understanding this delicate balance is essential for maintaining both positive emotional states and deep cognitive processing.
Third, we lack a comprehensive understanding of how topic complexity, emotional engagement, and cognitive engagement collectively impact academic achievement (Dillon et al., 2016; Wei et al., 2021; Yang et al., 2022). This gap offers an opportunity to develop an integrated model explaining how these factors work together to promote learning success in MOOCs.
In addition, analyzing the discussion topics and their influence on emotional and cognitive engagement in large-scale discussions poses substantial methodological challenges. For example, current methods for analyzing discussions typically rely on isolated engagement classification or topic detection. These isolated approaches have proven insufficient to accurately capture patterns of engagement and to model the complex interactions between these multifaceted factors (Dalipi et al., 2021). This critical gap requires innovative methodological solutions capable of conducting analysis of MOOC forum discussions at scale. To address these theoretical and methodological challenges, this study aims to: 1) Develop a novel two-step approach involving a deep neural network and a Bayesian probabilistic topic model to identify latent triggering relationships between topic complexity, emotional engagement, and cognitive engagement. 2) Investigate the interplay relationships between topic complexity, emotional engagement, and cognitive engagement and how they jointly predict academic achievement. These findings will provide valuable insights for optimizing the design and intervention strategies in MOOC forum discussions.
Literature Review
Topic Complexity in Discussions
Vygotsky’s social constructivist theory posits that knowledge is shaped, co-constructed, and learned within a knowledge community through language (Vygotsky, 1978). Lave & Wenger (1991) define learning communities as “the development of a shared identity around a set of challenging topics,” emphasizing the importance of discussion topics in fostering meaningful interactions. The complexity level of a discussion topic refers to the degree of effort and cognitive processing required in a collaborative discussion activity. High-complexity topics demand more cooperation, information extraction, and information synthesis from learners (Malinen, 2015).
Bloom’s taxonomy of educational objectives provides a widely used framework for analyzing topic complexity (Bloom, 1956; Krathwohl, 2002; O’Riordan et al., 2021). Building on this foundation, Ertmer et al. (2011) developed cognitive key verbs aligned with specific cognitive objectives to classify discussion topics into low-complexity (knowledge, comprehension, and application) and high-complexity topics (analysis, synthesis, and evaluation). Recent studies by Mohammed and Omar (2020) and Shaikh et al. (2021) have combined Bloom’s taxonomy with machine learning techniques to further refine topic complexity detection. They found that assessing topic complexity requires considering both cognitive key verbs and the difficulty of knowledge terms within the topic. These theoretical frameworks have been increasingly applied in computational approaches to analyze the topic complexity of discussions. For example, Iqbal et al. (2023) used Bloom’s taxonomy to categorize topics mined by Latent Dirichlet Allocation (LDA). Rivera et al. (2024) proposed a four-level classification scheme of discussion topics: asking questions, expressing opinions, analyzing content, and creating materials. This classification scheme offers a more nuanced approach to understanding the cognitive demands of different discussion topics. Chen et al. (2020) focused on evaluation and knowledge content in their topic complexity classification, highlighting the topics of distinguishing between factual recall and critical analysis in online discussions.
While multiple factors, including instructor presence, platform design, and course structure, contribute to learning engagement in MOOC discussions (Liu et al., 2024; Rivera et al., 2024), topic complexity emerges as a key focus for both practical and theoretical reasons. From a practical perspective, topic complexity serves three essential functions: first, it acts as a fundamental scaffold that structures and guides meaningful discussion (Chen et al., 2020); second, it represents a uniquely controllable and measurable variable that instructors can manipulate; third, it exhibits remarkable consistency across diverse discussion environments in online learning platforms. Furthermore, topic complexity has strong theoretical foundations and provides robust evidence of its pivotal role in shaping both cognitive processing and emotional states (Wei et al., 2021). These reasons make topic complexity an ideal factor for investigating the research objectives of this study.
Emotional Engagement and Cognitive Engagement in Discussions
Learning engagement in forum discussions encompasses behavioral, emotional, and cognitive dimensions that predict academic achievement (Sun & Rueda, 2012). Behavioral engagement refers to the duration and frequency of learning behaviors, and the role of topic complexity on behavioral engagement has been widely studied (Chen et al., 2020; Lee & Recker, 2021; Peng & Xu, 2020). Cognitive engagement reflects deeper levels of mental processing such as analyzing, synthesizing, and constructing knowledge (Chi & Wylie, 2014). Emotional engagement refers to learners’ affective responses during learning (Demir et al., 2023; Sun & Rueda, 2012). Topic complexity can lead to cognitive imbalance and affect learners’ sense of control over the learning process, resulting in a wide range of emotional responses (Buhr et al., 2019).
This study examines emotional engagement through the framework of epistemic emotions, which are intrinsically connected to cognitive processes and extensively studied in educational research (Han et al., 2021; Liu et al., 2022). Epistemic emotions function as key facilitators of cognitive processes, encompassing positive emotions (e.g., enjoyment), confusion (e.g., doubt), and negative emotions (e.g., frustration) (Chevrier et al., 2019). The phenomenon of confusion in learning represents a particularly nuanced case, as it occupies an intermediate position between positive and negative emotional states (D’Mello et al., 2014). When learners successfully resolve their confusion, it often leads to positive emotions, whereas unresolved confusion tends to generate negative emotions. Research indicates that confusion is a common experience during learning, with studies documenting its prevalence in 15%–50% of learning processes (Dillon et al., 2016). Based on these theoretical and empirical foundations, and following recent MOOC forum studies (Liu et al., 2022; Liu et al., 2024), the present study analyzes emotional engagement in discussions across three primary dimensions: positive, confused, and negative.
Cognitive engagement refers to the depth of mental processing that occurs when learners interact with learning materials. This study employs the Interactive, Constructive, Active, and Passive (ICAP) framework to measure cognitive engagement (Chi & Wylie, 2014). Numerous studies have substantiated the effectiveness of this framework (Atapattu et al., 2019; Liu et al., 2022; Wang et al., 2016). The ICAP framework posits that cognitive engagement and its contribution to academic achievement progressively increase from passive to active, constructive, and interactive cognition. Passive engagement involves merely acquiring information, which is not typically observable in discussion processes. Active engagement entails extracting information from existing knowledge. Constructive engagement indicates the integration of new knowledge with prior understanding. Interactive engagement represents a co-inferring process where learners collaboratively generate knowledge. Liu et al. (2022) further refined the coding scheme developed by Wang et al. (2015) based on the ICAP framework and applied it to analyze MOOC discussions, providing a robust methodology for assessing cognitive engagement in online learning environments.
Recent research suggests that topic complexity simultaneously influences learners’ cognitive and emotional engagement (Geng et al., 2020). According to cognitive load theory (Merriënboer & Sweller, 2005), topic complexity directly affects intrinsic cognitive load - low-complexity topics may induce boredom and lower-order cognitive engagement due to insufficient cognitive challenge, while highly complex topics may lead to frustration and excessive cognitive load that overwhelms working memory capacity, potentially limiting higher-order cognitive engagement (Ding, 2019; O’Riordan et al., 2021). However, the interplay between topic complexity, emotional and cognitive engagement, and their subsequent impact on academic achievement in online discussion remains to be fully explored.
Automated Analytics of Topic, Emotion, and Cognition in Discussions
The manual annotation of massive discussion data is both costly and impractical. The development of natural language processing (NLP) provides promising solutions for automated analysis of forum discussions. Recently, Bidirectional Encoder Representations from Transformers (BERT) have dramatically improved the performance of automated detection of emotion and cognition (Han et al., 2021; Liu et al., 2022; Liu et al., 2024). These findings highlight the robustness and efficiency of BERT in analyzing complex linguistic patterns in educational contexts. For example, Liu et al. (2022) employed BERT to detect emotional and cognitive engagement, observing enhancements in F1 scores by 18% and 8%, respectively, in comparison to text classification models based on machine learning, such as random forest (RF) and support vector machine (SVM). Although large language models (LLMs) such as GPT-4 and LLaMA-3 perform well with zero or few data samples, several evaluation studies have shown that BERT with sufficient data samples (>1000 samples) significantly outperforms LLMs in multi-class text classification tasks of emotion and cognition (Hou et al., 2024; Liu et al., 2024; Wang et al., 2023).
In addition, analyzing discussion using only emotional or cognitive indicators without detecting their co-occurring topics may hinder a refined understanding of the discussion process (Pan et al., 2022). LDA is an unsupervised machine learning algorithm based on the Bayesian network model for topic modeling (Blei et al., 2003), allowing automatic extraction of latent topics from discussion texts and coding their complexity. For example, Gašević et al. (2018) used LDA to identify topics and code them into critical thinking, social cohesion, and course logistics. Onan and Toçoğlu (2020) employed LDA to cluster discussion questions into three topics: technique, course content, and course logistics. These applications demonstrate the versatility of LDA in uncovering meaningful patterns in educational discourse.
Moreover, as a Bayesian network, LDA has good interpretability and a complete mathematical foundation. Thus, fusing LDA with feature variables (e.g., emotion or cognition) can effectively model complex discussion structures, providing a comprehensive framework for analyzing the multifaceted nature of online learning interactions. For example, Qiu et al. (2013) introduced a behavioral topic model (B-LDA) that jointly models discussion topic content and behavior patterns. Peng et al. (2020) developed an emotion and behavior topic model with an emotion lexicon to calculate emotional and behavioral patterns of topics, finding that course content topics elicited positive emotions, whereas course logistics topics triggered negative emotions. These studies highlight the potential of integrating emotional and behavioral factors into topic modeling for a more holistic understanding of online discussions.
To achieve our research objectives, after careful consideration of alternative methodologies, we determined that a single-step approach would be insufficient for two main reasons. First, while deep learning-based text classification algorithms excel at detecting emotions and cognitions, they lack the ability to effectively uncover latent topic structures within the text. Second, traditional topic modeling approaches are limited in modeling the intricate relationships between topics, emotions, and cognitions, resulting in reduced accuracy for both engagement detection and topic discovery (Birjali et al., 2021).
To address these limitations, we developed an integrated two-step approach that synergistically combines deep learning text classification with advanced topic modeling. This hybrid methodology effectively fulfills our research requirements through two steps: Step 1 leverages BERT’s powerful capabilities to achieve high-precision classification of emotional and cognitive engagement, while Step 2 uses a model to explore the complex relationships between topics, emotions, and cognition. Together, these components enable the discovery of meaningful and interpretable patterns in how engagement varies across topics.
Research Questions
To achieve the research objectives, we propose a two-step approach to investigate the following research questions (RQs): RQ1: What are the discussion topics and their levels of complexity in the MOOC forum? RQ2: How does topic complexity influence emotional and cognitive engagement in MOOC discussions? RQ2a: How does the topic complexity influence emotional engagement? RQ2b: How does the topic complexity influence cognitive engagement? RQ2c: How does the topic complexity jointly influence emotional and cognitive engagement? RQ3: How do topic complexity, emotional engagement, and cognitive engagement predict academic achievement?
Methods
Research Design
This study employs an evidence-centered learning analytic framework (see Figure 2) that systematically maps observable textual data to meaningful educational psychology indicators (Gašević et al., 2017; Wise et al., 2021). The research design comprises three distinct phases with six main steps: Research design. Note. Joint Emotion and cognition topic Model (JECTM).
Instructional Design and Achievement Measurement
In steps 1 and 2 of Figure 2, experimental data was collected from a national-level quality course titled Modern Psychology. The study included 2857 Chinese undergraduate learners majoring in non-psychology fields (e.g., management and computer science), all with similar levels of psychological knowledge and undisclosed gender information. A team of psychological instructors from a university designed and conducted the course from May to September 2019, hosting it on China’s largest MOOC platform (https://www.icourse163.org). This nine-week course comprised nine units featuring lecture videos, course materials, and forum discussions. Instructors and learners contributed 523 root posts with varying topic complexity, with learners encouraged to express their viewpoints on each post.
The selection of Modern Psychology, a general education introductory course, was motivated by two considerations. First, the course encompasses diverse and balanced discussion topics of varying complexity levels, integrating conceptual knowledge comprehension, real-world case analysis, and innovative solution development. This range of topics provides an optimal environment for comprehensively examining the research questions. Second, psychology’s inherently interdisciplinary nature encompasses multifaceted content, which substantially enhances the generalizability of our research findings across multiple academic domains, such as statistics, neuroscience, and social psychology.
The achievement assessment, totaling 100 points, consisted of two components: lecture video viewing duration and quiz performance (50 points) and a final exam (50 points). Passing certificates were awarded to 1153 learners (40.36%) who scored 60 or above. Conversely, 1704 learners (59.64%) did not receive a certificate, as their scores were below 60. We consider the learners who received a passing certificate as the high-achievement group and those who did not receive a certificate as the low-achievement group. Many studies have adopted this approach to measuring academic achievement and dividing learners into high and low achievement groups (Doo & Kim, 2024; Liu et al., 2022; Liu et al., 2024). The certification process for this course is free, resulting in a higher certification pass rate than courses offered on paid platforms such as Coursera and EdX, effectively reducing the impact of financial factors on measures of academic achievement. In addition, forum participation was entirely voluntary with no mandatory requirements or course credits, ensuring that participants’ contributions stemmed from intrinsic interest rather than external motivations.
To optimize our topic modeling quality, we implemented a content word retention strategy that effectively reduced textual noise. This strategy has been extensively validated through multiple studies (Peng et al., 2020; Yao et al., 2017). We utilized HanLP’s part-of-speech tagging functionality to identify essential content words (including nouns, verbs, and named entities) while filtering out low-frequency words (occurring fewer than 5 times) (He, 2021). Subsequently, posts containing less than 3 words were flagged as spam because they could not be used for topic modeling. Our analysis encompassed 27,502 learner discussion posts, of which 27,428 were determined to be valid contributions, with only 74 posts classified as invalid. Notably, the average number of posts by the high-achievement group (mean = 11.15, SD = 8.24) significantly surpassed that of the low-achievement group (mean = 3.30, SD = 3.87).
Detecting Emotional and Cognitive Engagement Using BERT
The Coding Scheme of Emotional Engagement and Cognitive Engagement (S. Liu et al., 2022).
Sample Distribution of Labeled Datasets of Emotional and Cognitive Engagement.
For automatic engagement detection, we trained and tested the BERT-based classification models using 10-fold cross-validation. BERT has demonstrated superior performance over traditional approaches like random forests, CNNs, and LSTMs in analyzing educational forum discussions (Devlin et al., 2019; Kong et al., 2025; Liu et al., 2022). We utilized the bert-base-Chinese model architecture (12 layers, 768 hidden dimensions, 12 attention heads) and fine-tuned it with optimized hyperparameters: 2 × 10−5 learning rate with linear decay, 32 batch size, 512 maximum sequence length, and 50 training epochs. The model training was conducted on a single NVIDIA V100 GPU utilizing mixed precision computation.
The classification performance of BERT demonstrated exceptional performance across both engagement dimensions (macro-F1 > 0.70) (Kong et al., 2025; Zou et al., 2021). For emotional engagement detection, the model achieved strong metrics with 0.837 accuracy, 0.793 macro-precision, 0.776 macro-recall, and 0.778 macro-F1 score. Similarly, robust performance was observed for cognitive engagement classification, yielding 0.754 accuracy, 0.758 macro-precision, 0.754 macro-recall, and 0.748 macro-F1 score.
Modeling Latent Relationships Between Topic, Emotion, and Cognition by JECTM
In step 4 of Figure 2, a Bayesian network model called JECTM was developed to extract latent topics that trigger emotional and cognitive engagement from large-scale discussion texts. As shown in Figure 3, the model incorporated emotional and cognitive features detected from BERT as heuristic information to guide the topic-generation process. The Overall structure of JECTM.
Figure 3 illustrates JECTM’s three-layer structure: learners (L), documents/posts (D), and words (W), with arrows indicating triggering relationships between variables. The model’s generation process follows these assumptions: (1) When a learner writes a post, a topic
JECTM’s design was guided by three key principles. First, we aimed to effectively capture the complex interplay between topics, emotions, and cognition in discussions - a capability traditional LDA models lack. Second, we incorporated Bayesian network principles to ensure result interpretability, crucial for educational stakeholders who need to understand the model’s reasoning. Third, we designed JECTM to handle MOOC discussion data’s characteristics, including its hierarchical structure (learner-post-word) and the co-occurrence of content and engagement features. These innovations enable JECTM to provide richer insights into discussions compared to conventional approaches.
To evaluate JECTM against baseline topic models, we employed normalized pointwise mutual information (NPMI) as our primary metric for topic coherence. NPMI quantitatively assesses how interpretable the topic-word distributions are from a human perspective (Hoyle et al., 2021). This metric has become the standard in recent literature, proving more effective than traditional measures like perplexity and entropy for evaluating topic model quality (Lau et al., 2014). We computed mean topic coherence across the top 10 words for each identified topic (Newman et al., 2010). We selected B-LDA (Qiu et al., 2013) and LDA (Blei et al., 2003) as baselines due to their widespread use in MOOC research (Gašević et al., 2018; Liu et al., 2019; Onan & Toçoğlu, 2020). B-LDA was specifically chosen for its similar structure in modeling behavioral patterns (such as response and like), while standard LDA served as a fundamental baseline that improved models should exceed.
Figure 4 displays the average NPMI scores of all models on our dataset, clearly comparing their performance. The results demonstrate that JCETM consistently outperformed both B-LDA and LDA across various topics. This superior performance can be attributed to JCETM’s incorporation of engagement information and ability to capture more nuanced relationships between words and topics. Furthermore, through a series of experiments, we determined that the optimal number of topics for JECTM on the dataset was 80. Topic Coherence Results of the Models. Note. JECTM is our proposed model.
The selection of 80 topics as the optimal number was based on comprehensive considerations across multiple dimensions. First, from a computational complexity perspective, this choice represents an optimal balance between topic granularity and processing efficiency. Experiments showed that when the number of topics exceeded 80, computational costs increased exponentially while model coherence improvements plateaued. Second, from a reliability perspective, the 80-topic model demonstrated significant stability with NPMI metrics reaching optimal levels across different topic models, indicating that this topic count effectively captures the fundamental discussion structure without overfitting. Third, from an interpretability perspective, 80 topics maintained both topic distinctiveness and reasonable costs for manual interpretation and coding. In conclusion, 80 topics not only achieved optimality in technical computational complexity metrics but also demonstrated good usability and interpretability in practical applications.
Coding Scheme of Topic Complexity
As illustrated in step 5 of Figure 2, we implemented a two-stage process for coding topic complexity. In the first stage, we classified topics into content-related and non-content-related topics following Wise et al. (2017). Content-related topics encompassed knowledge concepts directly from learning materials, while non-content-related topics covered course logistics like exam schedules. Two trained annotators independently classified the 80 JECTM-extracted topics using this scheme, achieving strong inter-rater reliability (Cohen’s kappa = 0.92, exceeding the 0.70 threshold) (Cohen, 1960). Any discrepancies were resolved by discussion until a consensus was reached between the two annotators.
Topic Complexity Coding Scheme based on Bloom’s Taxonomy of Educational Objectives (Bloom, 1956; Krathwohl, 2002; O’Riordan et al., 2021).
Data Analysis
As depicted in step 6 of Figure 2, we employed a comprehensive analytical approach to address our research questions. For RQ1, we utilized BERT and JECTM to extract discussion topics, while topic-word distributions informed our coding of topic complexity levels. For RQ2, we conducted cross-tabulation analyses and chi-square tests to examine the associations between topic complexity and different forms of engagement, with standardized residuals providing deeper insights into these relationships. For RQ3, we leveraged Epistemic Network Analysis (ENA) to visualize and quantify the co-occurrence patterns of topic complexity and engagement types across high- and low-achieving groups (Shaffer et al., 2016). We complemented logistic regression analysis to investigate how topic complexity and engagement patterns jointly predict academic achievement (Soffer & Cohen, 2019).
Results
RQ1: what are the Discussion Topics and Their Levels of Complexity in the MOOC Forum?
Complexity Levels and Their Representative Topics.
Note.
LCT/MCT/HCT = Low/Middle/High-Complexity Topic.
Two key findings of our dataset support the robustness of our subsequent analyses. First, the predominance of content-related topics (n = 69, 87.4%) indicates that discussions were primarily focused on course material rather than irrelevant material, providing a solid foundation for analyzing content-related engagement patterns. Second, the balanced distribution across complexity levels (LCT = 27.0%, MCT = 30.9%, HCT = 29.5%) ensures sufficient data in each category for meaningful comparative analysis of how different complexity levels influence engagement.
The content-related topics exhibited distinct characteristics across complexity levels. Low-Complexity Topics (LCT, n = 17, 27.0%) centered on fundamental cognitive processes of remembering and understanding knowledge concepts. Topic-2 exemplifies this category, where learners explained psychological theories and received responses elaborating on various theoretical frameworks. Middle-Complexity Topics (MCT, n = 30, 30.9%) emphasized concept application and analysis, as demonstrated by Topic-3 where learners applied psychological principles to enhance social skills. High-Complexity Topics (HCT, n = 22, 29.5%) required sophisticated cognitive engagement through evaluation and creation. Topic-4 represents this highest complexity level, showing how learners critically evaluated the influence of original family dynamics on child development through multiple psychological lenses.
RQ2: How Does Topic Complexity Influence Emotional and Cognitive Engagement in MOOC Discussions?
RQ2a: How Does Topic Complexity Influence Emotional Engagement in MOOC Discussions?
Statistical analysis using the Chi-square test revealed a significant association between topic complexity and emotional engagement ( Analysis of Topic Complexity Triggering Emotional Engagement: Conditional Probability and Chi-square Test Results. Note. LCT/MCT/HCT = Low/Middle/High-Complexity Topic. Standardized residual |z|* > 1.96.
RQ2b: How Does Topic Complexity Influence Cognitive Engagement in MOOC Discussions?
Statistical analysis revealed a significant association between topic complexity and cognitive engagement ( Analysis of Topic Complexity Triggering Cognitive Engagement: Conditional Probability and Chi-square Test Results. Note. LCT/MCT/HCT = Low/Middle/High-Complexity Topic. Standardized residual |z|
*
> 1.96.
RQ2c: How Does the Topic Complexity Jointly Influence Emotional and Cognitive Engagement?
Analysis of Topic Complexity Triggering Emotional Engagement and Cognitive Engagement: Conditional Probability and Chi-Square Test Results.
Note. LCT/MCT/HCT = Low/Middle/High-Complexity Topic. Standardized residual |z| * > 1.96.
RQ3: How Do Topic Complexity, Emotional Engagement, and Cognitive Engagement Predict Academic Achievement in MOOC Forum Discussions?
Figure 7 illustrates the distinct patterns of interaction between topic complexity, emotional engagement, and cognitive engagement across high- and low-achievement learner groups. The mean epistemic networks for these groups are depicted in Figure 7(a) and (b) respectively. A subtraction network analysis (high minus low) presented in Figure 7(c) reveals compelling differences between the high-achievement group (depicted in blue, mean = −0.15, SD = 0.21) and low-achievement group (depicted in red, mean = 0.10, SD = 0.23) along the x-axis (α = 0.05, t (2577.57) = −29.56, p < .001, Cohen’s d = 1.11). The analysis reveals that the high-achievement group demonstrates stronger connections to HCT, maintains positive emotion, and engages in both constructive and interactive cognition. In contrast, the low-achievement group exhibits stronger associations with both HCT and MCT, displays more confused and negative emotion, and primarily engages in constructive cognition. Notably, when engaging with high-complexity topics, the high-achievement group demonstrates significantly higher levels of positive emotion and interactive cognition while experiencing markedly less confused and negative emotion compared to low-achievement learners. Epistemic Network Analysis of Topic Complexity, Emotional Engagement, and Cognitive Engagement Between High- and Low-Achievement Groups. Note. The conversation window parameter is set to 1. LCT/MCT/HCT represent Low/Middle/High-Complexity Topics, respectively.
Logistic Regression Results of Topic Complexity, Emotional Engagement, and Cognitive Engagement on Academic Achievement.
Note. LCT/MCT/HCT = Low/Middle/High complexity topic. Nagelkerke R 2 . Non-significant predictor variables were excluded by the stepwise method.
The analysis revealed several significant findings across four progressive models in Table 6.
Discussions
This section discusses the key contributions and insights of our findings.
Methodological Advances in Analyzing MOOC Discussions at Scale
Our study makes several significant methodological contributions to the analysis of MOOC discussions at scale. First, we developed an innovative integrated two-step approach that combines automated topic modeling with engagement analysis, advancing beyond traditional methods that rely primarily on manual coding or basic text analysis techniques (Chen et al., 2020; Liu et al., 2022). This novel approach enables a rapid and comprehensive understanding of discussion topic trends, key question words, response content, and associated emotional and cognitive engagement patterns. The methodology provides a robust and efficient analytical framework for processing large-scale MOOC discussions.
Second, our research transcends the conventional single-dimensional focus on either cognitive or emotional engagement by integrating topic complexity analysis with both emotional and cognitive engagement measures (Liu et al., 2019, 2020). This multi-dimensional analytical approach reveals intricate patterns and relationships that would remain hidden in single-dimension analyses, offering deeper insights into the complex dynamics of online discussions.
Third, our study demonstrates the successful integration of advanced machine learning techniques with established educational theory in analyzing MOOC discussions. While previous researchers have employed algorithmic approaches to analyze online discussions (Onan & Toçoğlu, 2020; Wise & Cui, 2018), few have explicitly connected these technical methodologies to established theoretical frameworks of learning engagement. This synthesis of theoretical understanding and technological innovation enables a more nuanced and meaningful interpretation of discussion patterns and their implications.
Understanding the Interplay of Topic Complexity, Emotion, and Cognition
Our findings reveal several important relationships between topic complexity and learning engagement in MOOC discussions. While prior studies have acknowledged topic complexity’s role in influencing behavioral engagement (Lee & Recker, 2021), our research extends our understanding of its fine-grained effects on emotional and cognitive engagement. The results demonstrated that low-complexity topics generally elicited positive emotions but resulted in lower-order cognitive engagement. This may be because low-complexity topics are easier to understand and engage with, enhancing learners’ self-efficacy and positive emotions (Wigfield & Eccles, 2020).
In contrast, high-complexity topics promoted deeper cognitive processing while simultaneously increasing confusion and negative emotion. This aligns with previous research suggesting that high-complexity tasks place substantial demands on working memory resources, potentially leading to cognitive overload when the intrinsic cognitive load exceeds processing capacity (Merriënboer & Sweller, 2005). Insufficient cognitive resources may reduce learners’ sense of competence, generating negative emotions like confusion and anxiety (Pekrun & Linnenbrink-Garcia, 2012). Graesser (2020) further posit that confusion arises when learners encounter information conflicting with existing knowledge structures, necessitating deeper cognitive processing to resolve these cognitive conflicts. Therefore, this productive struggle may be necessary for developing higher-order thinking skills when appropriately scaffolded (Kapur, 2011). The key challenge lies in balancing cognitive challenge with emotional support to promote both deep thinking and positive emotional engagement.
The Role of Topic Complexity, Emotion, and Cognition on Academic Achievement
Our analysis reveals distinct engagement patterns between achievement groups. The discussion process of high-achieving learners demonstrated stronger associations between high-complexity topics, positive emotion, and both constructive and interactive cognitive engagement. The observed patterns align with past research suggesting that high-achieving learners typically demonstrate superior cognitive strategies and emotional regulation strategies, particularly in challenging learning tasks (Zimmerman & Schunk, 2011). These learners may be more adept at using peer interactions to construct knowledge, particularly when dealing with challenging content, and thus achieve higher academic achievement.
Conversely, the discussion process of low-achieving learners exhibited stronger correlations between middle-complexity topics, confusion, negative emotion, and primarily constructive cognition. A possible reason for this is that when encountering middle-complexity topics, these learners likely experience greater cognitive load. This increased cognitive burden leads them to express more confusion while promoting constructive cognition as they attempt to address the challenges (Graesser, 2020). The frustration experienced during the learning process often results in negative emotions (Merriënboer & Sweller, 2005). These findings suggest that emotion regulation plays an essential role in sustaining higher-order cognitive engagement (Dubovi, 2022). This underscores the importance of carefully designed discussion topics that provide sufficient cognitive challenge while fostering positive emotions to maintain sustained engagement.
Conclusions, Implications, and Future Works
In this study, a two-step approach was adopted to investigate the effect of topic complexity on emotional engagement, cognitive engagement, and academic achievement in MOOC discussions. This approach provides stakeholders with a quick and comprehensive overview of discussion topics and their emotional and cognitive engagement patterns. The evaluation showed that the approach outperformed known algorithms, which empowered efficient analysis of large-scale discussions.
This research provides practical implications for both instructors and learners. For instructors, our proposed two-step approach facilitates efficient monitoring of large-scale discussions, enabling instructors to track discussion topics and engagement patterns in real-time. We recommend three key implementation strategies: First, design topics progressively - begin with low-complexity topics to build confidence, then gradually introduce more complex topics while maintaining balanced emotional and cognitive engagement levels. Second, for challenging topics, provide targeted guidance, monitor emotional responses, and actively encourage peer interaction for collaborative problem-solving. Third, implement timely interventions by incorporating targeted guiding questions and prompts to stimulate meaningful discussion when topics become off-topic.
For learners, we propose several effective strategies to enhance their learning experience. First, our two-step approach serves as a learning analytics and self-regulation tool that generates discussion visualization reports (Chen et al., 2021; Ouyang et al., 2021), helping learners better understand their performance patterns and select appropriate discussion topics. Second, learners can track their engagement across topics of varying complexity, actively work to increase their cognitive engagement, and learn to regulate their emotions to maintain motivation. They can also enhance their learning strategies by observing and emulating the visualization reports of high-performing peers in complex discussions.
While our study makes significant contributions, several limitations warrant consideration. First, our analysis was confined to a single MOOC in psychology, future research should examine these relationships across diverse subject areas and learning contexts. Second, subsequent studies could investigate additional influential factors such as learner characteristics, prior knowledge levels, and motivational aspects that may moderate the relationships we observed. Third, our innovative two-step approach shows promise for application beyond MOOCs to other educational contexts, including classroom dialogues and agent-based conversations (e.g., AI tutors and chatbots). Fourth, while we have adopted achievement measurement methods commonly used in existing studies (Doo & Kim, 2024; Liu et al., 2022; Liu et al., 2024), future research could focus on implementing discussion-specific measures that can more accurately capture the unique characteristics of discussions. Fifth, the current three-dimensional emotional engagement framework could be expanded to include a broader range of emotional categories for more comprehensive measurement.
Footnotes
Author Contributions
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was financially supported by the National Science and Technology Major Project (2022ZD0117101), National Natural Science Foundation of China (62377016, 62293554, 62293550, 62407023), Hubei Provincial Natural Science Foundation of China (2023AFA020), the 2024 “AI + Examination and Evaluation” Special Teaching Reform Project of Central China Normal University (CCNU24JG10), and the Major Project of Philosophy and Social Science Research in Colleges and Universities of Jiangsu Province (2024SJZD074).
Data Availability Statement
Data is available on request from the authors.
Appendix
Equation (A1) denoted the joint probability equation of JECTM.
Equation (A2) presented the derivation process of the Gibbs sampling formula.
We employed Gibbs sampling for JECTM inference, with the detailed process outlined in Algorithm A1. Before analysis, we preprocessed the Chinese discussion posts through several steps: First, following LDA modeling assumptions, we combined root posts with their corresponding response posts to optimize topic-word distribution coherence and interpretability. Second, we utilized HanLP for word tokenization. Third, we created a domain-specific dictionary to ensure accurate word segmentation, incorporating key terms like “内在动机/intrinsic motivation” and “发展心理学/developmental psychology”. Fourth, we cleaned the text by removing stop words (e.g., “和/and,” “的/of”), low-frequency words (occurring ≤ 5 times), and noise elements (URLs and punctuation).
Algorithm A1: Gibbs sampling for JECTM
Pre-processed learner discussion posts.
The iteration times of Gibbs sampling.
topics number: T
hyper-parameter:
Learner-topics distribution
, topic-words distribution
Topic-cognitions distribution
, topic-emotions distribution
1: Begin procedure
3:
4:
5: Randomly assign a topic to
10:
11:
12: Gibbs sample topic
according to equation (A2)
13:
14:
16:
17: Estimate model parameters
,
,
,
, according to equations (A3–A6)
18: End procedure
