Abstract
ChatGPT has shown considerable potential for Automated Item Generation, but the quality of ChatGPT-generated items in language assessment remains insufficiently substantiated. This research recruited 121 participants to systematically compare the psychometric properties of the test items and the linguistic features of the reading passages in ChatGPT-generated and official CET-4 reading comprehension materials, using Item Response Theory and Coh-Metrix. Key findings are as follows: (1) generated items fell short in higher-order reading skills; (2) the generated items were less difficult than official ones, showing weaker discrimination and providing measurement information mainly for lower-performing students; (3) only 22.9% distractors functioned effectively, indicating insufficient distractor performance; and (4) ChatGPT-generated passages were characterized by irregular lexical distribution, higher lexical complexity, weaker cohesion but simpler sentences than CET-4 passages. Although ChatGPT-generated passages were less readable than CET-4 passages, the corresponding items were easier and showed lower discrimination. This discrepancy can be attributed to inadequate distractor functioning that facilitates option elimination without complete passage comprehension, as well as to the underrepresentation of higher-order reading skills. The findings corroborate the conclusion that ChatGPT may function effectively as a supplementary tool in low-stakes assessment; however, substantial refinements in item quality are imperative before its application in high-stakes testing.
Keywords
Introduction
Item development in language assessment faces multiple challenges. First, the development of high-quality test items is time-consuming and labor-intensive (Lane et al., 2015); as a result, teachers often encounter a shortage of resources when attempting to address students’ diverse needs. Additionally, the control of item difficulty and discrimination remains insufficiently precise (Young et al., 2025). Such issues not only compromise students’ learning effectiveness but also increase instructors’ workload. Accordingly, the realization of large-scale, cost-efficient, and quality-assured Automated Item Generation (AIG) has become an imperative. Amid the rapid advancement of large language models (LLMs), models such as ChatGPT have emerged as a promising solution to this imperative, as they demonstrate advanced capacities in natural language understanding and generation (OpenAI, 2023).
The College English Test Band 4 (CET-4) is a nationally recognized and widely adopted assessment in China for evaluating students’ English language proficiency (Zheng & Cheng, 2008). Its Reading Comprehension (RC) section is important, as it constitutes a substantial portion of the exam and has far-reaching implications for students’ academic development and social competitiveness. Prior research suggested that AIG can assist teachers in designing quizzes more efficiently, alleviating their workload (Benesch & Prior, 2023; Chinkina et al., 2019). Therefore, exploring the potential of ChatGPT for AIG in CET-4 RC tasks carries significant implications for language assessment.
Nevertheless, the extent to which ChatGPT-generated Multiple-Choice (MC) RC items can attain quality levels comparable to those of official items remains unresolved. Given that RC tasks consist of both passages and items, and both are essential in determining validity (Freedle & Kostin, 1999), we employed a comparative analysis of ChatGPT-generated and the CET-4 materials across two domains: (1) item quality and (2) the linguistic characteristics of the associated passages.
Literature Review
Theoretical Foundations of ChatGPT-Based AIG
ChatGPT is built on the Transformer architecture (Vaswani et al., 2017). Its training generally involves two stages: large-scale autoregressive next-token prediction during pre-training (Brown et al., 2020; OpenAI, 2023), followed by alignment through Reinforcement Learning from Human Feedback (RLHF) to ensure that model outputs better reflect human preferences (Ouyang et al., 2022). While this training enables ChatGPT to generate fluent and grammatically correct text, it exhibits several limitations that could affect the validity of AIG. (1) ChatGPT induces a strong dependence on high-frequency patterns in the training corpus, making the model more adept at reproducing surface-level linguistic regularities than at constructing the deeper cognitive structures required for assessment tasks (OpenAI, 2023). (2) Because negation appears relatively infrequently in natural language and often involves complex logical relations, ChatGPT is unstable in generating and interpreting negative constructions, resulting in their underuse in model-generated text (Kassner & Schütze, 2020). (3) High-quality distractors must be grounded in typical learner misconceptions (Gierl et al., 2017), yet RLHF explicitly optimizes for producing “correct and helpful” responses (Ouyang et al., 2022). Moreover, because ChatGPT primarily learn the statistical structure of fluent and correct language during pre-training (Brown et al., 2020; OpenAI, 2023), they struggle to generate distractors that exhibit the characteristic error features required for psychometric functioning. (4) Transformer-based generation relies heavily on high-probability syntactic templates combined with lexical substitution (Mizumoto et al., 2024), producing outputs that are lexically diverse but syntactically repetitive and lacking discourse cohesion, which increases cognitive load for readers.
In contrast, CET-4 items are developed by experts who follow a standardized process to ensure optimal balance among skill coverage, difficulty, discrimination, and distractor functioning (Bachman & Palmer, 2010; Lane et al., 2015; Zheng & Cheng, 2008). Although prompt engineering can narrow the gap between ChatGPT-generated and official items, it cannot replicate experts’ deep understanding of learner cognition, error patterns, and skill development trajectories (Lee et al., 2024).
Based on the foundations, we proposed the following hypotheses:
ChatGPT-generated items will overrepresent lower-order reading skills.
The overall reliability of ChatGPT-generated items is expected to be lower than that of CET-4 items.
ChatGPT-generated items will exhibit different IRT model parameters compared with CET-4 items.
Distractors generated by ChatGPT will exhibit weaker functionality compared with CET-4 distractors.
ChatGPT-generated passages will exhibit lower readability and a higher cognitive load.
ChatGPT-Generated Item Quality Evaluation
Evaluation of Measured Reading Skills
Balanced skill coverage is a fundamental concern in test development, as it ensures construct representation and supports test validity (Bachman & Palmer, 2010; Lane et al., 2015). In reading assessment, disproportionate emphasis on particular skills may bias score interpretation and weaken construct validity (Alderson, 2000).
Comparison of Skill Coverage in the CET-4 and ChatGPT Items
Automatically generated items often focus on explicit information retrieval while underrepresenting higher-order cognitive skills (Gierl & Lai, 2013). Although ChatGPT is capable of generating acceptable reading materials, it remains limited in aligning test items with the target skills (Zhang et al., 2025). More specifically, although ChatGPT-generated items show broadly comparable coverage of reading skills, they overrepresent lower-order skills, which underscores the necessity of explicitly defining skill requirements of prompt when leveraging ChatGPT for AIG (Lin & Chen, 2024; Sihite et al., 2023; Wen & Chu, 2025). Moreover, when guided by clearly defined contexts and prompts, ChatGPT can generate items that align well with textbook content and curricular knowledge (Bhandari et al., 2024), but such evidence is mainly confined to mathematics.
Evaluation of IRT Parameters
Item Response Theory (IRT) has become a cornerstone of modern psychometrics, evolving from the single-parameter Rasch model to the two-parameter model (2-PL) (Baker & Kim, 2004; Embretson & Reise, 2000). The advantage of IRT lies in overcoming the sample dependence inherent in Classical Test Theory (CTT), as it emphasizes the independence between item parameters and examinee ability estimates (Wright & Stone, 1979).
Within IRT, the assumption of unidimensionality, which posits that a test measures a single latent trait, is fundamental to ensuring accurate model fit and parameter estimation (Embretson & Reise, 2000). When this assumption is violated, both parameter stability and validity may be compromised (Ackerman, 1992). To address this, Monte Carlo simulation-based tests of unidimensionality, theoretically grounded in Drasgow and Lissak (1983) and later formalized by Kubinger (2003), have become widely employed. Confirming dimensional structure is thus not only a statistical requirement but also a prerequisite for establishing reliability (Ziegler & Hagemann, 2015).
Most recently, IRT has been increasingly applied to evaluate the quality of ChatGPT-generated items. Bhandari et al. (2024, 2023) compared ChatGPT-generated algebra items with textbook items and found no significant differences in items’ difficulty and discrimination. Young et al. (2025) demonstrated that ChatGPT-generated psychology items were easier but exhibited strong discrimination. Mendoza and Zúñiga (2025) found that ChatGPT-generated matriculation items were slightly harder but comparable to human-authored ones in terms of discrimination and model fit. Isley et al. (2025) reported that ChatGPT-generated Advanced Placement items were relatively easier but exhibited comparable levels of discrimination and reliability to human-authored items. In language assessment, Zhang et al. (2025) found ChatGPT-generated items and passages potentially usable but linguistically simple, with analyses showing that the items were easier and had acceptable reliability. Lin and Chen (2024) found that ChatGPT-generated CET-4 RC items were broadly comparable to official ones in psychometric properties, with some even showing higher discrimination. Overall, ChatGPT-generated items are appropriate for routine or low-stakes assignments.
Evaluation of the Quality of MC Distractors
The quality of distractors is a critical factor influencing the validity of MC items, as they directly determine item discrimination and measurement precision (Gierl et al., 2017; Ludewig et al., 2023). Traditionally, distractor performance has been evaluated through a combined analysis of the selection rate (SR) and discrimination index. According to Haladyna and Downing’s (1989) criteria, a functional distractor should meet two rudimentary conditions: the SR of at least 5% and the negative discrimination index, indicating greater attractiveness to lower-performing rather than higher-performing examinees. Tarrant et al. (2009) further refined this framework, showing that the joint consideration of SR and discrimination index provides a robust foundation for identifying functional distractors.
Medical education research showed that while ChatGPT-generated items were similar to human-authored items in difficulty and discrimination, the relatively low quality of distractors limited their ability to effectively differentiate higher-performing students (Lotto et al., 2024; Özer et al., 2025). Similar patterns have been observed in vocabulary testing, where Malec (2024) reported widespread distractor malfunctions in ChatGPT-generated items, resulting in substantially reduced reliability. In English listening comprehension, ChatGPT-generated stems were found to be broadly comparable to official items; however, distractor plausibility remained significantly lower (Chun & Barley, 2024). Research on RC showed that while ChatGPT generated passages comparable to human-authored ones, its distractors were less credible, limiting overall item quality (Shin & Lee, 2023).
Across multiple subjects, ChatGPT demonstrated the ability to generate high-quality stems; nevertheless, existing research suggested that its performance in MC AIG is constrained by multiple weaknesses, among which ineffective distractor generation is consistently reported.
Linguistic Feature Evaluation of ChatGPT-Generated Passages
The linguistic characteristics of passages directly affect both the cognitive load of RC and the validity of the assessment. Consequently, systematic linguistic analyses of ChatGPT-generated passages are essential for evaluating their educational applicability. Over the past 3 years, several studies have examined the linguistic differences between passages produced by ChatGPT and those written by humans (Ardeshirifar, 2025; Herbold et al., 2023; Mizumoto et al., 2024; Reviriego et al., 2024; Shen et al., 2025; Zhang & Crosthwaite, 2025; Zhang et al., 2025; Zhou et al., 2023). For instance, Zhang and Crosthwaite (2025) compared argumentative essays written by Australian undergraduates with those generated by ChatGPT-3.5, finding that ChatGPT used more formal, academic vocabulary, whereas L2 writers emphasized personal or social themes with context-rich language. Similarly, Reviriego et al. (2024) compared TOEFL essays and paraphrase tasks from humans, ChatGPT-3.5, and ChatGPT-4.0. They found ChatGPT-3.5 produced shorter passages with lower lexical diversity, while ChatGPT-4’s lexical diversity equalled or exceeded that of human writers.
Beyond lexis, other studies have also focused on syntax and cohesion. Ardeshirifar (2025) employed hand-crafted features and deep learning to differentiate between ChatGPT- and human-written passages, identifying significant differences in lexical diversity and sentence length variation. Herbold et al. (2023) compared 90 high school EFL essays with 90 ChatGPT-3 and 90 ChatGPT-4 essays. Teacher evaluations showed ChatGPT, particularly ChatGPT-4, produced higher-quality essays. Although syntactic complexity did not differ significantly across groups, lexical diversity was highest in humans compared with ChatGPT-3 but lower than ChatGPT-4. Inspired by this work, Mizumoto et al. (2024) compared essays written by Japanese EFL learners with ChatGPT outputs using NLP and random forest classification. Their results revealed that AI-generated passages exhibited greater syntactic complexity and lexical diversity, whereas humans used more modals and discourse markers.
Coh-Metrix, a Natural Language Processing (NLP) tool widely used in language assessment for textbook complexity classification (e.g., Kim & Yang, 2012; Ryu & Jeon, 2020), evaluation of test text quality (e.g., Crossley & McNamara, 2011; Ouyang et al., 2021), and diagnosis of learner language abilities (e.g., Abba et al., 2019; Chon & Shin, 2020), has also been employed on ChatGPT-human comparison. Based on a multilevel theory of discourse processing, Coh-Metrix automatically computes indices of lexical, syntactic, semantic, and discourse features. Coh-Metrix has been extensively used. Zhou et al. (2023), for example, analyzed narratives written by ChatGPT and Chinese intermediate English learners, finding that ChatGPT was superior in narrativity, word concreteness, and referential cohesion, but weaker in syntactic complexity and deep cohesion. Notably, syntactic simplicity improved only when prompts were optimized.
Shen et al. (2025) leveraged Coh-Metrix to demonstrate that passages generated by ChatGPT-4.0 and other AI tools were less grammatically complex than human texts at the same grade level. Other studies (Allaithy & Zaki, 2025; Ripoll Y Schmitz & Sonnleitner, 2025; Zhang et al., 2025) have examined the quality of AI-generated RC passages in multiple languages. Ripoll Y Schmitz & Sonnleitner, 2025 found that ChatGPT-4 performed better on informative passages, while humans excelled in narratives and coherence. Allaithy and Zaki (2025) evaluated AI-generated Arabic reading passages, and reported frequent errors, including verb misuse and syntactic inaccuracies. Likewise, Zhang et al. (2025) noted that AI-generated passages were fluent but lexically and syntactically limited, relying mainly on expert judgment rather than linguistic analysis.
Taken together, the literature remains inconsistent, suggesting that item quality may be contingent on factors such as disciplinary domain and item format. In addition, most prior research has focused on difficulty and discrimination, particularly in RC items, while neglecting evaluations of linguistic features in the generated passages (Lin & Chen, 2024; Sihite et al., 2023). Therefore, this study aims to investigate the potential and limitations of ChatGPT in generating items by systematically comparing their passages and MC items with their official counterparts, addressing five research questions:
To what extent do ChatGPT-generated and CET-4 items differ in their quantitative distribution of lower-order and higher-order reading skills?
To what extent do ChatGPT-generated items differ from CET-4 items in terms of reliability?
To what extent do ChatGPT-generated items differ from CET-4 items in terms of IRT model parameters?
To what extent do ChatGPT-generated items differ from CET-4 items in the number of functional distractors?
To what extent do ChatGPT-generated and CET-4 RC passages differ linguistically, as assessed through Coh-Metrix?
Methodology
Item Selection and Generation
Prior research suggested that providing reference passages to ChatGPT in advance enables it to produce outputs more closely aligned with human-authored passages (Ripoll Y Schmitz & Sonnleitner, 2025). Accordingly, prior to prompting ChatGPT-4o (December 2024 version) to generate items, we supplied it with the CET-4 Syllabus and 20 RC passages, each paired with five MC questions, answers, and explanations. Specific prompts to generate RC passages and items are shown in Appendix 1.
From the pool of ChatGPT-generated items that passed the quality check, seven passages with items were randomly selected as experimental materials. For comparison, seven official CET-4 RC passages, along with items, were also selected.
Data Collection
The participants were 121 first-year medical students from two classes at a university in central China, including 47 females and 74 males, aged 17–19 years. All participants were native Chinese speakers with English as a foreign language and had received at least six years of English instruction. According to records from the Academic Affairs Office, based on their national college entrance examination English scores, these participants’ English proficiency was estimated to fall within the CEFR A2–B2 range (Council of Europe, 2001), indicating that they are broadly comparable to typical CET-4 candidates (Zheng & Cheng, 2008).
The study comprised seven in-class tests, each administered by the course instructor. Each test included two RC passages: one passage with five MC items generated by ChatGPT 1 and another from CET-4. 2 The test duration was set at 20 minutes, aligning with CET-4 pacing and deemed sufficient to minimize fatigue. Furthermore, to mitigate potential bias due to prior familiarity with the content, the selected materials covered a diverse range of topics. In total, complete response data were obtained from all 121 students across the seven tests.
Dimensionality and Skill Coverage Analysis
To address the unidimensionality issue, we employed the unidimensionality testing procedure proposed by Kubinger (2003). It constructed confidence intervals for the test statistics through Monte Carlo simulation (Drasgow & Lissak, 1983).
To address RQ1, we adopted the descriptions of RC skills from the official CET-4 Syllabus. The syllabus outlines three domains, each comprising nine core skills. To further evaluate the cognitive hierarchy of skill coverage, we mapped each skill to the corresponding levels in the revised Bloom’s Taxonomy (Anderson & Krathwohl, 2001), classifying them as lower-order (Understand, Apply) or higher-order (Analyze, Evaluate) reading skills.
Two researchers independently coded each item in the seven test sets, referring to these nine skills. Classification was based on the type of comprehension reflected in the item stem and its correct answer. After independent coding, the two researchers compared their results, and any discrepancies were resolved through discussion until consensus was reached. Inter-rater reliability was subsequently assessed using Cohen’s κ coefficient (κ = .83), which confirmed the robustness of the coding procedure. The finalized coding results were then aggregated, and coverage proportions for each skill category were calculated (see Table 1). To address RQ2, the model-based reliability coefficient WLE-REL (Warm, 1989) was computed within the IRT framework. WLE-REL provides an estimate of the reliability of Warm likelihood estimates of ability (Davier, 2011) and is particularly sensitive to the alignment between item difficulty and examinee ability. Given that this study employed the test for low-stakes assessment purposes, values above .80 are considered acceptable.
Item Response Analysis
In psychometrics, item calibration refers to the empirical process of establishing the relationship between the measurement scale of an instrument and the latent ability it is intended to assess (Mari et al., 2023). To address RQ3, the study employed the two-parameter logistic (2-PL) model to calibrate the items. This model specifies the probability of a test-taker with a given ability level responding correctly to a particular item as follows:
The probability of examinee i answering item j correctly is determined jointly by the examinee’s ability (θ), the item’s difficulty parameter (b), and the item’s discrimination parameter (a). Difficulty indicates the ability level at which an examinee has a 50% probability of answering the item correctly. In contrast, discrimination represents the steepness of the slope of the item characteristic curve. A higher discrimination reflects greater sensitivity in distinguishing among examinees of different ability levels (Masters, 1988). High-quality item sets typically exhibit discrimination clustered around 1, with a slight rightward skew, indicating moderately high discriminating power (>1).
The validity of the newly developed items was examined following procedures proposed in prior research (Mislevy et al., 2003; Wilson, 2023). A concurrent calibration method was employed to calibrate the ChatGPT-generated and the CET-4 items simultaneously. Compared with separate calibration, this method is more efficient, yields smaller standard errors, and requires fewer assumptions (Wingersky & Lord, 1984). Given the assumption that both item sets measure the same latent construct, and considering that the missing data were missing at random, the resulting parameter estimates should not suffer from systematic bias (DeMars, 2002).
To evaluate the model–data fit, Infit and Outfit mean square (MNSQ) statistics, as well as standardized Z-scores (ZSTD), were analyzed to assess item–model alignment and identify potential aberrant response patterns. Regarding measurement precision, Wright maps were generated to visually compare the coverage of item difficulty with the distribution of examinee abilities (Torres Irribarra & Freund, 2014). The test information function (TIF) was plotted to compare the measurement accuracy of the two item sets across different ability levels. All IRT analyses were conducted in R (version 4.4.1; R Core Team, 2024) using the TAM package. Item parameters were estimated using marginal maximum likelihood estimation with the EM algorithm, and person abilities were estimated as weighted likelihood estimates (WLEs). Item-fit statistics, Wright maps, test information functions, and WLE-REL were derived from the fitted model.
Distractor Analysis
To address RQ4, we adopted the evaluative framework proposed by Tarrant et al. (2009), examining two indicators, selection rate (SR) and discrimination power (DP), to determine the functionality of distractors. According to the SR criterion, a distractor was considered non-functional if fewer than 5% of examinees selected it. This 5% threshold, a standard benchmark in the item analysis literature (Haladyna & Downing, 1989; Tarrant et al., 2009), ensures that distractors are minimally plausible by attracting at least a small proportion of responses while avoiding over-identification of functional distractors as non-functional due to random variation.
When the SR reached or exceeded 5%, its discrimination was further examined. DP was introduced to measure the extent to which low- and high-performing examinees differed in their selection of a given option. The index was calculated as follows:
PL denotes the proportion of low-performing examinees (bottom 27%) selecting the option, and PH the proportion of high-performing examinees (top 27%). Ideally, the correct answer should display PH>PL, resulting in a positive DP, whereas high-quality distractors are expected to attract more low-ability examinees, yielding PL>PH and a negative DP.
Based on this standard, distractors were classified as functional if their SR was at least 5% and their DP was negative. Conversely, distractors with a SR below 5% or a positive DP were categorized as non-functional. Applying these criteria, we evaluated and categorized all distractors from the ChatGPT-generated and the CET-4 items, enabling a comparative analysis of distractor quality.
Linguistic Feature Analysis
To address RQ5, we utilized Coh-Metrix 3.0 to analyze the linguistic features of RC passages generated by ChatGPT and those from the official CET-4. Coh-Metrix is a computational linguistics tool that quantifies various linguistic and discourse features, widely applied in second language research to assess text complexity and readability (Crossley et al., 2008; McNamara et al., 2014). The analysis focused on six main dimensions: discourse structure and syntactic complexity, which were measured by indices such as sentence count, word count, and syntactic features like noun phrase density and embeddedness; lexical features, analyzed through measures of diversity, frequency, and familiarity, where higher diversity combined with lower frequency indicates greater lexical complexity; discourse cohesion and coherence, assessed by lexical and semantic overlap, indicating how repetition and connections between sentences enhance comprehension; the situation model, which captures mental representations of causation, time, space, and protagonists, ensuring textual cohesion through cohesive devices; word information, where part-of-speech distribution, including the use of pronouns, reflects the communicative function and register of the text; and readability, measured by classic indices like Flesch Reading Ease, Flesch–Kincaid Grade Level, and the Second Language Readability Index, with additional Coh-Metrix measures evaluating text ease in terms of narrativity, syntactic simplicity, and cohesion.
Results
Comparison of Reading Skill Coverage
As shown in Table 1, the CET-4 and ChatGPT-generated items differed markedly in the distribution of lower-order and higher-order reading skills.
ChatGPT-generated items were predominantly concentrated in lower-order reading skills, particularly Skills (1), (2), (3), and (7). These lower-order skills occurred 41 times in the ChatGPT items, compared with 19 times in the CET-4 items. ChatGPT exceeded CET-4 in all four lower-order skills, most notably in Skill (2) and Skill (7), indicating a strong emphasis on understanding explicitly stated information and applying contextual cues for localized processing.
In contrast, CET-4 items demonstrated a clear quantitative advantage in higher-order reading skills, including Skills (4), (5), (6), (8), and (9). These higher-order skills appeared 58 times in the CET-4 items, substantially exceeding the 33 occurrences in the ChatGPT-generated items. The most significant discrepancies were observed in Skills (8) and (9), both of which require discourse-level integration and analytical processing beyond sentence-level comprehension.
Measurement Structure and Reliability
Unidimensionality was examined using a statistical test of dimensionality (Kubinger, 2003), with confidence intervals established through Monte Carlo simulation (Drasgow & Lissak, 1983), to verify whether the ChatGPT-generated and the CET-4 items satisfied the unidimensionality assumption required for IRT. The results indicated that the p-values for the ChatGPT items (p = .15) and CET-4 items (p = .25) were both greater than the significance threshold of .05, supporting the assumption of unidimensionality.
The WLE-REL for CET-4 items was .88, for ChatGPT items .82, and for the combined test .85; all values exceeded the .80 threshold, indicating satisfactory reliability. Notably, the CET-4 items yielded higher WLE-REL than the ChatGPT-generated items, suggesting that the CET-4 items provided more reliable measurement information across the examinee ability distribution.
Parameter Analysis Based on the 2-PL Model
As shown in Figure 1, the analysis results indicated that the difficulty of ChatGPT-generated items ranged from −4.79 to −0.77, with values concentrated in the negative logit region, suggesting that the items were generally easy. By contrast, the CET-4 items had a difficulty range of −2.20 to 0.44, covering a greater proportion of items around zero or in the positive range, which reflects a higher proportion of relatively difficult items. Results from independent-samples T-tests showed a statistically significant difference in mean difficulty between the two sets (p < .001). ChatGPT-generated items exhibited significantly lower difficulty levels compared to CET-4 items, suggesting an overall insufficiency of challenge. Distribution of examinee ability and item difficulties on a common latent scale. Note. All ChatGPT-generated items and CET-4 items were concurrently calibrated in a single 2PL model. The histogram in the upper panel represents the ability distribution of the full sample, and the lower panel displays the difficulty locations of CET-4 and ChatGPT-generated items on the same logit scale
Regarding discrimination, ChatGPT items exhibited a range of 0.17 to 1.02, which was relatively low overall. In contrast, CET-4 items demonstrated a broader and higher range of discrimination, from 0.60 to 1.70, with a significantly higher mean value compared to ChatGPT items. T-test results further confirmed this significant difference (p < .001), suggesting that CET-4 items were more effective in differentiating examinees across ability levels. 3
In terms of model fit, the analysis examined Infit/Outfit mean square (MNSQ) and standardized Z-scores (ZSTD). The Infit index primarily evaluates the consistency of examinees’ responses to items that are close to their ability levels, whereas the Outfit index is more sensitive to unexpected or aberrant responses. The MNSQ value reflects the degree of item–model fit, with a theoretical expectation of 1. Values between 0.7 and 1.3 are generally regarded as optimal. The ZSTD statistic represents standardized residuals, with an expected value of 0. Values of |ZSTD| < 2 are typically considered acceptable.
The CET-4 items demonstrated a more concentrated fit distribution within the acceptable range (MNSQ between 0.85 and 1.15), with fewer extreme values, indicating higher measurement stability. In contrast, the ChatGPT-generated items exhibited greater variability in fit indices, with some items displaying excessively high (e.g., Outfit MNSQ >2) or excessively low (<0.5) values. Correspondingly, several ZSTD values exceeded the absolute-value threshold of 2, suggesting the presence of unexpected response patterns among examinees. This instability is consistent with the relatively lower discrimination and easier difficulty observed in the ChatGPT items. 4
To further compare the measurement efficiency of the two item sets, Test Information Functions (TIFs) are presented in Figure 2, which reveal the precision of measurement across different ability levels (Lord, 1980). The results showed that the CET-4 items peaked around θ = 0, with a maximum test information value exceeding 11, and maintained high measurement precision within the range of θ = −2.5 to θ = 1.5, thus covering a broader spectrum of ability levels. By contrast, the ChatGPT-generated items peaked around θ = −1, with a maximum information value of approximately 4. The information was concentrated within the θ = −2 to 0, suggesting some measurement effectiveness for examinees at lower to medium ability levels but a clear lack of precision for higher-performing ones (θ > 1). Comparison of TIFs between CET-4 and ChatGPT items
Analysis of Distractors
Figure 3 illustrates the overall distribution of distractor classifications in the ChatGPT-generated and the CET-4 items. A chi-square test of the distribution revealed a significant difference between the two sets, χ
2
(1) = 39.00, p < .001, indicating a highly significant discrepancy in distractor quality. Specifically, of the 105 distractors generated by ChatGPT, only 24 (22.9%) met the criteria for functional distractors, while as many as 81 (77.1%) were judged non-functional. This finding suggests that most ChatGPT-generated distractors failed to function as intended, being either too transparent, implausible, or even more attractive to high-ability examinees. ChatGPT vs CET-4 Distractor Classification Comparison
In contrast, the CET-4 official items demonstrated a markedly higher level of distractor quality. Among the identical 105 distractors, 70 (66.7%) were classified as functional distractors, whereas only 35 (33.3%) were judged non-functional. This distribution reflects the effectiveness of the CET-4 test development and review process in ensuring that the majority of distractors are plausible. CET-4 distractors not only increase the difficulty of the items but also enhance discrimination by effectively misleading lower-performing examinees.
Results of Linguistic Feature Analysis
Comparison of Coh-Metrix Linguistic Features Between ChatGPT-Generated Passages and CET-4 Official Passages
Descriptive Indices and Syntactic Complexity
As shown in Table 2, the ChatGPT-generated passages exhibited a significantly lower total number of words (DESWC = 283.429) compared to the CET-4 RC passages (DESWC = 339.714). However, both the total number of sentences (DESSC = 16.429) and the average number of sentences per paragraph (DESPL = 3.205) were significantly higher in the ChatGPT-generated than in the CET-4 passages (DESSC = 12.571; DESPL = 2.301). Furthermore, the mean sentence length measured in words (DESSL = 17.489) and its standard deviation (DESSLd = 5.281) were significantly lower in the ChatGPT-generated passages relative to the CET-4 passages (DESSL = 28.864; DESSLd = 15.199). In contrast, the average number of syllables per word (DESWLsy = 1.860) and the average number of letters per word (DESWLlt = 5.595) were significantly higher in the ChatGPT-generated passages than in the CET-4 passages (DESWLsy = 1.674; DESWLlt = 5.152). These results indicate that, compared to the CET-4 passages, ChatGPT generates shorter passages and shorter sentences on average, but contains longer words.
As for syntactic complexity, the ChatGPT-generated passages received significantly higher scores than the CET-4 passages on SYNMEDwrd (0.936), SYNMEDlem (0.928), SYNSTRUTa (0.111), and SYNSTRUTt (0.110). SYNMEDwrd evaluates the average difference in word order and lexical choice between adjacent sentences and SYNMEDlem performs a similar comparison based on lemmas (ignoring inflectional variations). Higher values indicate greater dissimilarity in word order and lexical choice between sentences. Two additional measures, SYNSTRUTa and SYNSTRUTt, quantify syntactic similarity: the former between adjacent sentences, and the latter across all sentence pairs within a paragraph. Higher values reflect greater similarity in syntactic structures. This indicates that while adjacent sentences in ChatGPT-generated passages exhibit greater variation in word choice and order, their syntactic structures are more uniform compared to those in the CET-4 passages.
In addition, the incidence of negation (DRNEG) was significantly lower in the ChatGPT-generated passages (4.859) than in the CET-4 passages (14.338), suggesting that ChatGPT tends to use fewer negative constructions.
Lexical Features
Lexical features of all passages were analyzed through lexical complexity and word information. All four lexical diversity indices in Table 2, including LDTTRc (0.818), LDTTRa (0.655), LDMTLDa (175.608), and LDVOCDa (148.254), were significantly higher in the ChatGPT-generated passages than in the CET-4 passages. LDTTRc and LDTTRa represent the type-token ratio for content words and all words, respectively. Since these basic ratios are sensitive to text length, LDMTLD (Measure of Textual Lexical Diversity) and LDVOCD (Vocabulary Diversity) provide more robust, length-independent estimates of lexical variation. LDMTLD calculates the mean length of consecutive word strings maintaining a threshold diversity level, while LDVOCD compares the lexical profile of the text against a reference corpus. The significantly higher values across all four measures suggest that the ChatGPT-generated passages exhibit greater lexical diversity than the CET-4 passages.
As indicated in Table 2, the ChatGPT-generated passages scored significantly higher than the CET-4 passages on WRDADJ (137.576), which reflects the incidence of adjectives, and on WRDHYPn (6.656), which indicates the degree of nominal hypernymy. This suggests that the ChatGPT-generated passages use more adjectives and more abstract nouns. Regarding word frequency, three indices were analyzed: WRDFRQc (raw word frequency for content words), WRDFRQa (logarithm of word frequency for all words), and WRDFRQmc (logarithm of the minimum frequency of content words). The ChatGPT-generated passages showed significantly lower values than the CET-4 passages on WRDFRQc (2.028 vs. 2.243) and WRDFRQa (2.679 vs. 2.883), but a significantly higher value on WRDFRQmc (1.044 vs. 0.432). This pattern indicates that while most words (both content and function words) in the ChatGPT-generated passages tend to be high-frequency, there is also a notable presence of very low-frequency content words.
Cohesion
As shown in Table 2, the ChatGPT-generated passages showed significantly lower values than the CET-4 passages on four key measures of referential cohesion: local argument overlap between adjacent sentences (CRFAO1 = 0.321), global argument overlap across all sentences (CRFAOa = 0.362), local content word overlap between adjacent sentences (CRFCWO1 = 0.040), and its standard deviation (CRFCWO1d = 0.052). These consistently lower scores suggest that the ChatGPT-generated passages employ significantly less argument overlap and content word overlap as cohesive devices.
The ChatGPT-generated passages showed a significantly lower value on LSAGNd (0.111) than the CET-4 passages (0.129). LSAGN measures the proportion of given versus new information in each sentence relative to all prior text, using Latent Semantic Analysis (LSA) to compute semantic similarity. LSAGNd represents the standard deviation of LSAGN scores, reflecting the consistency in the distribution of given and new information across sentences. The lower value for ChatGPT-generated passages suggests a more uneven distribution of given information, indicating fewer semantic ties between sentences.
Situation Model
Significant differences were observed in causal cohesion between the two text types in Table 2. The ChatGPT-generated passages scored significantly higher than the CET-4 passages on two indices: the incidence of causal verbs (SMCAUSv = 42.407 vs. CET-4’s 23.474) and the incidence of both causal verbs and explicit causal particles (SMCAUSvp = 50.991 vs. CET-4’s 38.145). These results indicate that the ChatGPT-generated passages contain a denser presence of explicit causal markers, including both verbs denotating changes of state or action effects, and connectives such as “because” or “therefore.” This heightened causal explicitness may facilitate the construction of a more coherent and deeper situation model for readers.
Text Readability
Text Easability Components’ indices measure a text’s easability on different dimensions, each providing both a “z-score” (the “z” suffix, e.g., PCNARz) indicating its statistical deviation from a norm and a “percentile” (the “p” suffix, e.g., PCNARp) showing the percentage of passages it is easier than. PCNAR measures narrativity, or how story-like and familiar the content is. PCSYN assesses syntactic simplicity and PCREF evaluates referential cohesion. PCCONN measures connectivity, reflecting the use of explicit connectives to clarify relationships in the text. In all cases, a higher percentile score means the text is easier to process on that specific dimension.
Among the seven indices shown in Table 2, the PCSYNz (0.638) and PCSYNp (72.199) of the ChatGPT-generated passages are significantly higher than those of the CET-4 passages (−0.956 and 20.899). However, values of the other indices (PCNARz, PCNARp, PCREFz, PCREFp, and PCCONNz) of the ChatGPT-generated passages are significantly lower than those of the CET-4 passages. This result shows that while ChatGPT-generated passages are syntactically simpler, they are more challenging in terms of narrative structure and cohesion.
The L2 Readability index (RDL2) is a unified measure designed to estimate text complexity for second-language readers. It integrates factors such as content word overlap, syntactic similarity, and word frequency. As shown in Table 2, the ChatGPT-generated passages received a significantly lower RDL2 score (8.887) than the CET-4 passages (13.902), indicating that the ChatGPT-generated passages are less readable and pose greater comprehension challenges for L2 learners.
Discussion
Skill Coverage of ChatGPT-Generated Items
Tests of unidimensionality indicated that both ChatGPT-generated and CET-4 items met the unidimensionality assumption, satisfying the requirements of IRT. This is consistent with the findings of Bhandari et al. (2024), who reported that generated items did not violate the unidimensionality assumption.
The skill coverage analysis revealed that ChatGPT showed strengths in lower-order reading skills but was weaker in higher-order ones, confirming H1. This imbalance in coverage may weaken the validity of the test, as uneven representation of reading skills can bias score interpretation (Alderson, 2000). These findings support those of Zhang et al. (2025), who emphasized ChatGPT’s limitations in skill coverage. The uneven skill distribution aligns with the observation that ChatGPT still struggles to capture higher-order reading skills in the AIG (Wen & Chu, 2025). Such imbalances likely stem from ChatGPT training paradigms that prioritize high-frequency, explicit patterns via next-token prediction and RLHF (OpenAI, 2023; Ouyang et al., 2022), in contrast to expert item design grounded in learner cognition and developmental trajectories (Gierl et al., 2017; Lee et al., 2024), highlighting the need for human review (Lin & Chen, 2024).
Bhandari et al. (2024) noted that ChatGPT aligned effectively with target skills in mathematics. Mathematics often employs highly structured, quasi-computational frameworks, which facilitate alignment. English RC involves NLP, increasing the difficulty of achieving balanced skill coverage in ChatGPT-generated items. Therefore, the differences may be attributed to disciplinary characteristics.
Psychometric Measurement Properties of ChatGPT-Generated Items
The WLE-REL results supported H2: ChatGPT-generated items yielded lower reliability (WLE-REL = .82) than CET-4 items (WLE-REL = .88). This difference can be understood through the lens of IRT: reliability is maximized when item difficulty is well-matched to the full range of examinee abilities (Embretson & Reise, 2000; Lord, 1980). As revealed by the Wright maps and TIF analysis, ChatGPT-generated items were concentrated in the lower difficulty range, providing sufficient measurement information only for lower-ability examinees while leaving higher-ability examinees measured with greater error. CET-4 items, with a wider and more balanced difficulty distribution, provided more consistent measurement information across the ability continuum, resulting in higher reliability.
The mean difficulty of ChatGPT-generated items was significantly lower than that of CET-4 items. This finding aligns with some recent studies (Young et al., 2025) but stands in contrast to other empirical investigations (Bhandari et al., 2024; Lin & Chen, 2024; Mendoza & Zúñiga, 2025). The average discrimination of ChatGPT-generated items was significantly lower than that of CET-4 items, which stands in contrast to previous research reporting high discrimination for ChatGPT-generated items (Young et al., 2025), and is comparable to that produced by human test developers (Isley et al., 2025; Mendoza & Zúñiga, 2025). Interestingly, although the ChatGPT-generated passages were less readable, the IRT analysis revealed that the corresponding items were easier and exhibited lower discrimination. This paradox may stem from the combined effect of passage complexity and item construction. As noted by Freedle and Kostin (1999), the content and structure of reading passages can account for at least 33% of the variance in the difficulty of RC items; however, greater textual complexity does not necessarily lead to higher overall item difficulty, since the influence of passage complexity may be offset by weaknesses in item design. Malfunctioning distractors enable students to eliminate incorrect options with relative ease, without fully comprehending the passages. Furthermore, the overrepresentation of lower-order reading skills likely contributed to the items’ reduced difficulty. This agrees with the findings of Basaraba et al. (2012), who reported that items for higher-order reading skills were more difficult than those for lower-order ones. In contrast, Lin and Chen (2024) reported higher discrimination for ChatGPT items after removing outliers; however, their study included only 18 ChatGPT items and 10 CET-4 items, and the small sample may have made the results more susceptible to individual items, yielding relatively optimistic conclusions. Another reason is that their design relied on official passages with ChatGPT generating only the items, rather than requiring ChatGPT to generate both passages and items, which may have reduced variability.
TIFs analysis in this work highlighted a clear contrast between ChatGPT-generated and CET-4 items. While the latter provided precise measurement across a broad ability range, the former concentrated information in the low-to-mid range, offering limited precision for high-performing examinees. This finding resonates with Young et al. (2025), who similarly noted that ChatGPT-generated items were disproportionately easy and thus less effective for differentiating advanced learners. However, it diverges from the results of Lin and Chen (2024), who reported relatively stronger measurement precision and acceptable fit indices for ChatGPT items. This discrepancy may be attributed to the larger number of ChatGPT-generated items analyzed in this study, which likely produced more conservative estimates of test information, making limitations in measurement precision at higher-ability levels more apparent. Collectively, the findings on difficulty, discrimination, and TIFs analysis all support H3.
Therefore, ChatGPT-generated items may be more suitable for routine assignments and basic training. However, their use in high-stakes testing should be approached with caution, aligning with prior studies (Bhandari et al., 2024; Young et al., 2025) that recommend using ChatGPT-generated items primarily in low-stakes testing.
Limitations in the Distractor Design of ChatGPT-Generated Items
In this article, 22.9% of ChatGPT-generated distractors were functional, compared with 66.7% in CET-4 items, confirming H4 and underlining that distractor generation remains one of the most significant challenges for ChatGPT in AIG (Shin & Lee, 2023). This psychometric shortfall directly erodes item discrimination, as non-functional distractors flatten the ability spectrum, privileging low-order elimination over deep comprehension (Ludewig et al., 2023), reflecting the Transformer architecture’s bias toward fluent, high-probability outputs (Brown et al., 2020; Vaswani et al., 2017).
This result supports Malec’s (2024) findings in vocabulary testing, which demonstrated that many ChatGPT-generated distractors were non-functional and specific distractors mirroring the discrimination patterns of correct answers. The high proportion of non-functional distractors (77.1%) observed in this study further substantiates this limitation. Similarly, researchers in medical education reported that deficiencies in distractor quality constrained the capacity of ChatGPT-generated items to distinguish high-performing examinees (Lotto et al., 2024), underscoring the need for human review. Our study extends this line of evidence by quantifying the issue: ChatGPT distractors either failed due to the extremely low SR or weakened validity by showing reversed discrimination patterns, which also explains the consistently lower discrimination observed relative to CET-4 items. Lin and Chen (2024) reported only 22% problematic distractors in their study, lower than the 77.1% identified in ours. The discrepancy may stem from differing evaluation standards, as our study employed a stricter criterion, thereby offering a more accurate reflection of psychometric functionality and aligning with the scientific requirements outlined by Tarrant et al. (2009). Moreover, in our study, ChatGPT was required to generate both passages and items, rather than items alone. The increased complexity of this dual-generation task may have compromised the quality of the distractors, as the added load on ChatGPT likely constrained the model’s ability to maintain nuanced error simulation (Mizumoto et al., 2024).
Effective distractor design requires item developers to anticipate plausible error patterns in test-takers’ reasoning (Gierl et al., 2017; Ludewig et al., 2023). The training of ChatGPT is accomplished through next-token prediction during pre-training and an alignment process based on RLHF, both of which emphasize the truth and adherence to desired behavior (OpenAI, 2023). Accordingly, ChatGPT’s training objective is to generate coherent and contextually appropriate text rather than to simulate illogical or erroneous reasoning. This misalignment contributes to poor distractor quality and limits higher-order error modeling, as LLMs struggle with negation and logical nuance (Kassner & Schütze, 2020; Ouyang et al., 2022). Chun and Barley (2024) confirmed this view in their study: while ChatGPT-generated items matched human-developed items in linguistic clarity, they lagged significantly in distractor plausibility, which ultimately became the primary weakness of item quality. Consequently, when utilizing ChatGPT for the generation of MC items, it is essential to prioritize the quality of distractors as a critical component of the quality control process (Shin & Lee, 2023).
Linguistic Feature Differences and Their Impact on Test Validity
Coh-Metrix results reveal that ChatGPT passages differ from CET-4 texts at the lexical, syntactic, and cohesion levels. ChatGPT texts exhibit distinct lexical features, such as longer words, higher diversity, more frequent use of adjectives and hypernyms, and a mix of common and excessively rare vocabulary. These findings align with prior studies reporting increased lexical diversity in newer models (Herbold et al., 2023; Reviriego et al., 2024), irregular lexical patterns (Ardeshirifar, 2025), and neutral and abstract linguistic style (Ardeshirifar, 2025; Lew, 2023), while further highlighting the trade-off between sophistication and comprehension. These traits likely stem from ChatGPT’s training on large-scale corpora, which promotes lexical richness but also introduces discordant word choices. This blend enhances textual sophistication at the potential cost of readability, as abrupt shifts to low-frequency vocabulary may disrupt comprehension and increase cognitive load (McNamara et al., 2014).
ChatGPT passages also tend to include shorter sentences, fewer negatives, and relatively simple structures. Adjacent sentences may vary lexically but show limited structural variation, which can create a monotonous rhythm. Although this may seem counterintuitive, it becomes plausible given the benchmark of expert-authored CET-4 passages. Herbold et al. (2023) reported no major syntactic differences between ChatGPT-3 and learner essays revised by native speakers, implying that while ChatGPT may surpass many EFL learners in complexity (Mizumoto et al., 2024; Zhou et al., 2023), it may still fall short of experts. This repetitiveness also aligns with earlier findings that ChatGPT tends to recycle syntactic templates (Chakrabarty et al., 2025) and relies on high-probability grammatical frames (Ardeshirifar, 2025; Shaib et al., 2024), varying mainly through lexical substitution. In contrast, skilled authors deliberately alternate sentence patterns to enhance rhythm and rhetorical effect. Thus, ChatGPT passages may appear lexically rich but structurally flat, lacking the syntactic variation typical of expert writing.
Cohesion further distinguishes the two text types. ChatGPT passages appear less cohesive, using fewer referential devices and conjunctions, with less given information. Prior work similarly notes weaker cohesion in Gen-AI output (Herbold et al., 2023; Mizumoto et al., 2024). While CET-4 experts adjust cohesion density for balance, ChatGPT’s relative lack of cohesion likely increases readers’ cognitive load and, combined with lexical variation, reduces readability. ChatGPT passages also show an emphasis on abstract reasoning rather than storytelling. This is likely because its training data emphasizes informational and explanatory content, and its architecture is optimized for generating structured, coherent responses.
Overall, ChatGPT passages are characterized by syntactic repetitiveness, irregular lexical distribution, and weaker cohesion, contributing to lower readability compared with CET-4 texts. While ChatGPT texts may appear linguistically rich, they impose greater cognitive load and reduced accessibility. These differences raise validity concerns for high-stakes testing, where text readability must align with construct representation and fairness principles.
Conclusion
This study extended previous research by employing ChatGPT to generate both items and passages, and by using IRT and Coh-Metrix to quantitatively assess item quality and the linguistic features of the passages. The results indicated that ChatGPT demonstrated potential in generating both RC items and their corresponding passages. However, the generated items exhibited several limitations in item difficulty, discrimination, measurement precision, and distractor functionality. At the same time, the passages showed irregular lexical distribution, reduced cohesion, and lower readability than official CET-4 materials. Most notably, this study identified a previously undocumented paradox: linguistically complex passages are associated with relatively easy, weakly discriminating items. This disconnect can be attributed to weak distractor functioning, allowing test-takers to answer items through option elimination rather than through complete comprehension, and to the overrepresentation of lower-order reading skills, which may reduce cognitive demands and limit reliance on global, inferential processing of the passages.
This study made significant theoretical and practical contributions to AIG in language assessment. Theoretically, it advanced IRT applications by demonstrating that ChatGPT-generated items satisfy unidimensionality assumptions while exhibiting lower difficulty and discrimination than official items. The research contributed to validity theory by revealing imbalances in skill coverage: ChatGPT underrepresents higher-order reading skills. Practically, the findings provided evidence-based guidelines for educators and testing practitioners implementing ChatGPT in assessment development. For teachers, ChatGPT serves as a supplementary tool for low-stakes language assessment; however, human review remains essential for ensuring the quality of distractors and comprehensive skill coverage. The comprehensive evaluation framework, which combined IRT analysis, skill coverage assessment, and linguistic feature analysis, provided quality control protocols for testing practitioners while identifying specific algorithmic improvements for educational technology developers regarding distractor plausibility and difficulty calibration.
This research had several limitations that should be acknowledged. First, although the participants were broadly comparable to typical CET-4 candidates in age and estimated English proficiency, the sample was limited to 121 first-year medical students from a single university. As a result, the sample cannot be assumed to fully represent the broader national CET-4 population, which is more heterogeneous in institutional context, disciplinary background, and proficiency distribution. Therefore, the absolute psychometric estimates reported in this study should be interpreted cautiously as sample-bound. Second, the analysis focused exclusively on RC item types; therefore, the conclusions may not extend to other language skill domains. Third, although the ChatGPT-based AIG process employed carefully specified prompts and constraints, no systematic comparison was conducted across different prompt designs or model versions, potentially overlooking the role of prompt engineering in shaping item quality. Finally, because ChatGPT generated both the reading passages and the items, it was difficult to disentangle whether observed performance differences stem from passage-level linguistic complexity, item difficulty, or distractor quality. This confounding between passages and items limited causal interpretations regarding the specific sources of difficulty in the generated items.
Future research could (a) expand sample size to include learners from different regions, majors, and proficiency levels; (b) extend the scope to include other question types to evaluate ChatGPT’s potential in language assessment comprehensively; (c) explore the impact of different prompts on item quality; and (d) combine quantitative and expert qualitative analysis to understand distractor failures better.
Footnotes
Acknowledgments
The authors would like to thank all participants who contributed to data collection and the reviewers for their insightful comments and suggestions that helped improve this manuscript.
Ethical Considerations
Ethical approval for this study was obtained from the Ethics Committee of Hunan Normal University.
Consent to Participate
All participants provided informed consent prior to participation.
Funding
The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: The 2025 Graduate Research Innovation Project of the School of Foreign Languages, Hunan Normal University: Research on Automatic Question Generation for English Reading Comprehension Using Generative AI.
Declaration of Conflicting Interests
The authors have no competing interests to declare that are relevant to the content of this article.
Data Availability Statement
The data that support the findings of this study are available on request from the corresponding author, Li, Wanqing, upon reasonable request.
