
Book review
Select search scope: search across all journals or within the current journal
1-20 of 20 articles


Recent initiatives have sought to integrate dynamic assessment and diagnostic assessment to facilitate language learning. Extending this line of innovation, the present study uses the example of refusals to illustrate the development and validation of a computerized dynamic diagnostic assessment of pragmatic competence (CDDA-P). To enhance the authenticity of assessed pragmatic performance, a bottom-up approach was adopted in item design, in which response options were empirically derived from a corpus of productions by 335 language users. The effectiveness of the CDDA-P was examined from three key perspectives: its ability to diagnose learners’ strengths and weaknesses, to identify their zones of proximal development (ZPDs) and learning potential, and to promote the development of learners’ pragmatic performance. A pretest–immediate posttest-delayed posttest design was employed to track changes in 66 Chinese learners’ performance before and after the implementation of the CDDA-P. Findings reveal that the CDDA-P can provide a fine-grained diagnosis of learners’ strengths and weaknesses in performing L2 refusals and identify their diverse ZPDs and learning potential. Furthermore, learners demonstrated significant improvement after mediation. This study presents both theoretical and methodological implications for the development of integrated dynamic and diagnostic language assessment, while also offering insights into promoting L2 pragmatic competence through assessment.
As an important area of exploration within spoken dialogue systems (SDSs), empathetic spoken dialogue systems (E-SDSs) can provide emotional support for interlocutors, and this feature has the potential to be embedded in language teaching, learning, and assessment. However, the potential of E-SDSs in second language (L2) oral assessment remains underexplored, particularly regarding their ability to elicit interactional competence (IC) and learners’ perceptions of the two systems. To address these gaps, this study compares learners’ interaction with a neutral spoken dialogue system (N-SDS) with limited empathetic capability as a reference condition and an E-SDS, with an aim to explore its potential for L2 oral assessment. Twenty-five L2 learners completed two tasks (E-SDS and N-SDS). Their oral performances between the two tasks were examined in terms of IC features, and their perceptions of the E-SDS were also investigated through semi-structured interviews. Results indicated the E-SDS tends to elicit higher frequencies of certain IC features. Learners generally perceived the E-SDS as a competent, trustworthy, and emotionally supportive interlocutor, but certain concerns were also expressed concerning system design and technical limitations.
A key psychometric phenomenon in educational testing is differential item functioning (DIF), which evaluates whether specific test items function differently across subgroups of examinees who have the same level of the underlying (latent) ability. DIF happens when examinees with the same latent trait have different probabilities of correctly responding to a test item, influenced by their subgroup membership. This study aims to illustrate the application of the recursive partitioning Rasch tree model, shortly called the Rasch tree, to examine DIF in a large-scale German reading comprehension test for fifth-grade elementary students across gender, socioeconomic status, immigration status, personality disposition, and the need for cognition. Unlike the conventional DIF detection methods, the Rasch tree does not require a pre-specification of groups for exploring DIF, and continuous covariates can be easily included in the analysis. To investigate DIF of the test, item responses of 4,252 students were analyzed. The Rasch tree analysis generated nine non-predefined nodes, with slightly different patterns of item difficulties. Eleven items were flagged as exhibiting moderate and large DIF in the four splitting nodes. No splits were produced based on the need for cognition. The results indicated that the combination of the covariates impacted students’ test performance.
As a nationally accredited test, the Test of Proficiency in Korean (TOPIK) assesses general Korean proficiency as a foreign language for various high-stakes purposes, such as university admissions, employment, and visa issuance. For this reason, the social impact of TOPIK on test takers is significant and cannot be underestimated. Despite the growing number of international test takers and stakeholders, there is limited validation research evaluating TOPIK’s various uses, as well as critical evaluation of the test itself. Thus, this test review provides an overview of the history, test purposes and use, design, and administration of TOPIK and offers an appraisal of its strengths and challenges. While the test scores are broadly utilized for their intended purposes and provide some evidence of language development across four skills in Korean, the lack of publicly available information on the test construct, psychometric properties, and standard-setting procedures makes it difficult to fully evaluate the validity and reliability of the scores. Furthermore, additional evidence and validation research are needed to examine the generalizability of TOPIK scores beyond academic domains, given the test’s diverse applications. Addressing these gaps is critical to meet the needs of various stakeholders and to strengthen the overall validity of the test.
How to optimize ordering the parts of a language test (e.g., items, tasks, and sections) is often overlooked or assumed to be self-evident. However, ordering choices are an essential component of test design as they may impact test takers’ objective performance, perceived performance, or affective states. In this paper, we report on a systematic review of test component ordering research, focusing on 88 studies from 1933 to 2023. We provide a narrative synthesis, describing typical outcome variables (e.g., test-taker performance), independent variables (e.g., order of difficulty), and mediating variables (e.g., language proficiency). Key findings, mostly from higher educational contexts, indicate that easy-to-hard ordering may lead to better performance, though these effects are mediated by various test-taker characteristics, especially anxiety and proficiency. While section ordering by language skills is common practice in language testing, there is scant empirical support for this approach, or for organizing tests by content or format. We discuss the implications of these findings for language test developers and suggest avenues for future research, particularly the need for more studies on ordering effects in language assessment contexts.
While ethical and moral philosophy underpins core concepts in language testing, few publications in the field engage with moral philosophy in a sustained or systematic way. A notable exception is the International Language Testing Association’s (ILTA) Code of Ethics (COE), ratified in 2000, which continues to guide professional practice. As ILTA initiates a revision of the COE, this paper examines the objectives and philosophical foundations of the original COE. Drawing on the document itself and key writings by its principal author, Alan Davies, and contemporary authors, we offer an analysis that reveals both the strategic rationale behind the Code and the unresolved tensions within it. Our aim is to deepen the field’s understanding of the ethical frameworks that have shaped the language testing profession as ILTA traces a way forward in revising the COE.
Although research on young learners’ (YL) second language (L2) proficiency gains is on the rise, few studies have investigated YLs’ language progression in an English as a Foreign Language (EFL) context from a longitudinal perspective and the effects of background variables on their learning outcomes. This current study examined YL’s English proficiency gains over 6 months. It further investigated how the gains in proficiency were related to the learners’ educational context and background variables. In total, 110 YLs from three distinct educational settings in Mexico took the TOEFL® Primary tests at the beginning and the end of 6 months. They completed a set of background questionnaires at the start, midpoint, and end of the study. Results showed that the learners in the advantaged background with varying resources made significant gains in speaking proficiency throughout the 6 months, while those with limited resources did not. Also, educational context predicted learners’ proficiency gains over time, and low-proficiency learners exhibited more improvement than high-proficiency learners. The findings offer important implications to school administrators and testing agencies.
In most Western European countries, aspiring citizens must pass both a second language (L2) test and a knowledge of society (KoS) test, often administered in the L2. While several studies highlighted the social inequalities such tests may create, few have quantified these disparities. This study, anchored in the IMPECT project, analyses the test results of 79,794 migrants in Norway to explore the implications of these requirements, particularly for adult L2 learners with limited formal education in their first language (L1). Logistic regression analyses reveal that (1) learners with limited formal L1 education have significantly lower probabilities of passing the tests compared to those with higher education; (2) disparities are greater for KoS tests than for the oral A2 language test; (3) L2 scores strongly predict the likelihood of passing KoS tests, substantiating the assumption that KoS tests are implicit language tests; and (4) the impact of language skills on passing KoS tests varies by schooling level, with L2 scores being a stronger predictor for learners with limited education than for those with tertiary degrees. These findings highlight barriers faced by learners with limited L1 schooling in the citizenship acquisition process and underscore the need for targeted policy and support mechanisms.
Previous research has established the essential role of aural lexical knowledge (ALK) in second language (L2) listening comprehension; however, from a measurement perspective, it remains underexplored how best to measure ALK and which word frequency bands show the strongest predictive power for L2 listening comprehension. This study examines the effectiveness of three item formats in an aural vocabulary test (Yes/No, first language meaning-recall, or meaning-recognition) in predicting variance in L2 listening comprehension among 176 English as a Foreign Language (EFL) learners. Following a vocabulary levels test (VLT) format, 300 words representing the first 5000 flemmas of a television and movie transcript-based word frequency list were tested across the three item formats. The results indicated that the meaning-recall VLT was significantly more difficult than the Yes/No and meaning-recognition formats. Overall scores for all three tests significantly correlated with scores on a standardized L2 English listening comprehension test (TOEIC®), with the strongest correlations observed for the meaning-recall (
This study presents a methodological framework for applying corpus linguistics to systematically evaluate the authenticity of chatbot production in relation to (spoken) production in a general target language use domain. We demonstrate the approach through data drawn from the development cycle of a low-stakes formative assessment system in which learners interact with a ChatGPT-powered bot. A Chatbot Corpus containing approx. 290,000 words from 600 simulations of target ChatGPT production was created, representing two GPT versions (3.5 and 4), and three temperature settings. This corpus was then compared with relevant subcorpora in the British National Corpus 2014, which contains 100 million words of British English collected in naturalistic settings. Analyses were conducted at macro- (multi-dimensional analysis), meso- (comparative frequency analysis), and micro-levels (occurrence of specific pragmatic feature analysis). Results showed that the ChatGPT-powered chatbot production was systematically more similar to genres of written rather than spoken communication: output demonstrated higher lexical density and was characterised by a relatively low occurrence of features typical of spoken communication such as stance and pragmatic markers. We argue that the methodological framework is applicable across different chatbot models, allowing researchers and developers to use this approach with newer, more refined AI-powered conversational agents in the future.
Language testing is both an evaluative practice and a commercial enterprise shaped by market forces. Within this context, test developers have a responsibility to ensure transparency with test users, particularly when scores inform high-stakes decisions. This Viewpoint contrasts transparency in communicating the truth about language tests with the persuasiveness of argument-based validity, noting that persuasiveness, although central to such arguments, is not equivalent to transparency. Two measures are proposed to strengthen transparency. First, test developers should publish test-form-specific validity reports detailing content, psychometric properties, and interpretation guidelines for each form. Second, they should clearly explain a test’s limitations to the public, especially when scores are used in high-stakes settings, such as immigration or university admission, without adequate validation. The latter measure draws on regulatory norms in the pharmaceutical industry, where transparency can protect consumers from potential misuse. Specific steps are outlined to support these measures. Overall, these proposals aim to shift the emphasis from justification and persuasion toward transparency and align language testing practices more closely with open science principles.

When different tests are used for the same purpose, score requirements should be comparable so that examinees cannot obtain an unfair advantage simply because of the test they chose. Drawing on our experience conducting a large-scale concordance study to allow for an empirical comparison of IELTS Academic and TOEFL iBT test scores, we review Knoch and Fan’s evaluative framework, explore methodological best practices and challenges, and offer future directions for score concordance research. We emphasize the importance of methodological rigor in collecting test-taker score data, transparency in analyzing such data to build score concordance tables, and a reasonable degree of construct comparability as a prerequisite for conducting a score concordance study, while also highlighting the limitations of concordance tables as standalone tools for admissions decisions. We note that some aspects of Knoch and Fan’s good practice principles are more straightforward to implement in practice than others. The good practice principles could be updated or adjusted after real-world application, which we describe with a view to furthering best practice in concordance research. We conclude this viewpoint with recommendations for decision-making that are based on fair score requirements, irrespective of which test the examinees chose.

This study provides a bibliometric analysis of International English Language Testing System (IELTS) research from 1989 to 2024, incorporating 641 research documents. Among these, 482 were obtained from five online indices (Web of Science, Scopus, ERIC, EBSCOhost, Google Scholar) and 159 were identified manually (including 146 studies sponsored by the IELTS co-owners). The analysis focuses on patterns and trends in research topics, methodological approaches, publications and disciplinary areas, author-affiliated institutions and countries, and co-authorship. The results show that topics which were well-covered by researchers extended beyond psychometrics and validation of the testing system to include consequential factors, such as teaching and learning, academic/professional contexts, test preparation, and non-linguistic constructs (e.g., stakeholders’ attitudes, beliefs). Quantitative studies constituted the most common approach but were exceeded by (often well-cited) mixed-methods inquiries sponsored by the co-owners. Research authorship diversified from Anglophone to English as a Foreign Language (EFL)/English as a Second Language (ESL) contexts, reflecting a broader view of IELTS’s validity, warranting further exploration. Implications of the findings are discussed both for IELTS and language assessment research more generally.
This study investigates ChatGPT’s performance as an Automated Writing Evaluation (AWE) system by comparing its scoring with that of human raters and examining learners’ perceptions of its feedback. Six ChatGPT models were developed using different prompt configurations. Sixty English writing samples produced by Korean university English as a Foreign Language (EFL) learners were evaluated by two human raters and the six models. A multifaceted Rasch model, Spearman’s correlation, and intraclass correlation were used to examine reliability, severity, and bias. Learners’ perspectives on the models’ feedback were collected through open-ended surveys and analyzed thematically. The results indicate that prompt design plays a central role in shaping ChatGPT’s scoring behavior. Prompts combining Chain-of-Thought reasoning with Fill-in-the-blank scaffolding were associated with higher scoring consistency, while predefined personas and few-shot exemplars tended to moderate scoring severity. However, no stable patterns were observed for either bias or rating scale use, suggesting that prompt design alone cannot fully control domain-level bias. In particular, reasoning-intensive writing domains showed substantial divergence from human judgment, highlighting the need for human oversight. In parallel, learners generally viewed ChatGPT’s feedback positively, while also noting areas for improvement. Overall, the study demonstrates the potential of prompt-calibrated ChatGPT-based AWE as a supplementary tool for writing assessment and instruction.
Interactional competence (IC) is essential for oral communication assessment, yet human partner variability can introduce construct-irrelevant variance in paired speaking tests. As an alternative to a test with a human interlocutor, this study describes the development of a large language model (LLM)-driven Spoken Dialogue System and compares GPT-4o to Claude 3.5 Sonnet to inform model selection for IC assessment. Twelve international students completed paired discussion tasks with both LLMs in counterbalanced order. System performance was evaluated through breakdown activation consistency, stance maintenance, and persona adherence. Test-taker performances were analyzed using interactional discourse analysis to identify IC features across three dimensions: topic management, interactional management, and interactive listening. Semi-structured interviews explored test takers’ perceptions of the AI partners. Results showed Claude outperformed GPT-4o in eliciting IC features, successfully activating communication breakdown strategies and maintaining oppositional stance, thereby creating more opportunities for test takers to demonstrate key IC abilities. Test takers perceived Claude as more authentic and natural, while GPT was perceived as more artificial. These findings demonstrate that different LLMs create distinct interactional conditions affecting both IC elicitation and test-taker perceptions. The findings highlight the need for construct-driven evaluation criteria when selecting LLMs for language-assessment contexts.
The use of large language models (LLMs) in language assessment, particularly in spoken dialogue systems (SDSs) for assessing speaking, remains at an early stage. This study explored the use of an LLM-driven SDS to assess second language speaking ability. Using a within-participant design, we compared the paired discussion performance of 30 participants interacting with a self-built LLM-driven SDS (E-Talk) versus a human interlocutor, focusing on the interlocutor effect on test scores and fine-grained linguistic and interactional features. Results did not yield substantive differences in oral performance across the two conditions, although the interactions with the LLM-driven SDS displayed slightly lower scores for pronunciation and language use, slower speech rate, and reduced lexical diversity alongside increased syntactic complexity, and greater initiative in introducing new ideas, prompting responses, and guiding discussions toward negotiation, coupled with weaker interactive listening. These findings suggest that LLM-driven SDSs can serve as a potential, usable interlocutor in dialogic speaking assessments, eliciting key aspects of speaking ability. That said, further refinement is needed to better capture interactive listening and collaborative meaning construction, highlighting the importance of interpreting SDS-based performance relative to its specific interactional affordances.