Abstract
Second-language (L2) testing researchers have explored the relationship between speakers’ overall speaking ability, reflected by holistic scores, and the speakers’ performance on speaking subcomponents, reflected by analytic scores (e.g., McNamara, 1990; Sato, 2011). These research studies have advanced applied linguists’ understanding of how raters view the components of effective speaking skills, but the authors of the studies either used analytic composite scores, instead of true holistic ratings, or ran regression analyses with highly correlated subscores, which is problematic. To address these issues, 10 experienced ITA raters rated the speaking of 127 international teaching assistant (ITA) candidates using a four-component analytic rubric. In addition, holistic ratings were provided for the 127 test takers from a separate (earlier) scoring by two experienced ITA raters. The two types of scores differentiated examinees in similar ways. The variability observed in students’ holistic scores was reflected in their analytic scores. However, among the four analytic subscales, examinees’ scores on Lexical and Grammatical Competence had the greatest differentiating power. Its scores indicated with a high level of accuracy who passed the test and who did not. The paper discusses the components contributing to ITAs’ L2 oral speaking proficiency, and reviews pedagogical implications.
Since the 1980s, researchers have been studying issues related to international teaching assistants’ levels of oral English and how those levels impact the quality of undergraduate education (Bailey, 1983; Plakans, 1997; Tyler, 1992). To ensure that ITAs all possess a level of oral proficiency that is high enough for teaching, many universities in the United States provide ITA screening procedures and training programs (Choi, 2017; Ginther, 2003). Some universities have developed their own ITA screening tests internally (e.g., the University of California, Los Angeles, Purdue University, Michigan State University). One standardized test that is commonly used is the Speaking Proficiency English Assessment Kit (SPEAK), developed by Educational Testing Service (ETS). Although the SPEAK test is no longer supported by ETS, it is regularly revised by individual universities and administered locally to prospective ITAs. Michigan State University (MSU) currently uses a modified version of the SPEAK test called the MSU Speaking Test 1 (henceforth, the Speaking Test). Although the task types are similar to those used by the SPEAK test, the Speaking Test tasks are situated in academic contexts. Raters of the Speaking Test use a holistic rating scale, which was the rubric for the SPEAK test (Xi, 2005). The scale describes the typical profiles of the overall performance of examinees at different levels.
In this study, I investigate the use of an analytic rating rubric for the Speaking Test. The Speaking Test’s current holistic ratings are intended to be sufficient for administrators to decide whether a prospective ITA is proficient enough to teach undergraduate courses. The holistic descriptors, however, cannot by nature differentiate examinees with varied oral proficiency profiles that might exist within the holistic score levels. This observation has raised several issues related to score interpretation and score uses. First, the strengths and weaknesses of students’ speaking skills cannot be identified, and thus the question remains as to whether the students’ score profiles from the Speaking Test (which can be identified from their analytic scores) are aligned with the diagnostic feedback provided by instructors in the ITA support courses. 2 Second, it is difficult for ITA training programs to allocate available resources to best assist ITA candidates to improve their speaking without knowing which aspects of speaking play the most important role in ITA’s overall speaking performance assessment.
In an attempt to address the issues noted above, in this study, I explore the relationship between examinees’ overall speaking ability, as scored holistically, and their performance on four analytic rating criteria. In addition, I examine which analytic rating criterion has the greatest differentiating power among examines around the passing score threshold. The intention of this study is thus to contribute to a clearer understanding of the relationship between examinees’ analytic scoring profiles and their holistic evaluations.
Definitions of and comparisons between holistic and analytic rating scales
In second language (L2) testing, a distinction is often drawn between two types of rating scales for assessing language learners’ performances: holistic scales and analytic scales (Hamp-Lyons, 1991; Weigle, 2002). Comparisons between the two types of rating scales have been well documented by researchers (e.g., Bachman & Savignon, 1986; Knoch, 2009; Xi, 2007). Although there are still ongoing debates about which type of rating scale should be used in a given testing context, general guidelines suggest that holistic scoring is often employed in performance assessments where a test takers’ overall performance is more important than its particular features, whereas analytic scoring is used to separate relevant features of test takers’ performance, and to evaluate each subscale component independently (Taylor & Galaczi, 2011). In the ITA testing context, holistic scoring is often chosen because it can be undertaken relatively quickly. It promises efficiency in scoring and ease in score reporting (Xi, 2007). However, if an ITA assessment is intended to solely assess language skills, analytic scoring might be preferred as scoring is less likely to be influenced by construct-irrelevant factors such as teaching skills (Farnsworth, 2004).
Holistic
Holistic scoring is often preferred over analytic scoring in L2 speaking assessment when the speakers’ overall communicative effectiveness is of interest (Weir, 1990). On the other hand, holistic scoring also imposes some potential problems. First, in holistic ratings, raters may implicitly weight the features included in the rubric differently before arriving at their judgments. The difference may result from raters’ varied rating experiences and their perceptions of effective communicative skills. Second, holistic scores are not sufficient for teachers and practitioners to make meaningful score interpretations because a single holistic score cannot provide diagnostic information that may be useful for learners’ future learning (Alderson, 2005). Examinees receiving the same holistic scores do not necessarily share the same performance profiles. This argument is also in line with the finding of a study by Choi (2017) in which variance in test takers’ analytic scores was observed among those with the same test results, as determined by their analytically derived composite scores.
Analytic
Analytic scores are sometimes also referred to as subscale scores or subscores depending on the testing contexts. Analytic scoring allows raters to focus more narrowly on different aspects of performance, thus it contributes to higher rater agreement and rating reliability (Weir, 1990). In fields outside second language testing, such as educational measurement, there has been an increasing demand for analytic score reporting, which reflects a desire for diagnostic information, or a desire for more detailed information to facilitate placement or admission decisions (Haberman, 2008). However, some disadvantages of analytic scoring have been documented by researchers as well. Luoma (2002) argued that the task of analytic rating becomes burdensome (it imposes a high cognitive load) if raters are required to manage four or five criteria simultaneously and make multiple judgments accordingly. This causes potential rating inconsistencies. Analytic ratings have also often been found to be highly correlated not only among themselves but also with test takers’ holistic scores, thus rendering them psychometrically redundant (Bacha, 2001; Feinberg & Jurich, 2017; Lee et al., 2009; Xi, 2007). This observation is evidence of construct overlap. However, there are at least two arguments for the adoption of analytic scores over holistic scores: First, they provide a more accurate depiction of learners’ proficiency given the current multi-componential definition of language ability (Bachman et al., 1995); and second, academic institutions can use them to examine test takers’ performance-profiles to inform placement decisions and facilitate instructional effectiveness (Choi, 2017). The second point is especially relevant in some support courses offered by ITA programs. These support courses are devoted to different aspects of oral skills and aim to help ITA candidates function effectively as teaching assistants. Test candidates’ analytic test scores provide diagnostic information beyond that provided by holistic scores, and are therefore of more use to teachers of support courses.
Comparisons between holistic and analytic scoring of performance in L2 speaking assessments
Empirical research on the relationship between holistic 3 and analytic rating can be separated into different strands. In one line of research, researchers compared test takers’ holistic and analytic scores that are linked to an external reference framework such as the Common European Framework of Reference (Council of Europe, 2001). In these comparison studies (e.g., Barkaoui, 2011; Harsch & Martin, 2013; Khabbazbashi & Galaczi, 2020), researchers placed an emphasis on the measurement qualities (e.g., rater agreement, discrimination between candidates) of the rating methods (i.e., analytic rating versus holistic rating) rather than the relationship of test takers’ subscores on analytic criteria to an overall score. Although less common, another line of research touched upon this issue by exploring the relative importance of individual analytic rating scales in test takers’ global test performance (Choi, 2017; Farnsworth, 2004; McNamara, 1990; Sato, 2011; Xi, 2007). It is this strand of research that is of most relevance to the current study.
Research investigating holistic and analytic scoring of performance in L2 speaking assessments suggests that there seems to be no single answer to the question of “which aspects of speaking are most influential in examinees’ overall scoring?”. This is unsurprising considering that what is considered important in the overall evaluation of examinees’ performance is often specific to a given testing context and the needs of individual tests.
In the context of scoring validation in L2 speaking assessments, Xi (2007) compared the two scoring methods and investigated whether analytic scoring could provide additional information beyond what is provided by holistic scores for the Test of English as a Foreign Language Internet-based Test (TOEFL iBT) speaking. Results suggested that the use of analytic scales was not supported after an examination of the utility of analytic scoring in this large-scale holistically rated speaking test. Highly correlated scores on different analytic criteria, a great deal of overlap among different analytic dimensions that raters perceived, and test takers’ flat analytic scoring profiles led Xi to conclude that the holistic rubric would be sufficient given the testing context.
Unlike Xi (2007), in which overall communicative effectiveness was clearly specified in rubric descriptors, in McNamara (1990) and Sato (2011), raters were asked to rate candidates’ overall speaking skills without detailed information about which aspects of speaking were considered important in their evaluations. McNamara investigated the empirical relationship between test candidates’ holistic and analytic scores. Test candidates’ responses were collected from the speaking sub-test of the Occupational English Test and were rated on five analytical subscales in addition to being rated in terms of overall communicative effectiveness. McNamara found examinees’ scores on Resources of grammar and expression accounted for the most variance (R 2 = 67.8 – 69.5%) in their holistic ratings, highlighting the decisive role of grammar in the raters’ judgment on the overall scale. In Sato, to answer the research question as to which criterion is associated with general oral proficiency, 30 Japanese English learners’ responses were analytically and holistically rated. Findings suggested that the contribution of Content Elaboration Development made to overall test performance was the largest.
In a study conducted in the ITA context, Choi (2017) underscored an important feature of analytic scoring: the ability to purposefully weight different components of a speaking evaluation. He investigated the empirical scoring profile of the academic oral English proficiency of ITA candidates and identified seven distinct proficiency profiles based on 960 test takers’ scores on four analytic rating scales in an ITA oral screening test. The scale of Pronunciation was weighted more heavily (× 1.5) than Lexical Grammar, Rhetorical Organization, and Question Handling (× 1.0). Therefore, the test results favored test takers who scored highly on Pronunciation. This was particularly relevant for test candidates showing a non-flat analytic scoring profile. In another ITA study, Farnsworth (2004) investigated the relative influence of candidates’ teaching skills (assessed by an analytic rubric with five sub-scales) and language ability (assessed by an analytic rubric with four sub-scales) in holistic ratings of their performance (assessed by a holistic rubric). The results of the regression analysis suggested that both language ability and teaching skills were significant predictors of holistic scores. This might be because the holistic rating construct was defined as “the ability to work as a TA” instead of “sufficient oral language ability in English to work as a TA” (p. 56). The researcher argued that the effect of teaching skills needed to be adequately separated from the language ability scores in order for the holistic scores to be interpreted in a more meaningful and useful way.
The need for the current study
Although there have been some studies undertaken to determine the relative contribution that each analytic component made to raters’ assessment of speakers’ overall speaking ability, the studies described above have been constrained by a number of methodological limitations. First, in reporting the relative importance of individual analytic scales in test takers’ holistically rated performance, most studies often relied on multiple regression analysis. Even though the use of multiple regression is appropriate when the association between an outcome variable and several explanatory variables is of interest, speakers’ analytic ratings are often highly correlated with each other, so the multicollinearity inherently existing in these explanatory variables may render estimates for regression coefficients unreliable, and ultimately yield misleading results (Graham, 2003). Second, the research design employed in some studies did not allow for direct comparison between the two types of scoring methods. For example, Choi (2017) did not directly assess students’ overall speaking performance by giving them separate holistic scores in addition to their analytic scores. Instead, they obtained a composite score for each student by assigning weights to the analytic rating scales before combining them. The research design therefore arguably did not address the question of which analytic scale influenced examinees’ overall speaking performance the most, a judgement which is best captured using holistic ratings. Third, except for Xi (2007), the studies described above did not make direct comparisons between the two types of scoring. Therefore, the question remains largely unclear as to whether the use of analytic scoring is necessary, or in other words, if it provides information beyond what is provided by holistic scoring.
Taken together, in order to explore the relationship between holistic and analytic speaking test scores, and to address the issues noted above regarding the previous comparison studies, I designed a study employing separate holistic and analytic rating procedures. I also opted for statistical methods that are more appropriate in light of the features of the data. The research questions the study aims to answer were as follows:
Do analytic sub-scales scores discriminate test candidates in the same way as the holistic scores on the Speaking Test?
Which analytic subscale had the greatest differentiating power in examinees’ overall speaking performance on the Speaking Test?
Methodology
Participants
Test examinees
For this study, I obtained data from the English Language Center Testing Office at MSU. The data were drawn from 127 ITA candidates who took the Speaking Test at MSU between 2014 and 2016. The examinees were from diverse linguistic backgrounds, which included but were not limited to native speakers of Spanish, Chinese, Korean, Bengali, Arabic, and Hindi. Because the testing data was already de-identified, the exact number of examinees in each linguistic background group was not available. The requested Speaking Test data consisted of the recordings of the test candidates’ complete responses to 12 prompts in the test along with their holistic test scores given by official MSU Speaking Test raters. The responses were timed and the lengths of the responses varied from 45 seconds to 90 seconds depending on the prompt. To ensure that the test data used in this study was drawn from test takers representative of the actual examinee population of the Speaking Test, the distribution of holistic test scores was sampled to match the actual Speaking Test holistic score distribution. The holistic rating scale ranged from 20 to 60, with 10-point increments in between. 4 However, very few test takers were rated below 35, so the test scores obtained by the participants of the study ranged from 35 to 60. The demographic information about the test takers and their test scores are in Table 1.
Examinees’ information and holistic test scores on the Speaking Test.
The departments were grouped and renamed based on disciplines.
Raters
Twelve graduate students in an applied linguistics program were recruited to re-rate the test data using an analytic rubric. Information about the raters is shown in Table 2. Eleven of the raters were pursuing MAs in Teaching English to Speakers of Other Languages (TESOL), and one was a Ph.D. candidate in applied linguistics. Five of the raters were English native speakers, and the other raters were all highly proficient L2 English speakers. All raters were required to have at least one year of ESL or EFL (English as a foreign language) teaching experience by the time the study began to participate the study.
Raters’ background information (analytic ratings).
Materials
Test
The Speaking Test consists of 12 tasks, which are designed to elicit language functions such as apologizing, persuading, recommending, and giving and supporting opinions. The task types include discussions on topics of general interest, discussions on topics related to the duties of a teaching assistant, descriptions of information shown in a graph, and presentation of information from a revised assignment, announcement, or syllabus. The time allotted to each response typically ranges from 45 to 90 seconds depending on the type of task. Each test taker’s audio-recorded response to each task is rated holistically by the same two experienced ESL instructors. When there were discrepancies, a third rater may be called in to resolve them by providing a third independent rating. The final ratings are averaged across tasks and raters and rounded to the nearest five. These raters are carefully selected and trained by testing specialists from the English Language Center Testing Office at MSU. The (intraclass) correlation of the holistic ratings of the Speaking Test provided by two experienced ESL raters is approximately .80 (Director of the English Language Center Testing Office, personal communication, Oct 24, 2019), which suggests good rating reliability.
Holistic rating rubric
The holistic rating rubric for the Speaking Test was the rubric for the original SPEAK test (see Xi, 2005), which has been in operation since 2014. Raters utilized this rubric, ranging from 20 to 60 (20 = no effective communication, no evidence of ability to perform; 60 = communication almost always effective, task performed very competently) to evaluate an examinee’s overall task performance with respect to each task. The key features that raters need to consider while rating the responses include general speech comprehensibility and clarity, examinees’ command of vocabulary and grammar, and their ability to elaborate on ideas using details and examples. More detailed information about the raters’ holistic evaluation is available on the test website (see https://elc.msu.edu/tests/msu-speaking-test).
Analytic rating rubric
The analytic rubric used in this study was constructed through the following steps. The components included in the rubric were first drawn and modified from the rubrics for the TOEFL Speaking Test (see www.ets.org/s/toefl/pdf/toefl_speaking_rubrics.pdf) and the Test of Oral Proficiency (TOP; see https://www.teaching.ucla.edu/top/scoring), which is an ITA screening test developed and administered at the University of California, Los Angeles (UCLA). The rationale for the choice of the two rubrics was twofold: (1) Both tests are used as an ITA screening device; 5 (2) the key features included in the holistic rubric for the Speaking Test are adequately addressed in these two analytic rubrics. I then revised the rubric for clarity and specificity through communications with testing experts and colleagues, a consultation with the director of the English Language Center Testing Office at MSU, and a pilot rating session with three colleagues. The finalized analytic rating rubric has four subscales: Phonetic and Phonological Competence, Lexical and Grammatical Competence, Rhetorical Organization, and Topic Development. All subscale scores were on the same four-point scale ranging from one to four (see Appendix A).
Procedures
I gave all 12 raters a two-hour rater training session and a one-hour practice rating session where they learned how to use the analytic rubric for the Speaking Test. Specifically, in the rater training session, the raters read through the scale descriptors at each score level, and they then listened to the corresponding benchmark speech sample at each score level. In the subsequent online practice rating session, the raters individually rated 24 speech responses using the analytic rating scale. I computed the Intraclass Correlation Coefficient (ICC) of the practice ratings to assess the rater agreement on the four scales. The ICC exceeded .90 for all subscales except Rhetorical Organization (ICC Pronun = .92, ICC Lex Gram = .90, ICC Rhe Org = .87, ICC Topic Dev = .92), suggesting overall good rating quality. After the training and practice session, the raters individually rated audio-only speech responses from 127 test takers using the analytic rating scale on an online rating portal that I set up. Each time, the raters rated the speech responses to the same prompt by 127 test takers in random order. All of the raters completed the whole rating session within six weeks. I uploaded this study’s data and a code book with variable definitions to the Inter-university Consortium for Political and Social Research. See Ma (2021) to download the data and code book.
Data analysis
There were 12 task prompts in the Speaking Test, and the analytic scores were from 12 raters on four subscales. Therefore, each examinee received a total of 576 (12 raters × 12 tasks × 4 subscales) scores, which resulted in 73,152 (127 examinees × 576 scores) data points. I statistically analyzed these data points using Rasch measurement, correlation, cluster analysis, and one-way ANOVA.
In order to obtain a composite analytic score for each participant based on the analytic scores on the four subscales, I first used a many-facet Rasch analysis, in which (a) examinee ability, (b) task difficulty, (c) rater severity, and (d) rating scale measure were coded as four facets. In addition to examinees’ composite analytic scores, their scores on four subscales were generated based on ratings on 12 tasks from different raters. Initially, the rater infit and outfit statistics indicated the extent to which the raters’ ratings matched those expected from the Rasch model, and this information allowed me to examine the quality of the rating scores. Outfit is calculated using all the data, and therefore it is more sensitive to the influence of outlying ratings, whereas infit statistics are an information-weighted indicator of misfit, calculated using trimmed data, with the outliers removed (Knoch & McNamara, 2015). I applied Bond and Fox’s (2007) standard for accurate raters, which is that raters applying criteria inconsistently will have infit and outfit mean square indices greater than 1.30, and that raters overusing certain score levels (and thus rating more predictably than the model expects) will have infit and outfit mean square indices less than 0.7. I identified two raters with abnormal rating patterns after conducting the multi-facet Rasch analysis with ratings from all 12 raters (see Table 3). As shown in Table 3, the infit and outfit indices indicated that the ratings given by the two raters were either too predictable (less than .7, overfit the model) or too unexpected (larger than 1.3, underfit the model). Therefore, I excluded these two raters and conducted a second round of multi-facet Rasch analysis with the data collected from the remaining 10 raters.
Rater fit statistics and SR/ROR correlations.
Note:
MNSQ = mean square; SR/ROR correlation = single rater-rest of the raters (SR/ROR) correlations.
Rater 2 and Rater 6 were removed from the second round of analysis.
Next, to answer the first research question, I performed Spearman correlation analyses to explore the relationships among the holistic scores of the Speaking Test and the four analytic subscale scores. The holistic scores reflect the ordinal ranked values given the nature of the scoring scale, a situation where the Spearman correlation is deemed more appropriate.
To answer the second research question, which is to examine whether the four analytic subscale scores discriminated the examinees in the same way as the holistic scores, I first employed a clustering approach to classify examinees into groups based on their analytic scores. The examinees belonging to the same cluster were considered as having similar analytic score patterns, although they may receive different test results as shown by their holistic scores. I subsequently scrutinized their holistic score patterns to investigate how similar or dissimilar they were. The cluster analysis allowed me to (a) examine which analytic scale(s) differentiated participants around the passing score threshold, and (b) attempt to explain the within-cluster holistic score differences (i.e., why some examinees showing similar analytic score patterns ended up with different holistic scores).
I made the decision of how many clusters to have by calculating the average silhouette width, a cluster quality index measuring how similar an observation is to its assigned cluster (cohesion) compared to other clusters (separation). The silhouette value ranges from −1 to +1, where −1 indicates that most observations are better fit in neighboring cluster, and +1 indicates that most observations are well matched to the assigned cluster.
I conducted the cluster analysis in two steps. First, I pulled out a subset of data from examinees who received holistic scores of 40, 45, and 50 on the Speaking Test. Second, I performed k-means cluster analysis (distance measure: Euclidean distances) with these examinees using their four analytic subscale scores generated from the previous Rasch analysis. My motive to focus on this group of examinees is twofold. First, the cut score for the Speaking Test is usually 50, but an examinee who obtains a score of 40 or 45 can have the hiring department request an appeal, which can lead to success on the overall exam. In an appeal, a Review Board evaluates the examinee independently, and if the Board ascertains that the candidate would be able to perform TA duties specific to the course in question, the student may be given a teaching assignment. Oftentimes, certain restrictions are placed on the type of assignment that the student is eligible for (e.g., one-on-one instruction in a chemistry lab with no dangerous materials, or in a classroom setting with no new material being introduced, as opposed to having the TA teach a stand-alone course). Therefore, an examinee who is rated 40 could potentially be comparable to one rated 45 or even 50 in certain aspects of speaking ability, but may not be evaluated as strong as them when rated holistically. Second, most examinees fall within this score range, as can be observed in Table 1. Therefore, differentiating these groups of examinees can have a meaningful and crucial impact on the great majority of examinees of the Speaking Test. To determine the best value of the number of clusters (the value of k), I calculated the average silhouette width of different numbers of cluster solutions and selected the one generating the highest silhouette value.
Results
Reliability
I used Cronbach’s α to estimate the internal consistency of raters’ use of each analytic scale across tasks. Table 4 contains the reliability coefficients of the four analytic ratings for each rater. The reliability coefficients of four subscales all fell within an acceptable range (above .70), indicating that the way the 10 raters assessed speakers’ speaking skills was internally consistent. It was also worth noting that Phonetic and Phonological Competence and Lexical and Grammatical Competence ratings were comparatively more consistent across tasks than Rhetorical Organization and Topic Development. The lower reliability of Rhetorical Organization and Topic Development might be related to the lower amount of variation observed in the two scales (see Table 6).
Reliability (Cronbach’s α) of analytic scores.
Note: Pronun = Phonetic and phonological competence; Lex Gram = Lexical and grammatical competence; Rhe Org = Rhetorical organization; Topic Dev = Topic development.
Rasch analysis
I conducted a many-facet Rasch analysis to compute test takers’ analytic scores. The analytic subscale measurement report is shown in Table 5 (see Appendix B and Appendix C for more detailed statistics summary results for tasks and raters). The measure of each subscale indicates that among the four analytic subscales, Phonetic and Phonological Competence was relatively easy (−0.61 logits), whereas Topic Development and Rhetorical Organization showed the highest difficulty (0.32 and 0.31 logits respectively). The infit and outfit mean square indices of all four subscales were above 0.7 and below 1.3, demonstrating that all the analytic subscales had independent but not totally unexpected scoring patterns. The reliability of separation index reported in the subscale measurement report is 1.0, suggesting that the four subscale difficulties were widely separated and different.
Analytical subscale measurement report.
Note: SE = standard error; MNSQ = mean square; Estimated task discr. = Estimated task discrimination.
Pronun = Phonetic and phonological competence; Lex Gram = Lexical and grammatical competence; Rhe Org = Rhetorical organization; Topic Dev = Topic development.
Correlation analysis
Table 6 contains descriptive statistics of the scores and the coefficients of the Spearman correlations between examinees’ holistic scores on the Speaking Test, the four analytic subscale scores, and composite scores from analytic scores. The composite scores indicated examinees’ speaking performance assessed by the analytic rubric. Examinees’ scores on Lexical and Grammatical Competence had the strongest correlations with holistic scores (rho = .85). In other words, examinees who were rated high on Lexical and Grammatical Competence were primarily high on the holistic scale and vice versa. In addition, the relationship between examinees’ holistic scores and their composite scores was positive and strong (rho = .79), indicating that high performing examinees when assessed using the holistic scale were also ranked among the top when they were assessed using the analytic rating scales. This positive, strong relationship between holistic scores and composite analytic scores can also be observed in Figure 1’s scatterplot. Another noteworthy observation from Figure 1 is the varying widths of the analytic composite-score range obtained by examinees with the same holistic score. More specifically, high performing examinees (those who received 55 and 60 on the holistic scale) and low performing examinees (those who received 35 on the holistic scale) obtained analytic composite scores within a rather narrow range (Score Range35 = 1.13; Score Range55 = 1.56; Score Range60 = 0.82). The examinees who obtained a holistic score of 40, 45, or 50 had larger analytic-composite-score ranges (Score Range40 = 1.81; Score Range45 = 2.17; Score Range50 = 2.17).
Descriptive statistics and correlations between holistic scores, Rasch composite scores, and Rasch scores on four subscales.
Note: Hol Scores = Holistic scores; Rasch Comp = Rasch composite logit scores; Pronun = Rasch logit scores on phonetic and phonological competence; Lex Gram = Rasch logit scores on lexical and grammatical competence; Rhe Org = Rasch logit scores on rhetorical organization; Topic Dev = Rasch logit scores on topic development.

Relationship between holistic scores on Speaking Test and Rasch analytic composite scores.
Figure 1 also shows that there was a great overlap in the analytic composite score range among the examinees who received a score of 40, 45, and 50. These three groups of examinees shared the same analytic composite score range on the logit scale between 0.73 and 1.45. In addition, the widths of the composite score range shared by the 40 and 45 groups and by the 45 and 50 groups are 1.64 and 1.25 respectively, revealing that distinct test results could be obtained by these students despite their similar analytic composite scores. This finding also suggests a need to take a closer look at this group of examinees (with holistic score of 40, 45, and 50) and investigate which aspect of speaking ability plays a more important role in raters’ perceptions of examinees’ overall speaking ability relative to other aspects.
Cluster analysis
I identified the optimal number of clusters by calculating the average silhouette width of the solutions with different number of clusters. Figure 2 displays the relationship between cluster number and the average silhouette width associated with it. The three-cluster solution yields the highest silhouette value, indicating that most subjects were better-matched to the assigned cluster as compared to the neighboring clusters. Table 7 is the summary of descriptive statistics for the three-cluster solution. For ease of reference, the three clusters are labeled A, B, and C. Table 7 displays the number of examinees at each holistic score level assigned to each cluster, and this information is provided in Figure 3.

The relationship between Silhouette value and number of clusters.

Scatterplot displaying the three-cluster solution.
Summary of descriptive statistics for the three-cluster solution.
Note: Hol Score = Holistic score; Rasch Comp = Rasch composite logit score.
Cluster A comprised approximately 15% (16) of the sample. Except for one person, examinees who were classified into cluster A all obtained a holistic score of 50, and their analytic composite scores were concentrated around the top of the scale. Not surprisingly, the average of the holistic score of cluster A is the highest among the three clusters. Cluster B accounted for 43% (46) of the sample. The members of cluster B were mostly from the 45 and 50 groups, and their analytic composite scores were distributed in the middle of the scale. Cluster C, whose members accounted for the remaining 42% (44) of the sample, mostly consisted of examinees who received a holistic score of 40 or 45. The analytic composite scores obtained by those in Cluster C were spread out at the lower end of the scale, which resulted in the cluster with the lowest average composite score.
The scatterplots displaying the relationship between holistic scores and each analytic subscale score are shown in Figure 4. There are two layers of information contained in the Figure 4 scatterplots. First, direct comparisons can be made among examinees at different holistic score levels regarding their analytic scores. One can see a general analytic score pattern of examinees with the same holistic score. Second, the results of cluster analysis allowed me to make comparisons between examinees assigned to the same cluster but receiving different holistic scores. This would (a) enhance my understanding as to why these examinees obtained different test results despite their similar analytic composite scores, and (b) enable me to explore which aspect of speaking was valued more in the assessment of examinees’ overall speaking performance.

The relationship between holistic scores and analytic scores.
In terms of the general score pattern of each holistic score group, similar to what was seen in Figure 1, for each analytic scale, the range of analytic composite scores associated with each holistic score largely overlapped. The only exception was in Lexical and Grammatical Competence between the group of 40 and 50, in which almost all the examinees (30 out of 33) receiving a holistic score of 50 scored above the average (1.3) on this subscale, whereas the examinees in the group of 40 mostly scored below the average on this subscale.
The characteristics of analytic score patterns are consistent with the pattern observed in Figure 3 across the three clusters: the four analytic scores obtained by examinees in cluster A were concentrated at the upper end of the scale, cluster B in the middle, and cluster C at the lower end. One characteristic feature that sets examinees in cluster A apart was their distinctively high scores on Phonetic and Phonological Competence. In addition, the composition of cluster A is more homogeneous compared to other clusters: only one member did not get a holistic score of 50. The analytic score patterns of cluster B and C are more complex. Differences seemed to exist in these four analytic scores among the holistic groups within one cluster (see Table 8). To explore whether the difference was meaningful and significant, I performed a one-way ANOVA analysis with examinees in cluster B and C; the results are shown in Table 9.
Cluster number crosstabulation for the three clusters.
Note: Hol Scores = Holistic scores; Rasch Comp = Rasch composite logit scores; Pronun = Rasch logit scores on phonetic and phonological competence; Lex Gram = Rasch logit scores on lexical and grammatical competence; Rhe Org = Rasch logit scores on rhetorical organization; Topic Dev = Rasch logit scores on topic development.
Summary of the one-way ANOVA.
Note: Pronun = Rasch logit scores on phonetic and phonological competence; Lex Gram = Rasch logit scores on lexical and grammatical competence; Rhe Org = Rasch logit scores on rhetorical organization; Topic Dev = Rasch logit scores on topic development.
The scores on Lexical and Grammatical Competence obtained by examinees in both cluster B and C with different holistic scores were significantly different (see Table 9), and the effect sizes were large (cluster B: p < .001, η2 = .30; cluster C: p < .001, η2 = .40), as I determined by adopting Cohen’s benchmark for interpreting η2 (Cohen, 1988, p. 283). In addition, for cluster C, the difference in Rhetorical Organization scores and Topic Development scores among examinees in the three holistic groups were significant at .05 level. Therefore, I performed a post-hoc analysis using a Bonferroni adjustment to locate the source of the difference (see Table 10).
Results for post-hoc tests.
Note: NS = Not significant; Pronun = Rasch logit scores on phonetic and phonological competence; Lex Gram = Rasch logit scores on lexical and grammatical competence; Rhe Org = Rasch logit scores on rhetorical organization; Topic Dev = Rasch logit scores on topic development.
In cluster B, examinees who received a holistic score of 50 scored significantly higher on Lexical and Grammatical Competence than examinees in the other groups, and the mean differences were large (50–45: p < .001, d = 1.26; 50–40: p = .03, d = 1.64) according to Plonsky’s (2015) guidelines for interpreting effect sizes. I found similar results in cluster C, but significant differences existed among all three groups (50–45: p = .01, d = 1.64; 50–40: p < .001, d = 2.78; 45–40: p < .01, d = 0.99). In cluster C, examinees in the 45 group scored significantly higher on Rhetorical Organization (p < .01, d = 1) and Topic Development (p = .01, d = 0.9) than those in the 40 group.
Discussion
My goal for this study was (a) to explore the empirical relationship between ITA candidates’ analytic score profiles with their overall speaking performances, and (b) to examine which analytic scale differentiates examinees around the passing threshold better. An investigation of the importance of each analytic subscale’s relation to the holistic score across examinees at different performance levels promotes a better understanding of the underlying components of effective speaking skills that may otherwise be obscured by holistic ratings. More specifically, the use of analytic scores unveils the relative importance of each underlying aspect of speaking to speakers’ overall speaking ability.
The high Spearman correlations between examinees’ holistic scores and their four analytic subscale scores indicate that the two types of scores ranked these examinees similarly to varying degrees. This is not surprising, given that oral language proficiency is a multi-dimensional construct (Bachman & Savignon, 1986), and L2 learners, who have been exposed to the target language in different learning contexts (as a foreign language or as a second language) or have different language learning experience or goals, may progress at personalized rates on different aspects of speaking ability. The coefficient of the relationship between the holistic scores and the analytic composite scores computed by Rasch analysis is .79, suggesting that about 62.4% of variance found in examinees’ holistic scores could be explained by the variance of their composite scores. The strong relationship between holistic and analytic composite scores suggests they are largely comparable, however, they are not interchangeable. In addition, the four analytic scores were highly correlated with each other, as often found in similar studies (e.g., Xi, 2007), which lends support to not using multiple regression to analyze the data. Worthy of mention is the two, highly correlated criteria Rhetorical Organization and Topic Development (rho = .96). This high Spearman correlation may indicate that raters might have problems differentiating these two subscales.
Among the four analytic components, Lexical and Grammatical Competence had the highest correlation with raters’ evaluation of examinees’ overall speaking ability. This result echoes the finding of the second research question: Significant score differences on this subscale were between examinees who received a holistic score of 50 (i.e., those who passed the test) and those who received a holistic score of 40 or 45 (i.e., those who did not pass the test). This result suggests that Lexical and Grammatical Competence appeared to be a differentiating component for passing the test. In addition, I closely examined the relationship between examinees’ analytic score patterns and their test results and found that almost all examinees who obtained a holistic score of 50 received an above-average rating on Lexical and Grammatical Competence. This indicates that, unlike other components included in the analytic rating scales, poor Lexical and Grammatical Competence performance could not be made up by high scores on the other subscales. However, it is worth noting that given the strong intercorrelations among the analytic scores, the importance of other analytic scales should not be overlooked.
The findings were in line with previous ITA research, in which the importance of ITAs’ grammatical competence to their effectiveness of teaching was outlined. For example, Ard (1989) suggested that ITA’s low competence in grammatical skills could lead to communication breakdowns when talking to undergraduate students. Salomone (1998) pointed out the importance of ITAs’ adequate vocabulary and grammatical knowledge for effective teaching by reporting how ITAs’ concerns about their English vocabulary and grammatical knowledge influenced their teaching effectiveness.
The importance of Lexical and Grammatical Competence also pertains to the instructions offered in ITA support courses. The current two courses offered for ITA candidates who do not pass the Speaking Test aim to help them improve their oral language skills with special emphasis on pronunciation and the use of English in the context of American university classrooms, respectively. 6 Although accurate vocabulary use and grammar are also listed as learning goals for one course (see https://elc.msu.edu/programs/ita/ita-course-offerings), the result of the study suggests that it is meaningful and rewarding for language instructors to devote more time and effort to these two aspects of speaking, especially for those students who were identified as low performing on Lexical and Grammatical Competence based on their analytic scores.
In addition, I found a significant difference in analytic scores of Rhetorical Organization and Topic Development between students who obtained a holistic score of 40 and 45. Previous ITA studies addressed the importance of successful discourse characteristics (e.g., Allison & Tauroza, 1995; Flowerdew & Tauroza, 1995; Rounds, 1987) and communicative strategies (e.g., Sato, 2011) to L2 communicative competence. However, in the current study, the score differences regarding these two analytic components did not impact whether an examinee would pass the test or not, because no significant score differences were found when contrasted with examinees who received a score of 50.
Phonetic and Phonological Competence did not differentiate among the examinees of interest, and this is an unexpected finding because pronunciation was commonly addressed as important in previous studies focusing on ITAs’ speaking ability (e.g., Choi, 2017; Hahn, 2004; Isaacs, 2008; Kang, 2010; Kang et al., 2010; McGregor, 2007; Schmidgall, 2013), and ITAs’ pronunciation skills have often been found as highly correlated with their teaching effectiveness (e.g., Pickering, 2001). This discrepancy in the findings may be accounted for by the different research design I adopted in the current study. First, the discrepancy may be owing to the holistic scoring method employed in the current study, which was different from generating analytically derived composite scores as done in some previous studies. For example, in Choi (2017), the calculation of the final scores favors test candidates who scored highly on pronunciation, owing to the higher weight assigned to this criterion. Second, whether raters had ESL teaching and rating experience or not may also be relevant to what features that raters attended to the most in the speech of ITAs. In the current study, although with varying years of experience, the raters of both holistic and analytic scoring were ESL teachers. The result that they attended to speakers’ Lexical and Grammatical Competence the most was aligned with the results of the study by Hsieh (2011), who found that the ESL teacher raters appeared to comment more frequently on the accuracy and complexity of ITA candidates’ speech, whereas undergraduate raters tended to provide more comments on speakers’ pronunciation and accent. Similarly, Schmidgall (2013), who recruited listeners from the undergraduate population (with no ESL teaching experience), found that pronunciation scores were comparatively better predictors of undergraduates’ judgment of communicative effectiveness than other analytic scales, such as lexical grammar, rhetorical organization, and question handling. Isaacs (2008), who highlighted the important role of pronunciation in the assessment of ITA, also recruited untrained undergraduates as raters.
Limitations and future research directions
Despite the attempt to address some methodological limitations found in previous studies, the current study is not perfect in its design. Although about 62.4% of score variance found in holistic scores were explained by the analytic components included in the rating scales, there was still about 37.6% of score variance that could not be attributed to the components included in the analytic rating scales used in the study. McNamara (1990) and Sato (2011) both included fluency as part of their analytic rating scales. It was not included in the current study because, as noted before, I modified the analytic rating scales from the holistic rating rubric for the Speaking Test along with the two existing rating scales, which are being used in testing contexts more relevant to the current study. Future researchers could explore the influence of fluency on raters’ evaluation of ITAs’ speaking ability by adding fluency to their analytic rating scales. On a related note, the discrepancy regarding the score variance between the two scoring methods indicates perhaps a more interesting line of thought: Are raters using the holistic scales drawing on elements of the speaking construct that are not captured in the available descriptors? This question could be examined by collecting and analyzing raters’ eye-tracking or think-aloud data during the rating process. The results would provide insights into raters’ cognition for researchers to understand what raters are and are not attending to in examinees’ performances.
Another issue with the current study pertains to the analytic rating process. The results show that examinees’ four analytic scores were highly inter-correlated, as is often the case with analytic rating (Bacha, 2001; Feinberg & Jurich, 2017; Lee et al., 2009; Xi, 2007; Xi & Mollaun, 2006). It is not clear whether these high inter-correlations resulted from examinees’ L2 speaking profiles (i.e., varying proficiency was found on different aspects of speaking skills but only to a limited extent), or from raters’ evaluations of four analytic components biased toward certain salient features of an examinee’s speech, thereby generating these strong associations among the scores. Multiple raters and multiple speech responses could alleviate this issue to some extent, but cannot eliminate them completely. Xi (2007) addressed this issue by performing a different rating procedure. In her study, raters rated each examinee’s task on one analytic subscale at a time with task and examinee order randomized. With the number of examinees’ speech responses being the same, to keep the rating process within the same time frame (given that the rating process of the current study was already long), the number of raters was decreased. Future research could explore whether this trade-off is necessary and meaningful.
Conclusion
In this study, by the use of many-facet Rasch analysis, I assigned four subscale scores to each examinee based on the ratings of his or her oral responses to 12 speaking tasks from 10 raters using an analytic rating rubric. I investigated the empirical relationship between students’ holistic and analytic scores by comparing 127 examinees’ analytic scores to their holistic scores. The results suggest that the two types of rating scales differentiated the examinees regarding their oral language proficiency in similar ways: Examinees’ competencies on different aspects of their underlying speaking skill were mostly in line with their overall speaking performance. In addition, the use of several statistical methods allowed me to explore which component of the analytic scales plays a more important role in raters’ evaluations of speakers’ overall speaking ability. Overall, I found that although higher scores in all of the analytic dimensions were necessary for test takers to receive a higher rating, the scores on Lexical and Grammatical Competence had slightly better differentiating power among examinees who passed the test and those who did not, providing empirical evidence that compared to others, examinees who were perceived as more competent L2 speakers demonstrated greater vocabulary knowledge and accurate grammar usage. In addition, although not to the extent as was seen in previous ITA research, examinees’ pronunciation skills were highly correlated with their holistic scores.
In conclusion, despite the various disadvantages in terms of practicality, the analytic rating scales used in this study demonstrated usefulness in providing information beyond the holistic scales. First, the diagnostic information provided by examinees’ analytic scores could be used to help ITA program coordinator make informed decisions about appropriate ITA support courses for failing examinees to take and/or aspects of speaking they should improve; second, such information can also be used to guide ITA training programs to better allocate available resources to improve program effectiveness.
Footnotes
Appendix
Rater summary statistics.
| Rater | Observed raw score average | Fair average | Difficulty measure (in logits) | SE | Infit mean square | Outfit mean square | Estimated task discr. | Point measure r |
|---|---|---|---|---|---|---|---|---|
| 1 | 2.94 | 2.96 | −0.21 | 0.02 | 0.72 | 0.73 | 1.32 | 0.53 |
| 3 | 2.98 | 3 | −0.11 | 0.02 | 1.18 | 1.18 | 0.81 | 0.52 |
| 4 | 3.15 | 3.17 | 0.33 | 0.02 | 1.14 | 1.17 | 0.81 | 0.33 |
| 5 | 3.03 | 3.05 | 0.01 | 0.02 | 1.01 | 1.03 | 0.98 | 0.47 |
| 7 | 3.37 | 3.4 | 0.95 | 0.02 | 1.12 | 1.17 | 0.83 | 0.36 |
| 8 | 2.86 | 2.88 | −0.4 | 0.02 | 0.88 | 0.89 | 1.14 | 0.56 |
| 9 | 2.93 | 2.95 | −0.24 | 0.02 | 1.18 | 1.18 | 0.79 | 0.44 |
| 10 | 3.13 | 3.16 | 0.29 | 0.02 | 1.3 | 1.26 | 0.71 | 0.61 |
| 11 | 3.38 | 3.42 | 1 | 0.02 | 0.74 | 0.76 | 1.31 | 0.56 |
| 12 | 2.34 | 2.35 | −1.63 | 0.02 | 0.75 | 0.75 | 1.29 | 0.55 |
Acknowledgements
I wish to thank Dr. Paula Winke and Dr. Dan Reed for their valuable feedback on an earlier draft of the article. I also would like to express special thanks to the anonymous reviewers for their suggestions, which helped in improving the manuscript. I am responsible for any errors or shortcomings that remain. This study was made possible by the support of the English Language Center Testing Office at Michigan State University, which provided access to test score data and test materials.
Declaration of conflicting interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/ or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
