Abstract
Rating scale development in the field of language assessment is often considered in dichotomous ways: It is assumed to be guided either by expert intuition or by drawing on performance data. Even though quite a few authors have argued that rating scale development is rarely so easily classifiable, this dyadic view has dominated language testing research for over a decade. In this paper we refine the dominant model of rating scale development by drawing on a corpus of 36 studies identified in a systematic review. We present a model showing the different sources of scale construct in the corpus. In the discussion, we argue that rating scale designers, just like test developers more broadly, need to start by determining the purpose of the test, the relevant policies that guide test development and score use, and the intended score use when considering the design choices available to them. These include considering the impact of such sources on the generalizability of the scores, the precision of the post-test predictions that can be made about test takers’ future performances and scoring reliability. The most important contributions of the model are that it gives rating scale developers a framework to consider prior to starting scale development and validation activities.
Keywords
Scoring of written and spoken responses is not only central to practical work in language assessment but also plays a key part in conceptualizations of validity (see, e.g., Kane, 2013; Knoch & Chapelle, 2018). This is because scoring provides a crucial link between the performance and a proficiency claim. Although best practice in test development has been described in much detail (see, e.g., American Educational Research Association [AERA] et al., 2014; International Language Testing Association [ILTA], 2007), proposed processes for rating scale development have received much less attention. In language assessment more specifically, a small number of researchers have proposed possible approaches and models of scale development (further described below), but these models may not be representative of actual test development processes and have also not been clearly linked to the broader test development literature. With this study we set out with two specific aims in mind: (1) to model, by drawing on a systemic review, the actual scale development processes used by researchers and practitioners in the field of language assessment; and (2) to propose how such a model could be integrated with the wider test development processes proposed in influential publications such as the Standards for Educational and Psychological Testing (AERA et al., 2014), henceforth Standards.
Literature review
Throughout the history of language testing, theories and conceptions have continually been replaced, refuted, and revised, but the centrality of the rating process in rater-mediated assessment remains unchallenged. Also in the current era of argument-based validity, scoring continues to occupy a central position (Knoch & Chapelle, 2018), and for good reason: the act of rating links a performance to a proficiency claim, making it a vital component of any validity argument (Kane, 2013; Kane et al., 2017). In rater-mediated performance assessment, rating scales form a key part of that link between a performance and a claim (or score) (Kane et al., 2017).
The Standards sets out clear guidance and standards for test development and validation. The manual advocates, for example, that test developers “be clear on the construct(s) to be measured, including the target of the assessment, the purpose for which scores will be used, the inferences that will be made from the scores, the characteristics of examinees and subgroups of the intended test population that could influence access” (AERA et al., 2014, p. 5). The Standards further specifies that the “specifications for such tests should include descriptions of the outcomes the test is designed to predict and plans to collect evidence of the effectiveness of test scores in predicting these outcomes” (p. 75). The importance of scoring criteria to the possible interpretation of test scores is also mentioned in the following extract from page 78: “Further, both theoretical and empirical evidence are important for documenting the extent to which performance assessments —tasks as well as scoring criteria—reflect the processes or skills that are specified by the domain definition” (italics in original).
Although the Standards and other publications setting out processes of test development and validation point to the importance of the scoring criteria, which are often also described as the de facto test construct (McNamara, 2001), there is little information available on how rating scales (also referred to as scoring criteria, or scoring rubrics) are developed, what sources may influence the way the rating scale reflects the wider test construct, and how scale development can be integrated with other test development processes. This is surprising, considering the centrality of rating scales in the scoring process, and the direct effect the design of such scales can have on score generalizability (how broadly the score can be generalized beyond the test-taking situation), the precision of the predictions that can be made about test takers, as well as on the reliability of the scores.
Kane et al. (2017) described how rating scale development typically involves “expert opinion based on research and other sources of information,” and how it can involve empirical performance data (p. 24). This description of rating scale construction as being an intuition-driven endeavor aligns with Fulcher’s dyadic perspective of rating scale development, which involves either expert opinion (measurement-driven) or a use of performance samples (or performance-driven; see below) (e.g., Fulcher et al., 2011). But whereas Kane stated an observation, Fulcher raises an objection. For Fulcher, rating scales that are based on expert opinion rather than on empirical performance data lack construct validity.
Fulcher (1987) has long objected to the fact that many rating scales lack an empirical foundation. His dyadic typology of rating scales hinged on how empirical performance data are used (or not used) in rating scale construction, and constituted a purposeful simplification of an earlier classification into three intuitive methods and three empirical methods (see Fulcher, 2003).
In Fulcher’s conception, measurement-driven scales (Fulcher et al., 2011) can originate from various sources, such as the intuition of the developers or of the experts they consult, or they can originate from approaches that re-scale existing descriptors, or from “can do” statements that are not themselves generated by discourse studies of actual performance. Based on this, the scale developers will construct a scale that feels appropriate to them. Often, this scale will be presented as a linear sequence of proficiency levels, with the assumption that each higher level subsumes lower-level language, an assumption that may have anecdotal challenges (see Isbell et al., 2019) that in contemporary second language acquisition (SLA) may need resolution. Fulcher and his colleagues argued that such scales may rely on theory indirectly, insofar as theory has shaped the developers’ beliefs, but they do not stem from an analysis of actual performance data. Analyzing performance data is “important for establishing context-approproiate interpretations and uses of test scores” (Isbell et al., p. 444). Examples of measurement-driven scales are found in tests that have rating criteria adopted or adapted from the Council of Europe Framework of References (CEFR) descriptors (Council of Europe, 2001), which have been criticized for lacking an empirical or theoretical foundation (Deygers et al., 2018; Hulstijn, 2007). Such criticisms exist regarding the American Council on the Teaching of Foreign Language (ACTFL, 2012) descriptors (see Isbell & Winke, 2019, pp. 472-474) as well.
Performance-driven scales, on the other hand, are constructed on the basis of real-world data and corpora (Fulcher, 2012; Fulcher et al., 2011). For this reason, Fulcher (1987, 1996b, 2012) argued, well-designed performance-driven scales will be more in line with real-world language use, and allow for the imperfections, hesitations and false starts that mark language use at any proficiency level. In his work on rating scales, Fulcher outlined various types of rating scales across the years, but he has consistently applied the measurement- versus performance-driven logic as a classification criterion (e.g., Fulcher et al., 2011). As one example of the development of performance-driven scales, Upshur and Turner (1995) reviewed eight student performances on a story-retelling task. In a first step, the researchers divided the performances impressionistically into four better and four poorer. They then formulated a question that allowed them to classify the performances into these two groups. They then worked their way further down, by formulating another question that could divide the top and bottom groups further into two groups. In this way, Upshur and Turner produced a decision-tree of questions that the raters could use to arrive at the final score. The two researchers produced two of these decision-tree scales (which Upshur and Turner referred to as empirically based, binary-choice, boundary-definition, or EBB scales), one for communicative effectiveness and one for grammatical accuracy.
In recent work, Fulcher’s objections to measurement-driven scales have been applied to test validation. In his earlier papers, he compared the rating criteria of communicative language tests to the characteristics of real-world communication, and reported important discrepancies (Fulcher, 1987, 1996b). Fulcher, for example, critically examined the concept of fluency in existing rating scales at the time, and was, based on a careful analysis of test taker spoken data, able to criticize the use of descriptions of fluency used. Such discrepancies, he argued, constituted a fundamental validity problem, since a test cannot maintain construct validity if its criteria are at odds with the construct it purports to measure. Translated to a Toulmin-based validity argument (Kane, 2013; Toulmin, 2003), Fulcher’s argument would be that the scoring inference cannot be upheld when the warrant (e.g., the rating scale used is valid and produces reliable ratings) does not offer the backing required to make the case that an observed performance can be translated to a score that has real-world meaning. He therefore argued for a stronger alignment between scale criteria and real-world language use.
Fulcher’s work indicated that rating scale construction must thus satisfy at least two criteria. First, rating scale design must rely on empirical performance data and descriptions of real-world language use rather than on intuition. Second, rating scales must be meaningful to raters and relate meaningfully to real-world language use. The former point has been described in the paragraph above, and exemplified by Fulcher’s work on fluency (Fulcher, 1996a). Regarding the latter, Fulcher argued that general measurement-driven descriptors may appear to facilitate generalizability, when in fact they may not be uniformly interpretable and may be too vague and decontextualized to apply to a real-world context (Fulcher, 2012, 2017). If that is the case, validation may be a futile endeavor, since the scores may be unreliable, meaningless, or both (Fulcher, 1996a). As such, Fulcher suggested that rating scale validation may only be possible when the test goals, the tasks, and the rating criteria are all specifically contextualized (Fulcher, 2017). If not, they may have the apparent benefit of generalizability, while lacking in precision or predictive strength.
Recognizing the tension between the predictive strength and generalizability resulting from the design of a rating scale, Bachman and Palmer (2010) differentiated between can do and ability tests. Ability tests (e.g., diagnostic tests) yield general language proficiency observations, typically within the context of language education, and therefore are more generalizable, but have weaker predictive strength. Can-do tests, on the other hand, have the purpose of making inferences regarding the linguistic performance within more specific target language use (TLU) domains (e.g., professional contexts) and therefore have lower generalizability and higher predictive strength. While Bachman and Palmer’s conceptualization does not completely match that of Fulcher, as Bachman and Palmer’s was not focused purely on rating scales, both theories stress how a test goal (e.g., determining overall ability, investigating workplace language proficiency) impacts the specificity and generalizability of the score-based inferences.
Fulcher’s work on rating scales has profoundly impacted language testing research. His conceptualization of measurement-driven rating scale construction has highlighted the shortcomings of certain types of level descriptors, has unearthed mismatches between rating criteria and real-world TLU characteristics, and has stressed the need for empirically founded criteria (but see also Alderson, 2007; Jacoby & McNamara, 1999). However, various publications in the field of language testing have shown that a dichotomous rating scale typology (i.e., measurement-driven vs. performance-driven) may not correspond to actual practice, as many rating scales emerge from a variety of sources, including expert input, empirical performance data, and existing language proficiency frameworks such as the CEFR (Deygers & Van Gorp, 2015; Galaczi et al., 2011; Harsch & Martin, 2012; Knoch, 2009). Even the CEFR (Council of Europe, 2001) too, criticized by Fulcher (2004, 2012) as an example of statistically driven, intuition-based design, describes rating scale development as the process of combining intuitive, qualitative, and quantitative methods. More recently, in their overview of rating scale design, Montee and Malone (2014) too moved beyond a dichotomous approach, distinguishing between four types of rating scale development. A priori criteria, which resemble Fulcher’s idea of measurement-driven scales, are adapted from or linked to existing language proficiency standards (e.g., American Council on the Teaching of Foreign Languages [ACTFL] Guidelines, 2012; the CEFR), and may, as a result, be perceived as vague or abstract (e.g., Galaczi et al., 2011; Harsch & Martin, 2012). Second, empirically derived scales have much in common with performance-driven scales, since they are based on the analysis of task performances. Two additional rating scale development models described by Montee and Malone are theory-based scales, which are rooted in theoretical models of language proficiency, and criteria based on learning goals and curricula. They also discuss a hybrid approach, in which various types of development are combined, and argue how such a combination may balance out the disadvantages associated with a single approach (see also Knoch, 2009).
Montee and Malone’s identification of approaches to rating scale construction is reflected in a conceptualization proposed by Fulcher et al. (2011) himself. Using the aforementioned dichotomy as a basic structural element in his model, Fulcher identified “intuitive and experimental” and “scaling descriptors” as measurement-driven rating scales (p. 6). The first method corresponds with the above-mentioned a priori criteria; the second can be seen as a (Rasch) scaled version of an a priori constructed scale. Under measurement-driven scales, Fulcher et al. identified “performance data-based,” “empirically derived, binary-choice, boundary definition,” and “performance decision Trees” (p. 6). In their essence, all three types are empirically derived scales, operationalized in different ways to facilitate the rating process.
As can be seen from our review of the literature on scale development above, manuals, such as the Standards, are surprisingly silent on best practice in rating scale design, although they describe test development processes more generally. This literature review has offered an overview of the prevailing conceptualizations of rating scale development in language testing and assessment. The purpose of this study is to build on these conceptualizations and test them by taking into consideration real-world rating scale development processes, as described in the growing literature on this topic. We then propose a model that incorporates various test-internal and test-external sources which influence scale development and discuss how this model can be integrated into the wider test development processes.
Research question
The primary objective of this paper is to determine to what extent real-world rating scale development practices fit a data-driven versus measurement-driven approach. As such, the central question we wish to answer is as follows: What sources of information are used in the development of rating scales within the published literature on applied linguistics and language assessment?
Based on this research question, we examine whether the applied linguistics and language assessment literature offers sufficient evidence to support a dyadic conceptualization of rating scale design and whether another model could be considered, and how this could be integrated into the wider test development processes.
Method
To examine the methods and sources that test developers and researchers have relied on when developing rating scales designed to be used in rater-mediated assessment, we conducted a systematic review (Petticrew & Roberts, 2006) of empirical, peer-reviewed publications that describe scale development processes. Systematic reviews are a form of research, and follow systematic methods to answer previously specified research questions (Newman & Gough, 2020). They constitute a form of secondary analysis of already published research, and can be defined as “a review of existing research using explicit, accountable rigorous research methods” (Gough et al., 2017, p. 4). Although there are different types of systematic reviews, they typically follow the same underlying processes: develop research question; construct selection criteria; develop search strategy; select studies using selection criteria; code studies; assess the quality of studies; synthesis of study results to answer research questions; and report findings. The following section includes details of the search and selection process, inclusion criteria, and methods for coding the selected studies.
Search and selection process
Data included in this study were drawn from peer-reviewed research journals in both applied linguistics and language assessment. The search for the articles was conducted in March 2018 and consisted of two steps. First, we conducted searches in the following databases: (1) ProQuest, (2) Science Direct, (3) the Educational Resources Information Center (ERIC), (4) Scopus, and (5) the Web of Science. We used the following keywords as search terms: Rating scale design / Scale design / Rating scale development / Scale development / Rating scale revision / Scale revision / Scoring rubric design / Rubric design / Scoring rubric development / Rubric development / Scoring rubric revision / Rubric revision. The list of keywords used in the searches was constructed based on key words and synonyms commonly used in the literature we reviewed, trial searches in the data bases, as well as feedback from colleagues. Abstract, title, and article keyword searches were used. We limited the search to the following years: 1990 to 2018. We chose the start date, as this is, to our knowledge, just before the first articles on rating scale development started appearing in the applied linguistics and language testing fields.
We also conducted a systematic search of the following journals from 1990 to 2018: Language Testing, Language Assessment Quarterly, Assessing Writing, Papers in Language Testing and Assessment, and Language Testing in Asia. Although the majority of these journals publish impact factors, two of the journals on this list (Papers in Language Testing and Assessment, Language Testing in Asia) do not have their impact factors reported because their publishers do not calculate them or they have not yet published sufficient volumes to be assigned an impact factor. Language Testing in Asia, while not publishing an impact factor, reports a citation impact. This is the only journal not included in the Web of Science “Master Journal List” because it is a journal that charges authors for publication. Papers in Language Testing and Assessment is included in the Emerging Sources Citation Index and for this reason does not yet receive an impact factor. Owing to the specialized nature of the journals, and the fact that they are peer reviewed, we decided to include these journals in the list. The search of the selected journals was conducted through the advanced search option in Google Scholar, using the same search terms as described above. Because these journals were identified prior to the search, we conducted the search one journal at a time and limited the inclusion criteria in Google Scholar to the relevant journal and the years from its inception to March 2018.
The database search initially retrieved 23,586 studies. Following the initial searches, we narrowed our search by excluding any studies focusing on scales designed to measure constructs not related to language. We also excluded duplicates and studies that were not peer reviewed. We only included studies that described the development or revision of rating scales, whereas studies focused on scale validation (without including a sufficiently detailed description of how the scale was developed) and purely theoretical papers were excluded. Studies that were not published in peer-reviewed journals (e.g., book chapters, research reports) were excluded from the pool of papers. Studies that were explicitly only focused on L1 language assessment were also excluded. Following this process, we arrived at a final pool of 36 research papers. Table 1 lists the journals that had published articles that were included in the final corpus.
Information about selected articles.
Included in the Web of Science “Master Journal List” (https://mjl.clarivate.com/home).
Coding
Following the search and selection process described above, we determined, based on careful reading and rereading of each article by two coders (authors 1 and 3), what shaped the scale constructs in each of the studies and which sources impacted the design of the rating scales. By relying on this inductive coding, we identified 10 possible sources that impact scale construct, as set out in Table 2. Coding was conducted in Microsoft Excel 2016. The coding categories were grouped into three categories. We used the coding categories test-internal and test-external, depending on whether they relied on sources of scale construct from within the test (e.g., raters, or performance samples collected from test takers), or from outside the test environment. The third coding category refers to the context in which scores are reported or used. The coding categories listed in Table 2 can be linked to broad conceptualizations of measurement-driven, performance-driven (Fulcher, 2013), or theory-based (Montee & Malone, 2014) rating scale design, but they are more specific (e.g., a measurement-driven scale might encompass both expert intuition and existing scales, but we coded them separately).
Coding definitions of scale constructs.
Finally, we also noted the methods of quantitative analysis that were used in studies that were measurement driven. The detailed analysis of all 36 studies can be found in the supplementary materials (Section 1). The detailed list of all studies can be found in the supplementary materials (Section 2).
To determine the coding reliability, a random selection of 20 papers was re-coded by another coder (author 2), using the 10 sources of influence originally determined through inductive coding. Coding reliability was calculated in the same version of Excel. The inter-coder agreement was established, and proved satisfactory (95.4% exact agreement, χ2(1) = 0, p = .9). Differences between the coders were found to be based on one of three reasons: First, where raters were also from the TLU domain, they occasionally were unsure whether to code their involvement to the category TLU domain or Rater feedback, performance, cognition, background. This problem was occasionally exacerbated by vague descriptions or a lack of justification in the papers why certain groups of raters were involved. We coded the raters as test-internal if they were usually the raters of the test performances and to the TLU domain if they were described as being recruited because of their knowledge of the TLU domain. Second, an additional coding discrepancy happened when a scale development project was based on a previous study (also in the corpus of studies) where one coder coded the development features of the scale in the first study also in the second study. This was later resolved to avoid any doubling-up of information. Third, we included initially a coding category of scale developer intuition, which resulted in some coding inconsistencies; however, a careful further review of the studies in question led us to decide that all studies involved an element of scale developer intuition, regardless of its design, and we therefore excluded this category altogether. Following these changes, a further five papers were double-coded with no further coding discrepancies occurring.
Results
The Results section is divided into two main sections. We first present the results of our analysis of the corpus of 36 studies by describing the sources of scale design that we identified. We summarize these sources into a model of these sources of scale construct. We then briefly describe a number of sample studies that illustrate the various use of the sources we identified when coding the studies. Due to length limitations, we are not able to provide detailed descriptions of each of these studies, but nevertheless aim to show the variety of choices that scale developers make.
Which sources typically impact rating scale design in published research?
Table 3 summarizes the number of times each of the 10 sources of scale construct were reported in the corpus of studies. Nearly all studies (N = 35) drew on performance samples, and 34 of the 36 studies involved raters in one or more of their scale development phases. The third most common source of scale construct was a literature review (n = 19), followed by the use of existing rating scales (n = 14), expert intuition (n = 11), and language frameworks and assessment tasks (n = 9). The least frequently reported source of scale construct was the score reporting/score use context (n = 4).
Sources of scale construct in corpus of studies.
The studies varied greatly in the number of sources of scale construct they drew on. This is summarized in Table 4 and shows that some studies drew on up to six or seven of the 10 sources, whereas four studies only drew on two sources.
Number of sources of scale construct across studies.
The coding sheet in the supplementary materials (Section 1) further shows that a dichotomous categorization of studies does not reflect reality well. A purely performance-sample-based approach was not identified; this method, although present in all studies but one, was almost always paired with rater feedback and/or scoring, as well as other sources from the list. A purely intuitive approach was also not identified in the sample of studies. In reality, most studies drew on a mixed approach to scale development.
We also examined our data sheet to see if there were any typical combinations of sources on which scale developers draw, but apart from performance samples and raters often used in combination, we did not identify any other clear trends in the data.
Figure 1 graphically represents the different sources we identified to influence scale construct based on our coding of the 36 studies in the corpus. In the figure, we have sorted these into test-external, test-internal sources, and the score-reporting and score-use context.

Sources of influence on the scale construct.
We also included the policy context in Figure 1, as all testing occurs in a policy context, and all design decisions for rating scales (unless purely designed for research purposes) are generally directly or indirectly influenced by a policy. We will examine this in more detail in the Discussion section.
In summary, our coding of the corpus of papers, as well as our visual depiction of the various sources of rating scale construct we identified in our coding show that a binary model of scale construct is not capturing the complexity of scale development practices. The next section includes five concrete examples of how rating scale development studies have utilized the various sources we present in Figure 1.
Sample scale development studies
We briefly describe five representative studies from our larger corpus to illustrate the complex array of design decisions scale developers make when designing or revising scales. We selected these studies because collectively they illustrate the various design decisions we identified in our larger pool of studies. Owing to space limitations, we only report the key details of each study that illustrate the choices of construct representation. Because some studies included more sources of scale construct, the descriptions differ in length. Our coding of all these five studies can be found in the supplementary materials.
Study 1: Development of a comprehensibility scale (Isaacs et al., 2018; Isaacs & Trofimovich, 2012)
Isaacs et al. (2017) aimed to develop a user-oriented second language comprehensibility scale which could be used as a formative assessment tool within English for Academic Purposes courses in English-medium universities. The study was conducted in two main stages. In the first stage, reported in Isaacs and Trofimovich (2012), the authors asked untrained raters to listen to a series of speech samples collected from speakers of one L1 background and provide a comprehensibility rating (on a numerical, non-defined scale). The researchers then compared these comprehensibility ratings with analyses of the speech samples based on a range of L2 speech measures identified in the literature (e.g., measures of phonology such as word stress error ratio and vowel reduction ratio; measures of fluency such as total number of unfilled pauses; measures of linguistic resources such as the lexical error ratio; and discourse measures such as story cohesion, breadth, and depth). Based on the results of this analysis, a draft rating scale was developed. The draft scale was therefore based on two sources. First, a careful review of the literature and theory on comprehensibility resulted in the selection of the measures. Second, naive raters provided comprehensibility ratings. This draft scale formed the starting point for the work done by Isaacs et al. (2017). They conducted five focus groups over 12 sessions in which they introduced various contextual variables and collected information from potential scale users in a principled design with the aim of introducing context into the thus far theory-based scale. The focus groups aimed to generate L1-neutral descriptors for the scale to be used in mixed EAP classes (i.e., to broaden the group of test takers to which the scale applied), to ensure that the descriptors are usable for any academic extemporaneous task (i.e., to broaden the tasks to which the scale can be applied), to present the scale to new raters evaluating a different set of tasks (i.e., to broaden the raters and tasks), and to confirm that the scale would be intuitive to and manageable for EAP teachers who were relatively unfamiliar with rating systems from standardized tests (i.e., to ensure the usability of the scale by end users). They drew on various aspects of context (such as making the scale suitable for a range of academic tasks, score users, and L1 backgrounds of students) in order to ensure the scale was relevant to the specific context of EAP preparation courses.
Study 2: Assessing and reporting performances on pre-sessional EAP courses (Banerjee & Wall, 2006)
Banerjee and Wall (2006) aimed to develop an assessment checklist which could be used to generate a profile of students exiting EAP courses and provide information to university admissions officers to make admission decisions. The development took place over several phases. The developers started by conducting a careful literature review of the TLU domain which resulted in the creation of a draft scale. This scale was therefore arguably based on the TLU domain (albeit a theoretical representation of the TLU domain). The developers then introduced various layers of context by having the draft scale reviewed by both tutors (the raters) and admission officers (the recipients of the score profiles and decision makers). The scale was also reviewed by academic experts. The scale was then trialed by drawing on performance samples and revised based on the trial. In a final stage, the authors compared the construct coverage of the scale against assessment tasks used in the university domain, further increasing the precision of post-test predictions made about students.
To summarize the design decisions made by Banerjee and Wall (2006), the authors first reflected on the TLU domain by reviewing the literature. They also consulted the literature on academic language more generally. They then focused on various contextual variables (such as the raters and the score use context) to refine the scale. Finally, they returned to the TLU domain to ensure that the scale mirrored the construct coverage of the assessments in the domain.
Study 3: Rating scales for a classroom speaking task (Hirai & Koizumi, 2013)
Hirai and Koizumi (2013) described the development of scales designed to be used to score performances on a spoken story-retelling task used in a classroom context. The overall aim was to compare the effectiveness of the different scales. The scale developers started with a careful review of the curriculum goals and existing scales to inform the construct of the scales. They then focused further on the assessment context by reviewing performance samples from students responding to the story-retelling task. The scales were developed by the team of researchers drawing on the decision-tree which largely relies on the review of performance samples and the creation of key guiding questions to be considered by examiners. The resulting scales were task sensitive and based on both student performances and the curriculum goals. The scales were trialed by the teachers and various statistical analyses were conducted to evaluate the functioning of each of the scales. Finally, teachers were asked to review the scales for usability in their classroom contexts. The authors drew on a local curriculum, existing scales, performance samples, teacher raters, and the specific task type to inform the scale construct.
Study 4: Rating scale for school exam (Harsch & Martin, 2012)
Harsch and Martin (2012) described the development of a rating scale to be used in an English as a foreign language context at the end of lower secondary schooling in Germany. The assessment was designed to show students’ proficiency levels on the CEFR (Council of Europe, 2001). The draft scale was developed based on various sources, including actual CEFR descriptors, as well as the assessment grid from Into Europe, a school leaving exam aligned to the CEFR (Tankó, 2005), and a rating scale based on the CEFR used in the Germany context (Harsch, 2007). Any descriptors targeting four assessment criteria (task fulfilment, organization, grammar, and vocabulary) for five CEFR levels (A1 to C1) were included, after deleting redundancies. The scale then went through various stages of trialing and revisions, in which teacher raters used the draft scale, and based on their feedback and the quantitative results, the scale was revised. In sum, the initial scale developed by Harsch and Martin (2012) was based on two sources of construct, a language framework (the CEFR), and existing scales. The scale developers then drew on the assessment context, by training teacher raters, who then rated actual performance samples from a range of assessment tasks and revised the scale based on the teachers’ feedback and on quantitative results from statistical analyses.
Study 5: Rating scale for course placement (Plakans, 2013)
Plakans (2013) aimed to develop a scale for a writing placement test in a language program. The author led a group of teachers who were participating in the language program through the review of writing samples completed by students taking the placement test. The teacher participants drew on their teaching experience, as well as their knowledge of the program in which they were teaching, to develop the rating scale. They developed a decision-tree rating scale. The scale was therefore developed entirely based on the assessment context (i.e., writing samples collected from students completing the test, and the intuitions of the teacher raters working in the program). This was one of the few studies we identified that did not draw on any test-external sources of scale construct.
Our brief introduction of the five studies above provides an insight into the various design decisions made by rating scale developers in different contexts. It provides a snapshot of the various sources that are used to influence the scale construct and the complexity and, in some cases, longitudinal nature of scale development. A common thread across these studies is that they are all carefully designed and provide sufficient insight into and justifications for the decisions made during the scale development process. The studies differ in the number of sources of scale construct on which they rely in their design processes, ranging from two to five sources. The first four studies drew on both test-internal and test-external sources, whereas Plakans (2013) only drew on test-internal sources. What we can also see is that none of these studies can be neatly fitted into the dyadic description of scale development that we described earlier in the paper. Finally, it is clear from these five sample studies that documents such as the Standards should provide more guidance to test developers as to the decisions involved in rubric development. In the following section, we discuss the findings of our study, and draw up a model based on the results, and situate this model within the wider test development processes.
Discussion
Even though, like all models, Fulcher’s dichotomous approach was a purposeful oversimplification (personal communication, 12 September 2019) and even though he (Fulcher, 2003) and others (Montee & Malone, 2014) did identify more fine-grained overviews of approaches to rating scale development, the conceptualization of rating scales as either intuition-based (measurement-driven) or data-driven still largely dominates the discourse around rating scale development. Nevertheless, there have been indications in the literature that real-world rating scale development typically relies on a variation of sources. As such, the purpose of our current study was to establish the sources that have influenced rating scale construction in published research in the past decades, and to determine whether a binary conceptualization of rating scale design matches real-world practice.
To this end, we undertook a systematic review. We identified a corpus of 36 published scale development studies, which were then coded according to the sources of scale construct described in each paper. The results convincingly show that relying on multiple sources in rating scale construction (i.e., the hybrid form in Montee & Malone, 2014) is the norm rather than the exception in rating scale development. The corpus shows that all studies drew on more than one source of scale construct and that several drew on as many as six or seven sources. The primary sources reported on were performance samples, raters (scoring, feedback, background, and cognition), as well as theory and literature review. We summarized the 10 sources of scale construct that we identified by our inductive coding method in a model (depicted in Figure 1), and organized the sources as test-external, and test-internal.
Since all testing occurs in a policy space that shapes the conditions for testing to take place (Shohamy, 2006; Spolsky, 2004), all sources that impact rating scale design are directly or indirectly influenced by the policy context in which they operate. As such, we added this to our model. Logically, this policy context thus also shapes the assessment context, the way the assessment is developed, and the way in which scores are reported and used. By definition, policy is constantly in flux, and as such, the policy context is represented with a dotted line in Figure 1.
Policy can influence many aspects of scale development. In some cases, policy can support sensible assessment and scoring decisions. In others, a policy context may obstruct sound measurement (for examples of this, see Deygers et al., 2018). One obvious setting in which policy has influenced scoring is aviation language testing. Language proficiency requirements for pilots and air traffic controllers were first published by the International Civil Aviation Organization (ICAO) in the early 2000s (ICAO, 2004, 2010). The policy document stipulates how the testing of pilots and air traffic controllers working in international air space should be conducted. The ICAO language proficiency requirements (LPRs) include a rating scale which is designed to guide this process. Unfortunately, however, the scale is not without its problems (Alderson, 2011; Kim, 2013; Kim & Elder, 2015), as it includes features of construct irrelevance and construct under-representation of the TLU domain. The ICAO LPRs are a policy document which has had a significant impact on test design decisions made by various testing organizations.
Quite possibly no other language policy document has impacted language assessment more than the CEFR (Council of Europe, 2001; Byram & Parmentier, 2012; Little, 2007). In many contexts, in both Europe and further abroad, the use of the CEFR has proliferated, meaning that test designers are forced to draw on this policy document to have their local tests accepted for various purposes. Test developers therefore are required to link their tests to the CEFR. In quite a few cases, the influence of the CEFR directly impacts rating scale design decisions (we identified nine such studies in our corpus of studies). Crucially, however, the CEFR is a policy document, not a policy. A specific policy can therefore significantly impact how the CEFR is used within an assessment context. For that reason, the CEFR and other central language proficiency frameworks such as ACTFL (2012) are included as a separate test-external source of rating scale influence, rather than being subsumed in the overarching policy context.
The model we have presented in Figure 1 may not be complete (other design choices may exist which were not represented in the corpus of studies we analyzed) but it is quite probably more in line with real-world practice than a dyadic or compartmentalized approach. Since the selection of sources relied on clear inclusion criteria, the studies in the corpus are neither special nor cherry-picked. They represent what has been published on this topic within the field of language testing. The results may, however, be somewhat skewed as a result of publication bias. Since the corpus only relies on published, peer-reviewed research, the results may not be generalizable to rating scale development in every context. Quite possibly, less exemplary rating scale development practices than the ones published in peer-reviewed journals may abound.
As such, the model in Figure 1 above could prove to be a useful starting point as a theoretical representation which may inspire future research in this field. But more than that, it also has practical and theoretical applications. In the following sections, we explain how this model could be used for both scale development and scale validation purposes.
Guidelines for best practice in testing such as the Standards (AERA et al., 2014), or the ILTA Guidelines for Practice (ILTA, 2007) state that test development needs to start by determining the test purpose and score use (see, e.g., AERA Standards 1.1 and 4.1). Although the Standards also mention that “regardless of the type of scoring procedure, designing the items and developing the scoring rubrics and procedures is an integrated process” (p. 79) and on the surface it seems obvious that the test specifications, test items, and scoring criteria are all developed with the same considerations in mind, it is in fact very often the case that the construct considerations of the wider test are not deliberated when the scoring criteria are developed. There are many examples of this in the literature and in practice. For example, scales for specific purpose tests are often entirely linguistic in nature without taking into consideration the assessment criteria valued by domain insiders or the specific communicative demands of the TLU domain. The ICAO rating scale discussed above is one such example. Similarly, classroom assessment scales are often developed without considering local curricula, syllabi, or learning objectives (for some examples from our corpus of studies, see Jeffrey, 2015; Struthers et al., 2013; Youn, 2015).
Apart from the test purpose and context, there are three concepts related to scoring and score use that are particularly crucial considerations when developing scoring criteria: Not all of these are mentioned in the Standards or ILTA Guidelines for Practice as important considerations when scale development or revision projects are initiated: score generalizability, precision of post-test predictions, and scoring reliability. We argue that these are important to consider when starting out on a scale development or revision project, and that these need to be considered in tandem with the various sources of scale construct we have presented in our model above. Score generalizability relates to the degree to which a score can be generalized to new situations not expressly covered by the test construct. The precision of the post-test predictions describes how precisely the test can predict how well a test taker will cope in post-test communicative situations for which the test scores are being used as predictors. Scoring reliability, a concept more widely discussed in relation to scoring criteria, relates to the consistency of scores across raters.
Before starting out on a scale development project, test developers need to take into account the test purpose and score use contexts. Based on this, they must then decide which of the sources of test construct to draw on to optimally influence generalizability, precision of post-test predictions, and scoring reliability, all of which should be in line with the test purpose and score uses. Figures 2 and 3 convey this idea. Figure 2 focuses on scale design, whereas Figure 3 shows the scale design considerations in light of the wider test design considerations, so it focusses out a little more. In Figure 2, the light grey box (sources of scale construct) indicates the design choices we modelled in Figure 1. The figure clearly shows that both test purpose and score use should directly influence the sources of scale construct that test developers should draw on. The sources of scale construct also directly impact generalizability, precision of post-test predictions and scoring reliability and should therefore be considered when selecting sources of scale construct.

Relationship between sources of scale construct and test purpose, score use, generalizability, precision of post-test predictions, and score reliability.

Relationship between scale development and wider test development.
In line with Figure 2, the scale developer would first consider the test purpose and the score use. In many cases, these have the same intention (e.g., a test designed to assess academic language proficiency for entry into an undergraduate course may be used for exactly this purpose). In other cases, test purpose and score use may differ (e.g., a test designed to assess academic language proficiency for entry into an undergraduate course may be used to make migration decisions). Ideally, scale design decisions should be influenced by both test purpose and score use, but in cases of mismatch, score use should be the most important guiding decision. The test construct should always be designed to match score use. For example, if the score is used to make decisions about test takers’ readiness to enter a specific TLU domain (e.g. academic studies), then the scale construct should draw on the TLU domain. This would then ensure that score generalizability is reduced to this context only, and the precision of post-test predictions that can be made about test takers is increased. Scoring reliability should always be an important consideration as well, depending on the stakes of the test. Similarly, a scale designed to assess high school students’ foreign language writing might draw on a model or theory of writing, or a standards framework (both sources of scale construct which result in high score generalizability), or a syllabus, as no specified score use context is available (apart from making decisions about class progression). In this context, it is important to keep in mind Fulcher’s (2012, 2017) caution that a high level of generalizability may result in descriptors that are highly decontextualized and vague and therefore do not apply to any real-world contexts.
Table 5 sets out the likely influences of the 10 scale constructs identified in our systematic review on generalizability, precision of post-test predictions and scoring reliability. We have used the word “likely” because the effect of each of these construct choices may be mitigated by other design choices, and by other influences, such as rater training. Table 5 shows, for example, that if the scale construct is focused on the TLU domain, then score generalizability is likely to be reduced to that domain, but the predictions that can be made about test takers are likely to be more precise (as they are focused on a specific TLU context). It is difficult to make predictions about the effect on rater reliability, because this would depend on the raters used, and how familiar they are with the language use in the TLU domain. On the other hand, if the scale developer draws on a model or theory from the literature review or standards frameworks, then the score generalizability is likely to increase, as these sources of construct are generally designed to be broadly applicable. At the same time, the precision of the post-test predictions is likely to be reduced, as it is difficult to make precise predictions to a very broad score use domain. The effects on rater reliability in this instance is difficult to predict, as it depends on other factors, such as the background of the raters and the rater training. The effects of the use of existing rating scales, or expert intuition are unclear and depend by and large on the validity of the existing scale, or the background of the experts and the role they take. Test-internal sources of scale construct are likely to reduce score generalizability. In particular, a focus on test-internal performance samples is likely to reduce how broadly a score can be generalized to other types of tasks.
Influences of sources of scale construct on generalizability, precision of post-test predictions, and scoring reliability.
As many scales draw on more than one of these construct choices, the various influences will also interact and therefore the effect may be changed or reduced.
Other authors, for example, Bachman and Palmer (2010) and Fulcher (2012, 2017) have also discussed how rating scales impact on the specificity of post-test predictions and generalizability of test scores. To Fulcher, more general tests with general rating criteria may appear to yield more generalizable scores but might be too vague to be useable. As such, Fulcher (1996a; Fulcher et al., 2011) and others (e.g. Galaczi et al., 2011; Harsch & Martin, 2012) have argued that general scales which are based on language proficiency frameworks such as the CEFR should be enriched with (or at the very least tested against) real-world performance data drawn from the TLU domain in order to increase the construct validity and the generalizability of the scores, but we have seen that this is not necessarily the case.
A further application of a combination of the two models presented in Figures 1 and 2 can be used to guide validation activities, as the same considerations in Figure 2 and Table 5 can be used in order to review the appropriacy of scale construct choices made in the design of rating scales. In a recent paper, Knoch and Chapelle (2018) outlined a framework drawing on the argument-based approach to validation which researchers can draw on when validating scoring processes of assessments. As part of this, they formulated warrants and assumptions relating to a range of inferences. One assumption, for example, which underlies the explanation inference, states that “the rating scale is based on a defensible theoretical or pedagogical model of proficiency” (p. 489). The discussion in this paper and the models we have drawn up, provide further details as to how backing for this assumption can be collected. In the same vein, one of the assumptions relating to the extrapolation inference states that “the scale criteria reflect the evaluation criteria used in the TLU domain” (p. 491). Again, direct application of the two figures presented above can show how backing for this inference can be collected. This study showed that there are numerous examples of scale development studies in which the test purpose and score uses described were at odds with the sources of scale construct on which the developers drew. It is not our intention to list these here, but to draw attention to how scale development practices in the future could be reviewed in order to enhance our current scale validation practices.
Conclusion
We set out to explore in more detail the various sources of scale construct on which test developers draw when developing rating scale for rater-mediated assessment. We were able to show that the dyadic model of rating scale development is not representative of current practices in scale development. Based on our analysis, the resulting model and the sample rating scale design studies, we showed the various complex decisions scale developers have documented in various language assessment contexts. We have also documented the impact of each of these sources on generalizability, the precision of post-test predictions and rater reliability, and argue that scale developers can use this model to guide decision-making when embarking on scale development projects. Although influential guides on best practice in test development stipulate that the test purpose and score use need to be considered during test development processes, this is often not the case. Many rating scales draw on sources of scale construct that are at odds with the test purpose and score use context, or at odds with each other. We hope that this paper draws attention to this issue and furthers theoretical considerations in rating scale development.
As described in the Discussion section, the model we have put forward has the potential not only to inform and guide future rating scale design decisions, but also to expand our conceptualizations of scoring validity. Rating scale developers need to consider the impact of their design choices on score generalizability, the precision of post-test predictions, and rater reliability when setting out on scale revision studies. We hope that future research will further refine the model of the scale construct we identified in our study and also investigate the impact the use of more than one of these sources has on the generalizability, precision, and reliability of test scores.
Supplemental Material
sj-pdf-1-ltj-10.1177_0265532221994052 – Supplemental material for Revisiting rating scale development for rater-mediated language performance assessments: Modelling construct and contextual choices made by scale developers
Supplemental material, sj-pdf-1-ltj-10.1177_0265532221994052 for Revisiting rating scale development for rater-mediated language performance assessments: Modelling construct and contextual choices made by scale developers by Ute Knoch, Bart Deygers and Apichat Khamboonruang in Language Testing
Footnotes
Author’s note
Bart Deyger is now affiliated with Department of Translation, Interpreting and Communication, Ghent University.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
Supplemental material
Supplemental material for this article is available online.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
