Abstract
This study investigated the assessment literacy required for K–12 educators to interpret score reports from a K–12 English language proficiency assessment. The assessment in concern is ACCESS for ELLs, which is an annual summative assessment that is delivered to nearly 2 million English learners (ELs) across 39 US states and territories. This study was conducted in two phases. In Phase 1, an online teacher survey, consisting of 15 items, was completed by 1437 participants; data were analyzed using descriptive statistics. In Phase 2, 18 educators were interviewed to obtain in-depth understanding of educators’ interpretation of the score reports. Interview data were qualitatively analyzed in terms of (1) the essential assessment literacy required for interpreting the score reports, (2) resources referred to by educators for score report interpretation, and (3) educators’ suggestions to enhance the score report and its interpretation. The findings of the study reveal the relevance of K–12 EL educators’ assessment literacy for score report interpretation. For example, educators referred to proficiency level scores the most, but experienced difficulty in understanding technical terms, such as “scale scores” or “confidence band.” The results provide implications for enhancing the quality of score reports and the development of resources to support educators’ assessment literacy.
Keywords
The number of K–12 English learners 1 (ELs) in the United States has been steadily increasing, with ELs making up approximately 9.4% of public school students in the 2014–15 school year (USDE NCES, 2017). This rate underscores the continued need for high-quality English language proficiency (ELP) assessments. These assessments serve multiple purposes. For example, they are used for (1) placing ELs into appropriate language instruction educational programs to provide them with necessary English language support and services, and (2) monitoring ELs’ progress in ELP for federal accountability purposes as described in the Every Student Succeeds Act (2015).
In order for K–12 educators to make sound educational judgments based on students’ performance on ELP assessments, they need to understand the information in score reports. This requires a certain type of assessment literacy, namely, the set of skills and knowledge required to carry out assessment-related activities (discussed in further detail in the next section). Although it has been observed that different stakeholder groups may require different types of assessment literacy, and that classroom teachers may have specific assessment literacy needs in carrying out routine classroom assessment procedures (Crusan, Plakans, & Gebril, 2016; Taylor, 2013; Xu & Brown, 2016), the specific type of assessment literacy required by educators to interpret score reports has been under-explored.
In the current study, we examine the assessment literacy required by K–12 EL educators for interpreting Assessing Comprehension and Communication in English State-to-State for English Language Learners (ACCESS for ELLs; hereafter ACCESS) score reports and explore ways to enhance them. We also examine the role of resources for assisting score report interpretation. This paper is part of a larger two-phase study that examined parents’ and educators’ interpretation and use of previous and newer ACCESS score reports with the purpose of enhancing the quality of score reports. The current paper presents findings relevant to educators’ interpretation of the new score reports. To this end, we used a two-phase mixed-methods sequential explanatory design (Creswell & Plano Clark, 2011), which involved collecting and analyzing online survey data, followed by interview data. In Phase 1 of the study, research team members collected data on educators’ assessment literacy for interpreting score reports by using an online survey; we quantitatively analyzed the survey data to understand educators’ assessment literacy, including how they interpret score reports and utilize resources (see the “Methods” section for details); in Phase 2, to confirm and expand the understanding of educators’ assessment literacy, we collected interview data and qualitatively analyzed them via content analysis. The interview findings from Phase 2 informed the interpretation of survey findings from Phase 1. The following three research questions were addressed in the study:
To what degree do K–12 EL educators demonstrate assessment literacy in interpreting English language proficiency score reports?
To what degree do K–12 EL educators utilize score report resources for interpreting English language proficiency score reports?
How could English language proficiency score reports be enhanced to meet K–12 educators’ needs?
Literature review
Language assessment literacy
Defining assessment literacy or, more specifically, language assessment literacy 2 is a challenge due to its complex nature. Pill and Harding (2013) defined assessment literacy as “a repertoire of competences that enable an individual to understand, evaluate and, in some cases, create language tests and analyse test data” (p. 381). Meanwhile, Inbar-Lourie (2008) considered assessment literacy as both knowledge and a set of competences. That is, assessment literacy not only involves what practitioners may know but also what they can do. Fulcher (2012) included awareness of “the role and impact of testing on society, institutions and individuals” (p. 125) as part of the definition, in addition to skills, knowledge, and abilities. As implied by these varying definitions, assessment literacy consists of multiple elements.
In one of the seminal works on the topic, Davies (2008) saw assessment literacy as having three main elements: (1) skills (expertise in test development and analysis); (2) knowledge (theories of language and measurement); and (3) principles (concepts in assessments such as validity). Taylor (2013) provided a more expanded list of assessment literacy characteristics, consisting of eight dimensions: (1) knowledge of theory, (2) technical skills, (3) principles and concepts, (4) language pedagogy, (5) sociocultural values, (6) local practices, (7) personal beliefs/attitudes, and (8) scores and decision making. Taylor acknowledged that depending on their position or role, the degree of an individual’s literacy may vary in each dimension. For instance, language testing experts need a high level of knowledge of theory in language assessment to develop and validate assessments. Likewise, higher education instructors in the field of TESOL or applied linguistics need knowledge of theory to teach language assessment courses to future educators. Meanwhile, EL instructors and educators in universities and K–12 settings require sufficient knowledge on the dimension of language pedagogy to develop classroom assessments that could inform their students’ learning or to interpret the score data from the standardized tests their students take. Therefore, a classroom teacher may need more knowledge of language pedagogy than knowledge of theory.
Adapting the works mentioned above, we define assessment literacy as the set of knowledge, skills, and abilities needed to work with various language assessments, such as administering assessments, scoring, interpreting scores, and making decisions based on the interpretations. In addition to the knowledge, skills, and abilities, we consider the social context of the assessment, and include in the definition familiarity with test processes and awareness of the principles and concepts that guide practice (Fulcher, 2012). As discussed above, assessment literacy is required for a wide range of assessment-related activities, and the type and degree of assessment literacy may vary for different groups of stakeholders (e.g., language testing experts, EL educators).
The topic of assessment literacy has received much interest in recent years, but there is a lack of research in a number of areas. Taylor (2013) listed four areas in assessment literacy that require more research: (1) defining the assessment literacy construct; (2) language and discourse needed when addressing a non-specialist audience; (3) identifying, evaluating, and responding to stakeholder needs in assessment literacy; and (4) growth of assessment literacy over time.
Taylor’s third area of identifying, evaluating, and responding to stakeholder needs in assessment literacy is particularly important in the K–12 context; although a limited number of educators such as district testing coordinators may have some background in measurement, the majority of educators may lack in-depth assessment literacy (Zapata-Rivera, Van Winkle, & Zwick, 2010; Zwick, Zapata-Rivera, & Hegarty, 2014). Nevertheless, they are required to engage in formative assessment of their students on a daily basis (e.g., giving feedback about classroom performance) and to interpret and use summative assessment scores (e.g., annual English language proficiency outcomes). As Taylor (2013) wrote, identifying the range of relevant stakeholders and evaluating their specific needs in relation to what test scores mean in their context and, consequently, how scores can or cannot be used, is becoming a priority in a world where assessment occupies such a central role. (p. 407)
Considering that the teacher is the most important factor affecting student learning, it is critical to understand educators’ assessment literacy (Crusan, Plakans, & Gebril, 2016; Darling-Hammond, 2000; Scarino, 2013; Vogt & Tsagari, 2014; Xu & Brown, 2016). Although educators are involved in making various assessment-related decisions, such as placement of language learners into appropriate instructional programs or levels, they are requested to do so often with little formal background or training (Xu & Brown, 2016). Vogt and Tsagari (2014) wrote that primary and secondary teachers often lack assessment literacy and feel the need for additional training in the area. The authors conducted a mixed-methods study to examine the assessment literacy of foreign language teachers. Findings from survey (n = 853) and interview (n = 63) data across seven European countries show that educators’ assessment literacy is under-developed due to their pre- and in-service training, and they are required to compensate for their lack of assessment literacy on the job. These results emphasize the need for assessment literacy among educators.
Assessment literacy can be beneficial as it serves a dual goal of enhancing educators’ understanding of assessment and their self-awareness as assessors (Scarino, 2013). Moreover, assessment literacy training should be long-term and sustainable, so that educators can readily connect their knowledge with practice as needed (Xu & Brown, 2016). For example, K–12 educators should have an understanding of the assessments their students engage with and know how to interpret score reports from these assessments so as to understand their students’ performance.
Language assessment literacy and score reports
Due to the lack of research on K–12 language teachers’ assessment literacy, there is a need to look further afield for relevant research, such as higher education. Previous research (e.g., Brown & Bailey, 2008) indicated that university instructors need a relatively high degree of assessment literacy, especially those who are in charge of language assessment training. These studies have examined the characteristics of the language assessment instructors, courses, and students. For example, Kleinsasser (2005) discussed the challenges instructors face when teaching language assessment courses, such as helping students apply theories in practice.
Similarly, a number of studies have been conducted to understand university-level ESL educators’ or administrators’ assessment literacy (e.g., Baker, Tsushima, & Wang, 2014; Coleman, Starfield, & Hagan, 2003; O’Loughlin, 2011, 2013). For instance, O’Loughlin (2013) examined how university educators, administrators, and admissions staff interpreted and used International English Language Testing System (IELTS) scores. Findings from an online survey (n = 50) and follow-up interviews (n = 15) indicated that educators focused on the minimum scores needed for university entry and used that information for advising students regarding English proficiency requirements and making admission decisions. Moreover, results show that educators referred to the university’s English language entry regulations more than the IELTS Guide or resources.
In a similar study, Baker et al. (2014) conducted a survey study to examine Canadian university admissions officers’ (n = 19) assessment literacy. The findings from this study were more positive than those of other research discussed above, as the participants demonstrated their knowledge of basic concepts related to language assessment, such as cut-off scores for making decisions, and showed a strong interest in developing their assessment literacy.
Results from both O’Loughlin (2013) and Baker et al. (2014) highlighted the importance of assessment literacy related to score reports in the interpretation and communication of scores (Xu & Brown, 2016). In other words, educators need to know how to interpret evidence generated from assessments; they should also know suitable methods for communicating assessment scores to relevant stakeholders, including students, parents, and administrators.
Compared to research on educators’ assessment literacy in higher education, there has been less research on the topic in the K–12 setting. Existing literature suggests that when interpreting score reports, teachers search for information that could be useful for instruction (Luecht, 2003; Underwood, Zapata-Rivera, & VanWinkle, 2007). In a framework for designing and evaluating score reports for different audiences, including educators, Zapata-Rivera (2011) suggests that score reports for educators may include information on task levels, formative evidence to inform instruction, performance levels, and scale scores.
However, research shows that K–12 educators often lack assessment literacy and therefore experience difficulty understanding terms in score reports, such as those relating to measurement error concepts (e.g., Lukin, Bandalos, Eckhout, & Mickelson, 2004; Zapata-Rivera et al., 2010; Zwick et al., 2014). According to Hambleton and Slater (1997), K–12 educators frequently misunderstand or ignore statistical jargon such as standard error or significance, symbols, and technical footnotes. Some users even feel intimidated by these terms. In a more recent study, Zapata-Rivera (2011) conducted a usability study of a prototype score report with 12 sixth- and eighth-grade teachers. Although educators generally understood concepts such as item difficulty, scale scores, and raw scores, they struggled to interpret standard error of measurement. Therefore studies (e.g., Wainer, Hambleton, & Meara, 1999; Zwick et al., 2014) have explored ways to provide enhanced visual representations of statistical information so that users can draw appropriate inferences.
In addition, Stiggins (1995) described how teachers lack sufficient knowledge to assess their students in classroom settings. The development of standardized assessments rarely involves educators, meaning they miss an opportunity to develop assessment literacy. This skills gap also occurs in English language education, which has raised much concern (Inbar-Lourie, 2008).
Overall, few studies have examined assessment literacy in the K–12 EL context, let alone the assessment literacy required for interpreting score reports. In addition, few have explored how educators use score report resources for score interpretation (Zenisky & Hambleton, 2012), suggesting more research is needed on these topics. Considering that educators directly affect student learning, a better understanding of their uses and interpretations of score reports could have a significant impact on students’ learning.
Methods
Context of the study
In this study we examined a score report from ACCESS, an annual summative test designed to measure English language proficiency in ELs. ACCESS is jointly developed by WIDA and the Center for Applied Linguistics and is delivered to nearly 2 million K–12 ELs in the United States each year (see https://wida.wisc.edu/ for more information regarding ACCESS and its research and development). As indicated in the Every Student Succeeds Act, the federal government requires states to annually assess their ELs using an English language proficiency test; to fulfill this requirement, currently 39 states use ACCESS. The test assesses the four domains of listening, speaking, reading, and writing. WIDA recommends its stakeholders use ACCESS scores for four primary purposes: (1) for assessing students’ English language proficiency; (2) for placement of students into appropriate programs or levels; (3) for evaluating the effectiveness of EL programming; and (4) for measuring progress in students’ English language development.
After students complete the test, WIDA provides several score reports to stakeholders, including teachers, EL students, and parents of the students. These include the following: (1) the Individual Student Report for educators, parents, and students, which became available in the 2016–17 school year; (2) the Student Roster Report for educators; and (3) Frequency Reports for schools, districts, and states for administrators. With the release of the online version of ACCESS in the 2015–16 school year, the score reports have been redesigned with the goal of effectively communicating test-takers’ performance on the online versions of the test. One of the main changes was to combine the features of the previous Teacher Report and the Parent/Guardian Report into a unified score report—namely, the Individual Student Report. Thus, rather than having two separate score reports for teachers and parents, WIDA developed a single score report for both the teacher and parent audiences.
Participants
In Phase 1 of the study, 1437 EL educators from Grades K–12 completed an online survey. Participants were recruited via convenience sampling by distributing the survey using an existing WIDA listserv for K–12 EL educators. During the time of data collection, approximately 37 states and territories were part of the WIDA Consortium. Participants were from 35 US states and territories. Among the participants, 60% taught students in Grades K–5, and the remaining 40% taught students in Grades 6–12. Educators had varying years of teaching experience: 10% had 0–2 years, 20% had 3–5 years, 25% had 6–10 years, and 45% had more than 10 years of teaching experience.
In Phase 2, 18 educators participated in interviews: 12 were in-service EL teachers and six were EL coordinators, all recruited during a national conference attended by K–12 educators. This conference was selected due to its popularity with the large numbers of educators who were the specific focus of this study. All educators taught or supported ELs ranging from pre-Kindergarten to Grade 12. While EL teachers interacted with ELs on a daily basis, EL coordinators had experience as EL teachers and were supervising EL teachers at the district level at the time of the interview. As summarized in Table 1, educators’ backgrounds varied: they were from 13 US states and had on average 12.6 years of experience in the field. All but two had at least a master’s degree; two had doctorates.
EL educators.
Instruments
Three documents played roles in this study: the student score report, the online survey for educators, and an interview protocol.
ACCESS for ELLs score report
The score report used in this study was a draft of the Individual Student Report, which became available as a new report to ELs and their parents and educators in the 2015–16 testing year. This study used a draft of the report because final versions of the report were not available at the time the study was conducted in the 2014–15 testing year. The draft closely matched the final version of the score report.
The Individual Student Report includes the following: (1) student background information; (2) the purpose of the report; (3) information on the student’s English language proficiency level, scale score, confidence band for each domain (listening, reading, speaking, and writing), and combined domain scores (literacy, comprehension, oral language, overall scores); and (4) a description of English language proficiency levels (see Appendix A for an example of the report).
In comparison to the Teacher and Parent/Guardian Report from previous years, the Individual Student Report contains enhanced features. It is printed in color, whereas previous reports appeared in black and white. The new report includes enhanced visuals, such as bar graphs of test-takers’ proficiency levels for each language domain. Moreover, a new description of English language proficiency levels supports the interpretation of each proficiency level.
Online survey
An online survey was developed using Qualtrics (https://www.qualtrics.com) to collect information regarding teachers’ perception of the meaningfulness and usefulness of the score reports. The survey comprised 15 multiple-choice items regarding (1) participants’ background information, (2) teachers’ perceptions of the Teacher Report (from previous years), and (3) teachers’ perceptions of the new Individual Student Report. Some of the items allowed individuals to add comments. The current study focuses on the survey items regarding the Individual Student Report.
Semi-structured interview questions
Semi-structured questions guided interviews with educators. Compared to the online survey items, the interviews involved more in-depth questions regarding educators’ interpretation of score reports. The interview questions were largely categorized into (1) participants’ background information, (2) the previous Teacher Report, and (3) the new Individual Student Report. Interviewers were permitted to ask additional questions when needed. The current study focuses on the interview items regarding interpretation of the Individual Student Report.
Procedures for data collection and analysis
In Phase 1 of the study, research team members distributed the online educator survey to K–12 EL educators on the WIDA Consortium listserv. The survey was kept open for two weeks in May 2015. Survey data from the 1437 respondents were analyzed using descriptive statistics. Because respondents had the option to skip items, certain items had fewer than 1437 responses. Individual comments from the items were qualitatively analyzed.
The Phase 2 interviews with 18 educators were conducted individually or in groups. During the interviews, educators were shown a sample student’s Individual Student Report and asked questions regarding their interpretations of it. Interviews lasted 30 minutes on average. They were audio recorded and then transcribed for further analysis.
Interview data were qualitatively analyzed for content in terms of three initial themes reflecting the research questions (see the coding scheme in Appendix B): (1) interpretation of score report information; (2) use of score report resources; and (3) suggestions for improvement in the score reports. The first and third themes were further elaborated into subcategories. Interpretation of score report information included the following subcategories: (a) understanding of each section of the score report; (b) helpful information from the score report; and (c) unclear information from the score report. The latter two subcategories were examined because the type of information that is helpful signals which information educators can comprehend and process. In addition, it is unlikely educators may consider certain pieces of information to be helpful if they struggle to comprehend these data. Meanwhile, aspects of the score report that are unclear to educators could indicate gaps in assessment literacy. The third main theme of suggestions for improvement included the following subcategories: (d) information that should be added to or deleted from the score report; and (e) suggestions for improving the score report. Two independent coders reviewed the interview data and assigned codes to data that reflect the three main themes. In the case of discrepancies in their coding, the independent coders discussed until they reached agreement. 3
Findings
Study findings are organized in relation to the three research questions. To address these questions, relevant findings from the survey and interview results are presented below.
Research Question 1: To what degree do K–12 EL educators demonstrate assessment literacy in interpreting English language proficiency score reports?
K–12 EL educators’ assessment literacy was examined by first identifying the degree of educators’ understanding and interpretation of ACCESS score report information. This analysis not only involved examining what educators understand from the score report, but also exploring the score report information that they perceived to be helpful or unclear. Specifically, findings from Survey Item 10 and Interview Questions 10–12 were used to respond to this research question, as shown below:
Findings from Survey Item 10 indicate sections of the score report that educators found helpful and meaningful. These sections indicate the assessment literacy required for interpreting ACCESS scores, as educators can likely find the graphs, scores, and descriptions helpful only if they have an understanding of the student information presented. As seen in Figure 1, approximately half of the participants rated the description of English language proficiency levels to be “very helpful,” suggesting it to be the most helpful.

Helpfulness of the new Individual Student Report for understanding students’ performance (n = 1424).
The description of English language proficiency levels provided information on what students can do at each of their proficiency levels. Survey findings show that it was not only important to identify the numeric scores that students received, but also to understand what skills students can actually demonstrate. The description directly connects to how ELs may be able to perform in classroom settings and provides instructional support to educators. This finding suggests that educators can easily digest score report information that is connected to concrete language skills of ELs. It also reveals what educators value, that is, verbal descriptions that connect numeric scores to actual student demonstrations and performances.
Moreover, interview findings added details regarding the degree of assessment literacy educators demonstrated in interpreting ACCESS score reports. The survey findings indicated the importance of understanding language domain scores for listening, reading, speaking, and writing, together with composite scores. The interview results specifically indicated the type of scores, such as proficiency level, scale score, and confidence band, that were particularly helpful or challenging when interpreting ACCESS results. As detailed below, although the majority of educators were able to interpret proficiency level scores, some struggled to understand scale score or confidence band information.
According to findings from Interview Question 11, five participants (AM, LS2, IT, EG, and VU) found that the proficiency level information (on the four language domains and the combined domains) was necessary and useful when interpreting the score report. For example, EG stated, “You can see what modality your students are going to perform better in, so you can visually see if your students are better at one [language domain] or another.”
In addition, four participants (Participants CS, EG, JP, and NW) indicated that the proficiency level was useful for interpreting further the students’ growth in their English language ability. They used proficiency levels to check for improvements in ELs’ scores. For example, CS compared the proficiency level index in the current report with that from the previous year’s report; districts often provided previous year’s score report information on their online database systems. She believed that “if there’s a growth … a child should … at least grow 0.5 [in their proficiency level].”
Together with the proficiency level information, findings indicate that the description of English language proficiency levels, located at the bottom of the score report, helped EL teachers understand what their students can do with their current English proficiency levels. This assistance also applies to cases in which EL teachers share student information with mainstream classroom teachers or content subject teachers in the same school who have ELs in their classes, as EG describes below.
I really like the box at the bottom with what students can do, especially when trying to explain to [other mainstream or content] teachers, when the teacher is trying to figure out “well, how do I deal with this student, I don’t know what to try,” so I think it’s really nice to have concrete things that [a] student at or around [that] level generally can do. (EG)
On account of content teachers’ lack of training in interpreting ACCESS data, they are often uncertain about what ELs in their classrooms can do using English. EL teachers can function as advisers for content teachers by helping them better understand the score report and what they should expect from ELs in their classrooms.
Findings on Interview Question 12 revealed that several EL educators (AM, CK, CS, EG, LG1, and LS2) do not fully comprehend the technical terms in the score report, for example scale score or confidence band. These educators struggled to understand scale scores. AM specially addresses his uncertainty below.
seeing the number 367 [on the scale score], I don’t know what I’m supposed to do with that number versus 320. All I know is that 368 is higher than 320. And I don’t know what to do with those. (AM)
He added, “[t]his is lack of my background knowledge about looking at scale scores … I just don’t know.” As suggested by AM, this knowledge gap could be partly attributable to the lack of training that EL educators receive or not having the time or access to score report resources. Attributing issues identified in terms of score report interpretation is generally challenging as there may be multiple reasons underlying the problems, ranging from the assessment literacy of the educators, to the design of the score report, to the complexity of the assessment.
Overall, findings offer a general picture of K–12 EL educators’ assessment literacy in relation to ACCESS scores. They have the ability to interpret the proficiency levels and connect the numeric levels with descriptions of what students can do using their English language. The description of English language proficiency levels are meaningful as they provide concrete instructional implications. However, educators still lack knowledge of technical terms, such as scale scores and confidence bands.
Research Question 2: To what degree do K–12 EL educators utilize score report resources for interpreting English language proficiency score reports?
Research Question 2 was addressed by examining the materials and resources that K–12 EL educators used to support their interpretations of ACCESS scores. To this end, findings from Survey Item 9 and Interview Question 19 were analyzed.
The Interpretive Guide for Score Reports is available to the public on the WIDA website. This guide is designed to be the main resource for supporting educators’ interpretation of ACCESS score reports. We used response data from Survey Item 9 to examine how often participants referred to the Interpretive Guide to understand the score report information. Results showed they accessed it with varying degrees (Table 2). Approximately 60% of the participants referred to the guide (i.e., sometimes, often, or all the time), whereas 40% of them rarely or never used it.
Frequency of use of the Interpretive Guide.
In response to Interview Question 19, several participants indicated that they were aware of the Interpretive Guide. Two educators (LG2, VU) used it as a reference, while three others (CK, EG, LL) rarely used it. The reasons for teachers’ limited use of the Interpretive Guide vary, but as CS said, the main reason was a lack of time.
No, no, I think it’s not hard to comprehend, but what I miss is that … having someone – explain the nitty-gritty, the highlight, rather than me going through it [using the Interpretive Guide], and not being an expert in data, and reading through it, it would take me a lot longer, I understand it. But if we have a person [with expertise in data, such as an EL coordinator,] that can just say, “okay this is what you need.” And [teachers] go through it, and they, it clicks for them, and they transfer it in the student … and but no, it’s not difficult. (CS)
CS acknowledged that she is not “an expert in data” and her lack of expertise means she must devote a considerable amount of time to interpreting fully the score report. This response showed that she prefers to have someone explain the scores to her because the technical vocabulary used in the Interpretive Guide makes the guide less user-friendly. In addition, lack of time is an issue for educators who need to interpret and understand their students’ score reports. When planning revisions to the Interpretive Guide, WIDA should consider the fact that educators (and especially EL educators) have a limited amount of time to read the guide.
Six educators (AM, CS, EG, JP, LS1, and VU), who share score reports with content teachers, reported that content teachers, because of their lack of training and limited knowledge about how to interpret language-related data, rely on EL teachers to interpret the reports. Because EL teachers reported that content teachers depend on them for interpretation, it is critical to have a clear Interpretive Guide that first considers the needs of EL teachers.
Moreover, four educators (JE, JP, LS2, and NW) referred to the WIDA Can Do Descriptors (WIDA, 2016), a separate document from the score report, which provide detailed examples of what students at various levels of proficiency can do with their English language in school settings. JE responded as follows: “Teachers look specifically at these language domains and correlate them to the Can Do Descriptors to get some ideas for differentiation.” She used the proficiency level index in conjunction with the Can Do Descriptors.
These results indicate that educators referred to the Interpretive Guide for Score Reports and the WIDA Can Do Descriptors to interpret score reports. However, they had limited time to refer to these resources, suggesting the need for more concise and user-friendly materials.
Research Question 3: How could English language proficiency score reports be enhanced to meet K–12 educators’ needs?
Findings from Survey Items 14 and 15, and Interview Questions 14 and 20 were examined to respond to Research Question 3:
In response to Survey Item 14, the majority of participants (91%) wanted to know about students’ scores from previous years (Table 3). According to individual comments, participants were very much in favor of having previous scores included in the score report. They wanted information on students’ growth (from the previous year), not only regarding their overall language ability, but also in individual language domains of listening, reading, speaking, and writing. Comments (individual notes written in the “other” option) also showed the need for an additional score index that indicates students’ performance in comparison to other EL students. This information could allow teachers to understand how students are performing compared to the norm.
Suggestions for the new score report.
Note: Participants were allowed to select multiple options; therefore the total does not add up to 100% and is not reported.
In addition, findings from Survey Item 14 suggest including actionable information to improve students’ language ability (65%; Table 3). For instance, individuals suggested creating enhanced online videos and documents to guide/train teachers how to interpret and utilize score reports to inform instruction. Findings from the same item indicated the need for more details regarding students’ ability in each language domain (57%). Some respondents requested more fine-grained details, such as students’ language performance and strength/weakness in various language standards, including the languages of mathematics and science. Other individual comments indicated the need for exit criteria information in each domain.
In addition, the response to the open-ended Survey Item 15 revealed the need to provide score reports soon after test administration. Currently, it can take up to 4 months to receive the score report. Since children make progress quickly in their language development, a report received months after the test administration may not be very informative. At the latest, score reports need to arrive before the new school year to grant teachers the time to plan for next year. Other individual comments indicated the need for a section to mark students’ placement in language instruction educational programs based on their test scores, and score reports in students’ English language. These comments included having a separate translated version or having the translated version on the back page of a double-sided report.
Results from Interview Questions 14 and 20 generally confirmed the findings from the survey items. Although the new score report is meant to serve as an enhanced version that replaced the old score report, participants offered various suggestions for improvement. First, several participants (CK, CS, LG1, LS2) suggested that WIDA include test-takers’ scores from the previous year. As demonstrated by the response of CS, having the previous year’s scores in the score report would make it more convenient for educators easily to interpret student growth: “I think it’s just a dream of mine, but it would be nice to have … the previous year, maybe … Then we don’t have to do a lot of digging … so it’s just convenient for us [and] would be nice.” According to LG1, parents and EL teachers can benefit from seeing at first glance how much their students’ English language proficiency has grown over the year. CS specifically suggested that using different colors for the past vs. current year’s scores may be helpful.
Educators also indicated that they wanted to see the full description of English proficiency levels to have a clearer idea of what students can demonstrate in their proficiency levels. The descriptions on the current score reports are limited to the level of the students’ performance in the four language domains (listening, reading, speaking, and writing). In addition, NW wanted “more information in … for reading and writing tasks, like some … like more types of things that students can and cannot do … I think that would be helpful for classroom teachers.”
Educators (CS, LG1, LS1, MC, and SC) also wanted more information to support their instruction, especially for teaching content areas. For example, LS1 wanted information on the specific lessons/tasks that she could have her students work on to improve their weaknesses. However, it may not be appropriate to include such detailed amounts of instructional materials in a score report. Also, WIDA does not necessarily recommend that educators use ACCESS scores to inform their daily instruction as ACCESS is not formative in nature. Rather, due to the summative nature of ACCESS, WIDA recommends that its stakeholders use ACCESS scores for making placement and programming decisions.
Other suggestions included the following: (1) adding students’ home language information; (2) adding information regarding student’s test completion rate; (3) enlarging the font of essential information, such as composite scores; and (4) delivering the reports more quickly. Although ACCESS score reports may be shipped to districts and schools in advance, educators may not have access to them. SC, commenting on a need to have score reports available more quickly, put it this way: Part of the challenge, if I could also add, is that we get this report for us probably a week and half before we start school. And so by then, you know, class lists are already made, things are already set up, and it’s very difficult, especially [at] the middle school, high school level, we try to put all of our ELs on one team in each grade level, so they can be better supported. And we end up having to move kids around, last minute … So, we were all hoping with the … newer tests starting earlier and being more computer-based, we might get things earlier so that a lot more prep could be handled, conversations could take place to get better use of the data earlier rather than once the kids are back. (SC)
Overall, educators’ suggestions provide guidance in further improving the quality and delivery of the score reports, which could ultimately enhance educators’ assessment literacy to interpret the reports.
Discussion and conclusion
In this study, we investigated the assessment literacy that K–12 EL educators need for interpreting score reports and identified the resources they use to interpret those reports. As this study was part of a larger project at WIDA that aimed to enhance the quality of ACCESS score reports, findings provided guidance for refining the score reports and enhancing score report interpretive guides. Results show that educators find interpreting technical terms, such as scale scores or confidence bands, to be challenging, which suggests the need for describing these terms with more clarity and ease.
These study findings provide theoretical implications by confirming previous research which indicates that K–12 educators often lack assessment literacy and experience difficulty understanding certain terms in score reports (Lukin et al., 2004; Zapata-Rivera et al., 2010; Zwick et al., 2014). The current study found that K–12 EL educators struggle with conceptually understanding scale scores and confidence bands. This could be partly because educators frequently refer to proficiency level scores, which are used to make high-stakes decisions such as EL identification and reclassification, as required by federal and state policies. It may also be partly because proficiency level scores are easily connected to the Standards framework in which the assessment is situated, providing straightforward score interpretation for educators. However, scale scores are less intuitive for educators and are more challenging to connect to the Standards framework, whereas confidence bands are seen as overly technical and confusing by many educators, making it challenging to comprehend the terms.
In addition, the findings help fill a gap in the literature regarding K–12 EL educators’ score report interpretation and use of resources. During the interviews, educators suggested including more information in the score reports, such as information about student growth. This suggests that educators are not only interested in measures of student performance at a single point in time, but whether the students are making progress over time. Regarding the use of score report resources, due to time pressures, educators did not have sufficient time to utilize resources made available to support score interpretation. In addition, the existing materials may have been too lengthy and complex for educators. This reveals a dilemma in promoting score report interpretation: educators need more resources for score report interpretation, but they also lack the time to read the materials. These findings suggest that the resources offered to educators should be easily accessible and digestible.
Results also provide practical implications for further supporting K–12 EL educators’ assessment literacy by indicating the areas in which educators need more support. Based on the feedback from the educators, the following suggestions need to be considered for future English language proficiency assessment score reports:
Clarify technical terms, such as scale score and confidence band, to enhance interpretation of score report information. Educators experienced difficulty in interpreting scale scores and confidence bands owing to their limited understanding of these terms. These terms should be explained concisely, as educators often lack the time to refer to supporting resources.
Include information on student growth to monitor students’ language development in comparison with previous years. Educators wanted to know how their students were performing, but also if they were displaying growth in comparison to previous years. Schools and districts often keep previous years’ reports on file or create a database for this purpose. However, in order to monitor student growth more easily, educators requested that this information be on the score report itself.
Improve the delivery timeline of score reports to reduce the gap between the time of test administration and delivery of the scores to stakeholders. Currently, for many large-scale English language proficiency assessments, there is a relatively large gap (up to a four months) between the time of test administration and delivery of the scores to stakeholders. This is the case for ACCESS mainly because human raters, not computers, score students’ speaking and writing responses. Learners, especially at a young age or with lower English proficiency, may make considerable gains during a short time period. Therefore, when score reports are delivered months after the students took the test, their scores may no longer accurately reflect their English language proficiency level.
One limitation of the study was not having detailed information regarding the educators, including the type of EL instruction they provide, such as pull-out or push-in instruction, or the type and amount of training participants received in relation to score report interpretation (e.g., online webinars and professional learning opportunities). Although most participants had at least a master’s degree in education or a related field, not all master’s programs offer courses in language assessment. Therefore, it is unclear how much previous assessment literacy educators had and to what extent their training influenced their test score interpretation.
Another limitation was the sampling method to recruit survey participants. Owing to the difficulty of recruiting a large number of respondents for the survey study in Phase 1, we used a convenience sampling method by distributing the online survey link to educators on the WIDA listserv. Although this method led to the collection of more than 1400 samples, it weakens the generalizability of the findings. Using a stratified sampling method may have led to better data collection.
Study findings provide implications for future research. Research team members specifically examined the assessment literacy required by K–12 EL educators for interpreting score reports from ACCESS, which is an annual standardized English language proficiency assessment. Although it is crucial to interpret accurately large-scale language assessment scores in order to examine EL students’ English language proficiency and their language growth, other types of assessment literacy may be needed. Educators engage in various assessment-related activities in K–12 EL settings, such as creating classroom-based assessments to measure and interpret students’ in-class performance. These specific assessment literacy skills are crucial in everyday class settings, yet educators may not receive sufficient guidelines or support in order to develop them. Therefore, future studies need to investigate further the various assessment literacy skills (necessary for large-scale and classroom-based assessments) that are relevant to K–12 EL settings. In addition, the field may benefit from future research on educators’ assessment literacy, initiated by independent researchers or state education agencies. Moreover, future studies should examine how educators’ assessment literacy allows them to use the score report information. Although this study reports findings on the degree of educators’ assessment literacy in relation to interpreting score report information, educators do much more with score reports, such as making various decisions for EL students. Helping educators to interpret more effectively the technical data on their students will lead to better decisions for those students and, ultimately, better educational outcomes.
Footnotes
Appendix
Coding scheme for analyzing interview data.
| Category | Description or (sub)categories | Examples |
|---|---|---|
| Interpretation of score report information | Understanding of each section of the score report | Everything is very clear to me. I understand all of the information conveyed on the score report. |
| Helpful information from the score report | I find the proficiency level information to be most helpful. | |
| Unclear information from the score report | I’m not sure what scale score means. | |
| Use of score report resources | The type of resources used | When interpreting the test scores, I usually refer to the Interpretive Guide. |
| Suggestions for improvement | Information that should be added to or deleted from the score report | Not sure if confidence band information is really necessary. |
| Suggestions for improving the score report | Would it be possible to make the font a little larger? |
Authors’ Note
Akira Kondo is now affiliated with Rakuten, Inc., Japan and Carsten Wilmes is now affiliated with American Institutes for Research, USA.
Declaration of conflicting interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
