Abstract
We used multivariate generalizability theory to examine the reliability of hand-scoring and automated essay scoring (AES) and to identify how these scoring methods could be used in conjunction to optimize writing assessment. Students (n = 113) included subsamples of struggling writers and non-struggling writers in Grades 3–5 drawn from a larger study. Students wrote six essays across three genres. All essays were hand-scored by four raters and an AES system called Project Essay Grade (PEG). Both scoring methods were highly reliable, but PEG was more reliable for non-struggling students, while hand-scoring was more reliable for struggling students. We provide recommendations regarding ways of optimizing writing assessment and blending hand-scoring with AES.
Writing skills are critical to students’ success in K–12 and postsecondary settings (National Commission on Writing, 2003). Yet roughly two-thirds of students in Grades 4, 8, and 12 fail to achieve grade-level writing proficiency (National Center for Educational Statistics, 2012; Persky et al., 2002). Accordingly, present educational reform efforts emphasize writing proficiency, as indicated by dedicated writing standards within the Common Core State Standards and similar state-specific standards for English language arts (Graham & Harris, 2015; Shanahan, 2015). These standards include expectations that students will demonstrate proficiency in writing for multiple purposes and in multiple genres, such as narrative, informational, and persuasive/argumentative writing. Moreover, as of 2015, 46 states have adopted summative writing performance assessments (Behizadeh & Pang, 2016).
Within this broader context of increased accountability for students’ writing proficiency, teachers need access to reliable and efficient classroom-based writing assessments. Regular assessment of writing is likely to benefit elementary-aged children by improving the effectiveness of writing instruction within the context of multitiered systems of support (Graham et al., 2015; Philippakos & FitzPatrick, 2018). Writing assessment and feedback have also been shown to positively affect student writing in meta-analysis (Graham et al., 2011a). Moreover, writing scores may be used to adjust instruction, identify students who may benefit from supplemental intervention in multitiered systems of support, monitor students’ progress, or determine whether a student qualifies for special education services. Some of these decisions require educators to make relative judgments, in which a student’s performance is evaluated relative to that of their peers (e.g., screening), or absolute judgments, in which a student’s performance is evaluated relative to established performance criteria (e.g., summative evaluation).
However, writing assessment is quite challenging for teachers because it is difficult to assess reliably and efficiently. Reliability threats arise from multiple sources of error (e.g., imprecision), including rater effects, such as rater leniency, severity, and drift (see Wind et al., 2017); the prompt topic (Lee & Anderson, 2007); the prompt genre (Bouwer et al., 2015); interactions between those facets (Li, 2017); and other unaccounted-for sources of error (e.g., occasion or administration errors; Webb & Shavelson, 2005). Writing assessment is also time-consuming and laborious. Frequently, teachers cite the challenge of assessing writing as a barrier to increasing the amount of writing instruction and feedback students receive (Applebee & Langer, 2009, 2011; Graham, 2019).
Due to these sources of measurement error, findings from application studies of generalizability theory (G theory; Brennan, 2001; Cronbach et al., 1963; Cronbach et al., 1972; Shavelson & Webb, 1991) suggest that multiple writing prompts and raters are necessary to improve reliability and make accurate relative or absolute judgments for students (e.g., Bouwer et al., 2015; Chen et al., 2007; Graham et al., 2016; Kim et al., 2017; Schoonen, 2005, 2012). For instance, findings have indicated that three to five prompts each scored by four raters are needed to achieve a reliability of .80 for students in Grade 9 (Chen et al., 2007), seven prompts scored by two raters (or six prompts scored by four raters) are needed to achieve a reliability of .90 for students in Grades 3 and 4 (Kim et al., 2017), and eight to 14 prompts scored by two raters are needed to achieve a reliability of .90 for struggling writers in Grades 2 and 3 (Graham et al., 2016). There are different ways to define a struggling writer, but here we consider struggling writers to be students who fall at or below the 25th percentile for their age or grade level, placing them at risk for writing delays. Importantly, reliable judgments for struggling writers may require more extensive assessment procedures due to smaller differences between writers in that restricted ability range. Regardless, following such recommendations would require between 12 and 28 teacher ratings of student writing to make reliable relative or absolute decisions.
Unfortunately, such recommendations are not feasible or practical. Beyond the sheer volume of writing assessments for students, the time costs for teachers associated with hand-scoring the resulting responses are too great (time costs refers to the amount of time human raters must spend on scoring that may otherwise be used for instructional activities [e.g., planning or teaching]). More than 55% of teachers report limiting the amount of writing they have their students do because of the time it takes to grade it (Graham et al., 2014). Thus, it is important to continue to research and identify feasible methods that teachers may use to make accurate, generalizable inferences about students’ writing ability. Recent research on automated essay scoring (AES) shows promise in this regard.
Benefits and Challenges of AES
AES relies on (a) natural language processing (NLP) to identify features of student writing (e.g., vocabulary or syntax) that correlate with writing quality scores assigned by trained human raters and (b) machine learning algorithms to maximize prediction of human rater scores from those NLP features. AES returns scores that can distinguish students with a range of writing ability immediately, removing the time cost associated with rater training and hand-scoring. This automation makes it possible to score an entire classroom of students—or even an entire school or district—in a few seconds. Also, unlike human raters, AES applies its scoring algorithms consistently: The weighting of certain features of writing quality remains constant. The same essay always receives the same score, and scores retain their meaning over time. Consequently, AES engines are highly reliable (see Shermis, 2014).
For instance, a recent G theory study (Wilson et al., 2019) of an AES system called Project Essay Grade (PEG; Page, 2003) finds that educators seeking to make reliable (≥ .80) low-stakes decisions (e.g., instructional groupings) about students’ writing performance across three commonly instructed genres of writing (narrative, informative, persuasive) would need to administer one 30-minute prompt per genre (i.e., three total) for non-struggling writers and two 30-minute prompts per genre (i.e., six total) for struggling writers. Making high-stakes decisions (e.g., grade retention or referral to special education; reliability ≥ .90) would require administering two 30-minute prompts per genre (i.e., six total) for non-struggling writers and between four and five prompts per genre (i.e., 12–15 total) for struggling writers. Compared to the numbers of prompts and raters, and the total administration and scoring time required when relying on hand-scoring (see above), the results reported by Wilson et al. (2019) are quite promising.
Another benefit of AES is the ability to supplement summative scoring with diagnostic and formative scoring. AES systems commonly report scores that describe a student’s performance in an overall and a nuanced manner by assessing dimensions (i.e., traits) of writing quality (Sinharay et al., 2019; Wang et al., 2020; Wilson et al., 2017). Also, many AES systems complement scoring with automated feedback to help students revise—an application referred to as automated writing evaluation (AWE)—affording AES an instructional benefit.
However, some challenges accompany the use of AES. First, AES does not read and understand text as a human does (Attali, 2015). NLP features approximate elements of writing quality, such as “style,” and key features of a genre (e.g., a thesis or topic statement), but these approximations are not the same as a teacher evaluating those elements directly. Consequently, AES and human raters do not measure the writing construct identically (see Deane, 2013), and instructionally relevant information may be lost with exclusive reliance on AES.
Second, rater effects may end up being “baked into” the AES system if the underlying data used to develop and train the AES model are insufficiently reliable (Wind et al., 2017). If the training data underrepresent certain scores, the resulting AES model may be reliable overall but lack reliability at the tails of the scoring range (Raczynski & Cohen, 2018). As Bridgeman (2013) writes, “Just because a[n] [AES] system works well on average, there is no guarantee that it will work equally well in all population subgroups” (p. 227). Indeed, a prior G theory study of AES has shown that twice as many writing prompts were needed to obtain reliable estimates of struggling writers as compared to non-struggling writers (Wilson et al., 2019), although this result also is aligned with results of G theory studies of human-scored raters (see previous section).
Similarly, any bias in the underlying training data that may privilege certain racial, ethnic, or language backgrounds may also end up being baked into the AES system. Bias may be introduced via biased rater scores or the use of training data that are insufficiently representative of the target population. For instance, most automated scoring systems within AWE are built to accurately evaluate native English speakers’ writing ability, but not those of English language learners (ELLs; e.g., Hoang & Kunnan, 2016), despite AWE being used with that group. If biased, reliance on AES may result in promulgating educational inequities for minoritized students.
Benefits and Challenges of Hand-Scoring
Hand-scoring is regarded as the gold standard for writing assessment (Powers et al., 2015) because humans are the best judge of whether a student has achieved a rhetorical purpose (e.g., to entertain, to persuade, or to inform), content accuracy, and elements of writing that are difficult to quantify, such as style, humor, and creativity (see Deane, 2013). Furthermore, educators can use a student’s strengths and weaknesses identified during hand-scoring to inform instruction and feedback. However, hand-scoring threatens reliability, even with well-defined and rigorous rater-training procedures (Bridgeman, 2013; Wind & Walker, 2019). These reliability threats include:
halo effect, in which a single dimension of writing quality (e.g., conventions) improperly sways a rater’s judgment (see A. C. Johnson et al., 2017)
rater leniency and severity, in which a human rater systematically assigns higher or lower scores than warranted (see Wind, 2018, 2020)
narrowing of the scoring range (i.e., rater centrality), in which raters overly rely on the middle categories of the scoring range (see Wind, 2018, 2020)
rater drift, in which raters deviate from their application of a rubric over time (see Lottridge et al., 2013; Palermo et al., 2019)
One reason for rater effects is the format and subsequent interpretation of the scoring rubric (Lottridge et al., 2013). Rubrics that require raters to consider and weight many different aspects of writing quality are subject to inconsistent interpretation and application. Such inconsistencies threaten reliability and validity. Additionally, hand-scoring is time-consuming. Time is required for training, scoring, and resolving disagreements, which inhibits efficient decision making and makes writing assessment an onerous task.
Unfortunately, efforts to alleviate time costs compound existing problems with reliability and construct validity. For instance, using a single rater instead of multiple raters is more efficient but increases the likelihood of reliability threats. Educators could also limit the amount of writing they assign to reduce time burdens, but this option results in fewer writing samples available for decision making. Finally, educators may score shorter writing samples by using curriculum-based measurement (CBM) methods of writing assessment that limit students to 1 minute for planning and 3 minutes for composing. However, this option results in a narrowing of the writing construct and, unlike extended writing tasks, precludes opportunities to evaluate the extent to which students demonstrate control over key writing processes (e.g., reviewing, revising, or editing) regarded as central to writing proficiency (Graham, 2018).
Study Purpose
As the prevalence of AES grows, it is incumbent on researchers to identify ways to leverage the strengths of hand-scoring and AES to improve the feasibility and reliability of writing assessment. Accordingly, we seek to answer the following research questions:
How efficiently can we obtain a reliable estimate of writing ability when using hand-scoring and AES via a system called PEG for students in Grades 3–5?
Are results consistent for struggling and non-struggling writers across multiple genres?
We define efficiency in terms of the number of writing samples and scores needed to achieve specified reliability thresholds (i.e., .80 and .90). In our view, educators first need to understand how many writing samples and scores are needed before considering other dimensions of efficiency, such as training, administration, and scoring time. Although educators also need to consider cost efficiencies, such as material costs (e.g., AES licenses) and labor costs (e.g., costs associated with training and hand-scoring), we consider the present study to be an important first step toward examining efficiency and cost-effectiveness more broadly.
We used G theory to answer our research questions. G theory advances classical test theory by incorporating the algorithm of analysis of variance to disentangle the impact of all possible measurement errors. Thus, it is ideally suited for determining the contributions of different sources of error variance to observed score variance and how best to optimize a measurement procedure to achieve a given reliability threshold. We conducted hand-scoring following recommendations for improving the reliability of scoring (see Graham et al., 2011a). We conducted AES scoring by using the PEG system, a formative AES system available to educators via the MI Write AWE system.
The present study advances current understanding in two critical ways. First, prior G theory research on writing performance assessments has exclusively focused on either hand-scoring (Bouwer et al., 2015; Chen et al., 2007; Graham et al., 2016; Kim et al., 2017; Schoonen, 2005, 2012) or AES (Wilson et al., 2019). The present study applies G theory to compare both scoring methods applied to the same corpora of student writing.
Second, the present study is the first to apply multivariate G theory to simultaneously scrutinize the reliability of the two different scoring methods for assessing student writing in three different genres. Prior G theory studies have either confounded genre and prompt topic by eliciting only one prompt per genre (Graham et al., 2016; Schoonen, 2005) or treated genre as a fixed facet within a univariate G theory analysis (Bouwer et al., 2015; Wilson et al., 2019). Interestingly, prior univariate G theory studies indicate that genre contributes minimally to observed score variance (see Bouwer et al., 2015; Wilson et al., 2019). However, when genre is treated as a facet within a univariate G theory analysis, it is not possible to examine whether there are systematic differences in measurement precision across distinct genres. The use of multivariate G theory in the present study afforded this possibility.
Method
Participants
This study draws from the sample described in Wilson et al. (2019). The present study involved additional measures and different analyses and addressed different research questions.
The original sample included a total of 570 students (~190 students per grade) in Grades 3–5. In the original sample, struggling writers were defined as students falling below the 25th percentile using a standard score based on the aggregate of three measures of transcription skill: a 1-minute handwriting fluency task (inter-rater agreement = 97%) and the Sentence Combining and Sentence Building subtests of the Wechsler Individual Achievement Test, 3rd edition (test-retest coefficients for Grades 3, 4, and 5 = .86, .88, and .85, respectively). Non-struggling writers were identified based on scores at or above the 30th percentile of the aggregate standard score. The 25th percentile is commonly used to define risk status, and similar percentile splits have been used in prior research comparing groups of differing levels of risk (Etmanskie et al., 2016; Speece & Ritchey, 2005). We identified groups this way rather than based on demographic factors to align with academic screening methods used in multitier systems of support (Philippakos & FitzPatrick, 2018).
In the present study, we aimed to select 120 students with complete data—40 students per grade (20 non-struggling writers and 20 struggling writers * 3 grades = 120 students) who composed two writing prompts in three genres (six prompts total)—as this would provide for feasibility of hand-scoring (120 students * 6 essays = 720 essays) and enable the use of our desired analytic method, multivariate generalizability theory, and software (mGENOVA; Brennan, 2010) that requires complete data. However, due to the presence of missing prompt data, we were unable to select 40 students each in Grades 4 and 5. Thus, our final sample included 113 students: Grade 3 = 20 non-struggling writers and 20 struggling writers; Grade 4 = 19 non-struggling writers and 19 struggling writers; and Grade 5 = 20 non-struggling writers and 15 struggling writers.
Table 1 presents demographics of the sample. Free or reduced-price lunch status was not available to report. Chi-square tests indicated non-statistically significant differences across Grades 3–5 with respect to gender (χ2 = 0.67, p = 0.716), race (χ2 = 2.31, p = 0.889), ELL status (χ2 = 1.86, p = 0.394), and special education status (χ2 = 2.86, p = 0.239).
Sample Demographics
Note. ELL = English language learner. SPED = students receiving special education services via an individualized education plan.
Adequacy of Sample Size
There are no clear guidelines regarding the adequacy of sample size when conducting G theory analyses because G theory assumes persons as a particular population of interest (Brennan, 2001). In other words, G theory provides a methodology to describe the variation of different sources of error in a replicable measurement procedure (i.e., universe of admissible observations) within a specific population; in G theory, the word population is used for the object of measurement (Brennan, 2010), which are persons (i.e., students) in the current study. Therefore, concepts of statistical power and precision (i.e., confidence intervals) do not apply because no hypotheses are being tested. G theory is used simply to describe the generalizability of scores under specific, replicable measurement procedures (i.e., consistency). To provide contextual relevance for this study, interventionists usually have fewer than 15 students in a writing group. Moreover, sample sizes of several published G theory analyses of writing assessment data are similar to ours [e.g., Barkaoui, 2007 (n = 16); Eckes, 2011 (n = 21); Yamanishi, 2005 (n = 20)] and have withstood critique in a recent systematic review (In’nami & Koizumi, 2016). Finally, our sample included six writing samples per student; more than two observations per student improves adequacy of the sample size in G theory analyses (Webb et al., 1988). Thus, our sample sizes were sufficient for answering our research questions.
Measures
Writing Prompts
On six successive days, students composed responses to six randomly assigned prompts, (two in each of three genres: informative, narrative, and persuasive. Prompt topics were randomly assigned from a pool of six prompts per genre, or 18 total. To control for order and prompt effects, genre order was counterbalanced across classrooms. Prompt topics were developed by the third author with a team of graduate students certified in elementary education and selected based on the likelihood that students would possess sufficient background knowledge to respond. Example prompts included (a) Informative: “Think about your favorite animal. Teach your reader all about this animal”; (b) Narrative: “Write a story about what you would do if you could fly”; and (c) Persuasive: “Should kids get to pick their own bedtime?” The full list of prompts is available in Appendix A of the online supplementary materials.
Prompts were administered by trained research assistants who read from a set of standardized administration directions that explained the purpose of the writing activity and provided some information about the purpose and key features of the genre. Directions for administering prompts are presented in Wilson et al. (2019). All administration sessions (n = 217) were audio recorded. A random sample of 30% of the recordings was evaluated for fidelity of assessment administration, which we calculated as the percentage of directions read correctly and correct time provided for composing (i.e., 30 minutes). Fidelity was high: 98%.
Students composed their responses to the prompts by hand. Research assistants transcribed the responses verbatim into Word documents, maintaining any errors in spelling, grammar, or punctuation present in the original. Accuracy of transcription was evaluated on a random sample of 30% of the writing prompts and calculated as the percentage of correctly transcribed elements, including spelling, capitalization, punctuation, paragraph breaks, and punctuation marks. Transcription accuracy was high: 98.51% (SD = 1.54%).
AES via PEG
AES scores were generated by PEG (Page, 2003), which is the AES system used in the AWE system called MI Write, formerly known as PEG Writing. PEG scores student writing based on genre and grade-level band-specific scoring models rather than prompt-specific scoring models, enabling PEG to provide reliable scores for prepackaged and custom prompts. PEG uses 15 different scoring models that are modified from a core scoring model to score three genres (informative, narrative, and persuasive) in five grade bands: Grades 3–4, 5–6, 7–8, 9–10, and 11–12. Because PEG scores by grade-level band, similar scores across grade-level bands do not indicate equivalent writing proficiency. PEG scores are intended to support formative writing assessment (i.e., low-stakes decisions).
PEG provides ratings on a 1–5 scale for each of six writing quality traits: ideation, organization, style, sentence structure, word choice, and conventions (see Coe et al., 2011). Additionally, PEG provides an Overall Score to represent students’ writing performance holistically. The Overall Score (hereafter, “PEG Score”) is formed as the sum of the six traits (range = 6–30), with scores reported to a tenth of a point. Consistent with Wilson et al. (2019), we used the PEG Score rather than trait scores, due to collinearity among the trait measures and our desire to evaluate students’ writing performance holistically. Quadratic weighted kappa (QWK) ratings evaluating the consistency of PEG with human ratings applied to a held-out test set for the six scoring models used in the present study were high: Average QWK for the three genre models for Grades 3–4 was 0.87 (SD = 0.03) and 0.85 (SD = 0.04) for the three genre models for Grades 5–6. The average QWK across all six models was 0.86 (SD = 0.04).
Hand-Scoring
The same electronic essays scored by MI Write were also hand-scored by human raters to obtain a holistic writing quality score. We elected to score holistically because it required less time than analytic scoring and has been found to be more reliable (Graham et al., 2011b). Prior research has also rarely shown distinctions among analytic scores, indicating that they generally measure the same thing (Gansle et al., 2006).
Training and Procedures
Four master’s student graduate research assistants (GRAs) scored the writing samples. All GRAs had completed undergraduate teaching programs prior to their master’s program and held elementary (K–6) teaching certifications (three also held a special education teaching certification), although none had professional teaching experience. The GRAs were enrolled in a graduate seminar in writing research during the semester in which they scored the writing samples for this study (although this was not part of the training).
During training, the second author provided the GRAs with an overview of the three genres to be scored and the six writing prompts used for each genre. The team discussed the purposes and potential features expected for each genre. For example:
narrative writing prompts: first-person story, student as a character in their own story, settings, events, problems, outcomes, and emotions
persuasive writing prompts: topic sentence, reasons, elaborations and/or explanations to support the reasons, counterarguments with refutations, and a conclusion
informative writing samples: topic sentence, descriptive facts about the topic, elaborations and/or examples that provide additional factual information
The GRA team also discussed how to make holistic scoring decisions and the variety of criteria to consider. In addition to genre elements, the team discussed grammar, organization, spelling, punctuation, and word choice as factors to consider when scoring. Considerable discussion was given to reducing scoring biases, as identified by Graham et al. (2011a). For instance, we explained that we decided not to correct for spelling prior to scoring for this project, to be consistent with the way that PEG scores writing. This instruction allowed GRAs to consider spelling while trying to avoid unduly emphasizing it in the score.
The second author then provided an overview of the holistic scoring system to be used and modeled it using exemplars (not from the study sample). Similar to procedures used by Penny et al. (2000a, 2000b) and R. L. Johnson et al. (2003), the GRAs learned to score the writing in two steps. First, they used anchor papers to score the writing on a 1–7 scale, and then they reviewed all the papers within each score category to determine whether each writing sample should stay in the category or receive a plus or minus (e.g., 5, 5+, or 5–). A half point (.5) was added to scores with a plus and subtracted from scores with a minus. Thus, scores with a half point were awarded with a plus or minus (e.g., 4.5 points could be awarded using 4+ or 5–).
We trained GRAs to score the writing samples using anchor papers specific for each grade level (e.g., a score of 7 in Grade 5 was not equivalent to a score of 7 in Grade 4). This distinction was important for calibrating the human raters to PEG, which was, in turn, trained using human raters with unique grade-band scoring rather than scoring on a continuum across grade levels. To calibrate the rater scoring with PEG, we used “range finders” for the GRAs to use to practice scoring during training (a range finder is an essay with a known score that is not part of the sample of essays to be scored and that is used to assist in rater calibration, comparable to the use of anchor papers [see Osborn Popp et al., 2009]). For each genre and grade, 36 essays were selected as range finders (324 range finder essays total: 36 essays * 3 genres * 3 grades). The third author selected the range finders so that the 36 essays were distributed evenly across “low,”“middle,” and “high” ranges of the PEG Score observed for that genre and grade (i.e., 12 range finder essays per category). Essays in the “low” range received a PEG Score ranging from 6–9. Essays in the “Middle” range received a PEG Score ranging from 10–14. Essays in the “high” scored ≥ 15.
During training, the GRAs practiced scoring one genre at a time. For each genre, the GRAs individually scored a subset of writing samples and then compared scores with each other. When scores differed by 1 or more points on the scale, the GRAs discussed the scores. Specifically, they discussed their rationale for providing their score, including the writing components that influenced their decision. Finally, they compared their scores to the range finder scores from PEG. Although PEG used a different scoring scale, this comparison allowed the GRAs to see whether their scores matched the low, medium, and high ranges from PEG. When they did not match, the GRAs discussed potential reasons for differences. GRAs were retrained if less than 80% of scores were within 1 point; retraining occurred in only one instance.
The GRAs trained for a single genre, scored all the samples for that genre, and then trained and scored for the next genre. The GRAs mixed and scored writing samples for all prompts within a genre to reduce the influence of topic. They followed the same scoring procedures as completed in training (including the use of range finders), except that the GRAs did not compare their scores with a second rater. The GRAs wrote the final score on a copy of each writing sample and then entered their scores into the database. Data entry was checked for reliability by a second GRA to ensure that all data were entered correctly.
Time for Training and Scoring
The GRAs received a total of 8 hours of training: 2 hours of general training, followed by approximately 2 hours of training specific to each genre. All four GRAs then scored 678 total papers (113 students * 6 writing samples). The raters scored the writing samples over a 3-week period. We estimated that it took each rater approximately 35 hours to score all the papers, or approximately 3 minutes to score each paper. This included the time to read the paper, compare it with the anchor papers, make a judgment, revise the judgment during the scoring process, and take mental breaks. Overall, approximately 140 hours total across the four GRAs were required to score all the writing samples.
Inter-Rater Reliability
Inter-rater reliability was calculated among the four raters using the intra-class correlation (ICC), a method of estimating reliability and consensus in rating designs with more than two raters (Hallgren, 2012; Shrout & Fleiss, 1979). Our rating design was fully crossed (all raters scored all writing samples), and we considered the four raters to be randomly sampled from a population. Thus, we used a two-way random-effects model to calculate the ICC for absolute agreement. Results of the ICC models are presented in Appendix B of the online supplementary materials. The ICC for the average of all raters was ≥ .94 (range = .94–.98) for each of the six scoring tasks for each grade level, indicating a very high level of inter-rater reliability.
Analytic Method
We used multivariate G theory to answer our research questions. Within multivariate G theory, items (e.g., prompts) are nested within a fixed facet (e.g., genre), and levels of the fixed facet (e.g., different genres) are treated as separate dependent variables. Treated in this way, the model specification of random facets remains the same across each level of the fixed facet and statistically connected through covariance components to obtain a multivariate design (Brennan, 2010). Therefore, in the present study, multivariate G theory has the advantage of estimating complete random facet models and examining differences and correlations between genres. Importantly, because we used multivariate G theory to examine the variability of writing scores across groups, and not to compare the equivalence of scores across groups, issues of measurement invariance were not applicable.
For the hand-scoring method, students (p) were rated holistically by four raters (r) on two writing prompts (i) within each genre. For the AES rating method, students (p) were rated by PEG on two writing prompts (i) within each genre. For both rating methods, the three types of genre (informative, persuasive, and narrative) were treated as the fixed facet levels in a table of specifications. Figure 1 in Appendix C of the online supplementary materials illustrates the parallel universes of admissible observations of the hand-scoring method, p• x i○ x r•, and the AES method, p• x i○. (A superscript filled circle [•] means that the facet is crossed with the linked facet [i.e., fixed facet], and an empty circle [○] means that the facet is nested within the linked facet.)
For the hand-scoring method, because the four raters rated all students on all the prompts, the object of measurement (i.e., student) is defined as a full matrix
For the G studies, we calculated variance matrices, the percentage of variance across facets within each genre, and the covariance and correlations of writing scores between genres. We calculated this separately for student writing ability status (non-struggling writers and struggling writers), rating method (hand-scoring and AES), and grade level (Grades 3, 4, 5). In total, we estimated 12 G studies.
For the D studies, we calculated G and Phi coefficients within each genre as well as composite G and Phi coefficients. We calculated these separately for each scoring method by writing ability status and grade level. G coefficients describe the degree of reliability for making relative decisions. Phi coefficients describe the degree of reliability for making absolute decisions. Both coefficients range from 0.00–1.00, with 1.00 being perfect reliability.
Given that the measurement model for hand-scoring (4 raters * 2 prompts = 8 scores per genre) and AES scoring (1 rater [PEG] * 2 prompts = 2 scores per genre) was different, we estimated additional D studies wherein we statistically manipulated the number of raters and prompts for the hand-scoring method and the number of prompts for the AES method to facilitate more direct comparisons of the reliability between the two scoring methods. Specifically, we estimated the following additional D studies:
For the PEG scoring model: We estimated G and Phi coefficients after increasing the number of prompts per genre to eight. This allowed for matching PEG’s reliability to that produced based on the measurement design for hand-scoring (4 raters * 2 essays).
For both the PEG scoring model and the hand-scoring model: We adjusted the number of prompts and raters, such that both models would output four scores per genre. For PEG scoring, we increased the number of prompts to four per genre. For the hand-scoring model, we reduced the number of raters to two and maintained two prompts per genre.
Results
Descriptives
Table 2 includes descriptive statistics for the two scoring methods across grades, genres, and ability status. As expected, struggling writers scored lower than non-struggling writers for both scoring methods and across all grades and genres. Within each scoring method, students’ mean performance across genres was quite similar within a grade level. Finally, the full range of the human holistic scoring scale was used, as evidenced by Grade 5 non-struggling writers obtaining scores of 7.00 (ceiling) and Grade 3 struggling writers obtaining scores near the floor of the scale (i.e., near 1.00). However, the upper range of the PEG scale (ceiling = 30.0) was never reached; the maximum observed value was 21.45.
Descriptive Statistics
Note. Hand-scoring was conducted using a holistic score. AES (automated essay scoring) scoring was conducted using the PEG system. Range of holistic score = 0.0–7.0. Range of PEG Score = 6.0–30.0. Descriptive statistics for the hand-scoring method were derived using the average of the four raters’ scores across the two prompts within a genre (i.e., average of eight ratings). Descriptive statistics for the PEG rating method were derived using the average of the PEG Score across the two prompts within a genre (i.e., average of two ratings). M = mean. SD = standard deviation.
Correlations Between Scoring Methods
Before examining results of the G and D studies, we inspected the correlations between the average scores per student for each genre obtained from hand-scoring (i.e., eight scores per genre) and AES (i.e., two scores per genre) for non-struggling and struggling writers in each of the three grade levels. We did this to evaluate the similarity with which human raters and AES rank-ordered students based on their performance despite scoring on different scales. Table 3 includes these correlations; 15 of the 18 correlations were ≥ .70, and the remaining three correlations—all for performance within the narrative genre—ranged from .55–.69.
Correlations Between Hand-Scoring and AES
Note. Correlations were reported for the average score of the four human raters across two prompts (i.e., average of eight scores) and the average score of PEG across two prompts (i.e., average of two scores). *p≤ .05; **p≤ .01; ***p≤ .001.
Results of G Studies for Hand-Scoring
Non-Struggling Writers
Table 4 presents results of the multivariate G studies for non-struggling writers, scored using four raters for two prompts per genre (i.e., eight scores per student per genre). Student, the object of measurement (
Variance and Covariance Matrices for Non-Struggling Writers: Hand-Scoring Using Holistic Ratings
Note. Bold font indicates standardized correlation coefficients.
Table 4 also includes correlations of performance between genres within grade levels (indicated by bold font). Correlations of students’ average performance across genres ranged from moderate to very strong in Grade 3 (r = .77–.86), moderate to very strong in Grade 4 (r = .78–.90), and strong to very strong in Grade 5 (r = .80–.92).
Struggling Writers
Table 5 presents the results of the multivariate G studies for struggling writers scored on two prompts per genre by four raters (i.e., eight scores per genre). As with the non-struggling writers, the object of measurement (
Variance and Covariance Matrices for Struggling Writers: Hand-Scoring Using Holistic Ratings
Note. Bold font indicates standardized correlation coefficients.
As shown in Table 5, correlations of students’ average performance between genres within grade levels ranged from moderate to very strong. In Grade 3, correlations ranged from .79–.87. In Grade 4, correlations ranged from .81–.92. In Grade 5, they ranged from .65–.91.
Results of G Studies for AES
Non-Struggling Writers
The results of the multivariate G studies for non-struggling writers scored by PEG using two prompts per genre (i.e., two scores per student per genre) are presented in Table 6. The largest proportion of variance was attributed to the student (
Variance and Covariance Matrices for Non-Struggling Writers: AES Scoring Using PEG
Note. AES scores were generated by PEG. Bold font indicates standardized correlation coefficients.
As shown in Table 6, Correlations of students’ average performance between genres ranged from moderate to strong in Grade 3 (r = .61–.82) and from strong to very strong in Grade 4 (r = .82–.91). Correlations between genres were very strong in Grade 5 (r = .92–.97).
Struggling Writers
Table 7 includes the results of the multivariate G studies for struggling writers scored by PEG for two prompts per genre. In Grades 3 and 5, the object of measurement (
Variance and Covariance Matrices for Struggling Writers: AES Scoring Using PEG
Note. Bold font indicates standardized correlation coefficients.
As shown in Table 7, correlations of students’ average performance between genres differed across grade levels. In Grade 3, correlations were moderate (range r = .67–.75). In Grade 4, there were perfect correlations (r = 1.00) between students’ average performance in informative and narrative genres and in informative and persuasive genres, but only a moderate correlation between their average performance in narrative and persuasive genres (r = .51). Likewise, in Grade 5, there were perfect correlations between students’ average performance in informative and persuasive genres and in narrative and persuasive genres, but only moderate correlations between their average performance in informative and narrative genres (r = .38).
Reliability of Measurement: G and Phi Coefficients
Table 8 includes the G and Phi coefficients based on the multivariate G studies just described. Respectively, G and Phi coefficients detail the level of reliability for making relative (rank-ordering) or absolute decisions (criterion referenced). Generally, .80 is considered a desirable level of reliability for making low-stakes decisions, and .90 is considered desirable for making high-stakes decisions (Nunnally, 1967). Table 8 reports G and Phi coefficients individually for different genres as well as composite G and Phi coefficients, based on a student’s average score across all genres. For each grade level, composite coefficients for hand-scoring are based on the average of 24 scores (8 ratings per genre * 3 genres); for AES scoring, they are based on the average of six scores (2 ratings per genre * 3 genres).
G and Phi Coefficients for Original Measurement Design
Hand-scoring
As shown in Table 8, for all three grades and for non-struggling and struggling writers, the composite G and Phi coefficients exceeded .90. With one exception (Phi for Grade 4 narrative), G and Phi coefficients associated with individual genres (eight scores per genre) exceeded the .80 threshold and exceeded the .90 criterion for high-stakes decisions in many instances.
AES
As shown in Table 8, for non-struggling and struggling writers in all three grades, the composite G and Phi coefficients based on the average of six PEG ratings exceeded .80. However, those coefficients were higher for non-struggling writers: all composite G and Phi coefficients for non-struggling writers met or exceeded .90, except for the composite Phi coefficient for Grade 3 (0.88). That criterion was met for struggling writers only in Grade 3.
For non-struggling writers, G and Phi coefficients associated with individual genres (i.e., two scores per genre) met or exceeded the .80 criterion for all grades, with the exceptions of Grade 3 informative (G = 0.77; Phi = 0.73) and Grade 3 persuasive (G and Phi = 0.76). For struggling writers, G and Phi coefficients varied across grades and genres. For instance, in Grade 3, reliability was moderate for scoring writing in the persuasive genre (0.67) but generally acceptable for scoring writing in the other genres. In Grade 4, reliability was weak for informative writing (0.45), moderate for persuasive writing (0.67), and acceptable for narrative writing (0.80). In Grade 5, reliability was weak for narrative writing (0.54) and persuasive writing (0.58) and near criterion for informative writing (0.78).
D Studies: Optimization of Measurement Procedure
Given that the measurement models used for hand-scoring and AES were different, direct comparisons of the reliability coefficients necessitated conducting D studies.
Matching Measurement Models Using Eight Scores per Genre
The first D study estimated G and Phi coefficients for the PEG scoring method when increasing the number of prompts to eight per genre. We then compared the resulting G and Phi coefficients with those obtained from the original measurement design of the hand-scoring method that involved eight scores per genre (4 raters * 2 prompts per genre).
As shown in Table 9, increasing the number of prompts to eight per genre increased the composite reliability of the PEG rating method for all grades for non-struggling and struggling writers. Composite G and Phi coefficients for PEG were reliably equal to those of hand-scoring. Furthermore, when scoring the writing of non-struggling writers, G and Phi coefficients of PEG using eight prompts per genre in all cases exceeded the corresponding coefficients obtained using hand-scoring with four raters and two prompts per genre. Finally, increasing the number of prompts to eight per genre also improved the reliability of PEG for scoring the writing of struggling writers in individual genres, making PEG’s reliability coefficients for this subgroup generally equal with that of hand-scoring. However, even when the number of prompts per genre was increased to eight, PEG’s reliability coefficients did not attain minimum acceptability (i.e., 0.80) when scoring Grade 4 struggling writers in the informative genre (G and Phi = 0.76).
G and Phi Coefficients for Modified Measurement Design (D Study): Eight Scores per Genre
Matching Measurement Models Using Four Scores per Genre
A next set of D studies considered the comparative reliability of the scoring methods when using measurement models for four scores per genre. For hand-scoring, we examined reliability when using two prompts per genre scored by two raters. For AES, we examined reliability when using four prompts per genre.
As shown in Table 10, composite G and Phi coefficients exceeded the .90 criterion of reliability for making high-stakes decisions for both scoring methods. For non-struggling writers, the use of four prompts per genre with PEG resulted in higher reliability than hand-scoring with two raters and two prompts. For all grades and genres, PEG yielded higher G and Phi coefficients. However, for struggling writers, different scoring methods were superior for specific grade and genre combinations. For instance, (a) the scoring methods were equally reliable for scoring Grade 3 struggling writers in informative and narrative genres; (b) human raters were more reliable than PEG for scoring Grade 3 struggling writers in the persuasive genre, Grade 4 struggling writers in informative and persuasive genres, and Grade 5 struggling writers in narrative and persuasive genres; and (c) PEG was more reliable than human raters for scoring Grade 4 struggling writers in the narrative genre and Grade 5 struggling writers in the informative genre.
G and Phi Coefficients for Modified Measurement Design (D Study): Four Scores per Genre
Discussion
To help educators better understand the merits and ideal uses of common (i.e., hand-scoring) and novel (i.e., AES) methods for scoring writing performance assessments in a formative assessment context, we conducted the present study to examine, using G theory, the reliability and efficiency of hand-scoring and AES for assessing upper-elementary students of different levels of writing ability composing multiple genres of writing. Rather than pit one scoring method against the other, we aimed to surface ways that the different scoring methods might have complementary strengths that educators could leverage to address the challenges of assessing writing for instructional decision making.
How Similarly Did the Scoring Methods Rank-Order Students?
Generally, the scoring methods rank-ordered students quite similarly: correlations between scoring methods were ≥ .70 in 15 of 18 instances. Strong correlations might be expected given our rater-training methods that involved calibration with writing samples scored by PEG. However, in three instances, two of which pertained to struggling writers, correlations were only moderate (range = .55–.69). Thus, even with our calibration process, there do appear to be differences in how humans and PEG rank-order students.
How Much Did the Scoring Methods Distinguish Students?
When assessing non-struggling writers across grade levels and genres, student ability explained the largest percentage of variance in observed scores overall, with an average of 71.7% of the variance across both scoring methods. Student ability explained an average of 77.0% of the variance in PEG scores and an average of 66.4% of the variance in rater scores. However, when assessing struggling writers, student ability explained a lower percentage of the variance in observed scores overall, with an average of 66.3% (range = 28.7%–93.9%) across both scoring methods. The percentage of variance explained in the PEG models averaged only 58.4%. In human rater scores, however, student ability always explained the largest percentage of variance for struggling writers, with a mean of 74.2%. In fact, student ability explained a larger percentage of variance overall in human rater scores for struggling than non-struggling students. Thus, human raters seemed to better distinguish student ability in struggling writers.
What Are the Impacts of Prompts and Relationships Among Genres?
The students’ scores and raters’ judgments were not influenced by the writing prompts. Indeed, correlations of students’ scores between genres were large for human raters and PEG, but slightly more varied for PEG, especially for struggling writers. The range of correlations between genres across grades for hand-scoring non-struggling writers was .77–.92 and for hand-scoring struggling writers was .65–.92. The corresponding range of correlations for PEG scoring non-struggling writers was .61–.97 and PEG scoring struggling writers was .38–1.00. Thus, PEG may have been slightly more sensitive to differences between genres for struggling students, although moderate to large correlations still suggest that the students’ writing scores were largely unaffected by genre. This result was unexpected, as it conflicts with other research that resulted in correlations among students’ scores across genres ranging from .22–.60 for struggling writers (Graham et al., 2016), although it does align with results of a prior study using the same writing prompts (Wilson et al., 2019).
One potential explanation for large correlations between genres may be similarities in the administration procedures. Instructions signaled that the purpose of the writing assessment was to include specific genre features (e.g., reasons, evidence) rather than communicating broader and more authentic purposes for each genre (e.g., writing to entertain, writing to inform). Conveying a broader purpose may have led to more varied responses across genres. Also, we spent considerable attention to equating the prompts prior to the study, and we randomized and counterbalanced the prompts. Therefore, results are likely a product of this particular study that used highly controlled prompts and directions and should not be interpreted to mean that students and raters are never influenced by prompts (see Lee & Anderson, 2007).
How Reliable Are the Scores From Both Scoring Methods?
Human raters and PEG achieved high reliability coefficients overall. Human raters seemed to outperform PEG in the D studies based on the original measurement design, but differences in the measurement models likely explain those results: hand-scoring generated eight scores per genre, and PEG generated two. Thus, we conducted D studies to better match the two rating methods. In 8- and 4-score D studies, PEG’s reliability coefficients were more consistent with those of human raters. Both scoring methods achieved composite G and Phi coefficients above .90 for all grade levels for non-struggling and struggling writers. However, hand-scoring was generally more reliable than AES for assessing struggling writers, and AES did not meet a reliability criterion of .80 in some instances.
Given these results, we conclude that decisions about which scoring method to use and for which group of students should be made based on the purpose of the assessment and the benefits and challenges of each scoring method. Reliable hand-scoring may require fewer prompts per genre per student, but multiple raters (two or four) are required to spend a considerable amount of time hand-scoring. Conversely, using AES may require eliciting more writing samples from struggling writers, which reduces instructional time. Therefore, contextual strategies for using both scoring methods may be most appropriate. Thus, in the following section, we provide recommendations for leveraging each scoring method individually and in combination.
We present these recommendations with the caveat that they are based on the results of our study, which was an analysis of the psychometric properties of two assessment methods that affect one aspect of assessment efficiency: the number of writing samples and scores needed to obtain a reliable estimate of students’ writing ability. However, this is only one aspect that should be weighed in a more comprehensive decision-making process that includes a cost analysis that considers PEG license fees, time and wages for human raters, time and costs for training human raters, and the relative costs of accurate/inaccurate student placement into intervention or special education programs. Although we did not conduct such a cost analysis in the present study, we believe that the following recommendations have practical value and can serve as the foundation for future research that interrogates these recommendations in other ways, such as with a cost analysis or with qualitative research exploring teachers’ perceptions of the utility and validity of each scoring method.
Recommendations for Optimizing the Scoring of Elementary Students’ Writing
Recommendation 1: Base Decisions on Scores Averaged Across Multiple Writing Assessments
We strongly recommend that educators obtain multiple writing samples and/or use multiple raters when making decisions that have consequences for students, such as making instructional placements and disability identification. Within the context of multitiered systems of support, performance data are used to determine whether a student should receive supplemental intervention. Our findings suggest that teachers will not have sufficiently reliable data from a single writing performance assessment to make correct decisions, regardless of the scoring method. To ensure valid decision making, teachers should base such decisions off multiple writing scores.
Recommendation 2: Use PEG (or Another AES) for Periodic Class-Wide Formative Assessment
PEG provides immediate and reliable evaluation of student writing. Administering two prompts per genre was sufficient to make accurate genre-specific decisions for non-struggling writers, but educators must administer a greater number of prompts to make accurate genre-specific decisions for struggling writers: four prompts per genre were sufficient in all but three instances (Grade 3 informative, Grade 5 informative, and Grade 5 persuasive). Thus, if educators are administering periodic formative assessments to support class-wide instruction (i.e., Tier 1), PEG provides a viable option, especially for non-struggling writers. Caution should be taken, however, to ensure that students do not adopt strategies to “game” the AES system and artificially inflate their score (see Higgins & Heilman, 2014).
Recommendation 3: Blend AES With Hand-Scoring, Relying on Human Raters to Make the Most Consequential Decisions
PEG can save human resources, and it is highly reliable for scoring average or higher-performing writers (i.e., students at or above the 30th percentile). However, PEG is less reliable than hand-scoring for making decisions for vulnerable students. Thus, it seems that human raters may be better suited to reliably scoring struggling writers. It may be that the training data used to build PEG’s scoring algorithms lacked sufficient representation from the lower end of the scoring scale (see Raczynski & Cohen, 2018). Alternatively, it may be that human raters are more sensitive to seeing the potential in students’ writing even when it includes errors. Perhaps humans can infer students’ intentions despite minor errors in spelling, word choice, or organization, whereas AES does not. Regardless, it appears that human raters may be able to score the writing more reliably for struggling writers, which is the subgroup of writers for which the most consequential decisions are made.
Therefore, we recommend using both scoring methods in a complementary way. PEG can be used with a small number of writing assessments (e.g., two) for all learners. This approach will help educators distinguish non-struggling writers from struggling writers with a limited number of writing samples (see Wilson, 2018; Wilson & Rodrigues, 2020). Because the stakes are typically lower for decisions regarding non-struggling students, and PEG is very reliable for this group, no human scoring is needed. Educators can then use human raters to hand-score writing samples for the lower-performing writers; this step requires no additional assessment time for students and would provide the necessary scores for reliable decisions. Furthermore, through the scoring process, teachers may identify their students’ strengths and weaknesses, information that will assist them in planning instruction and intervention. For higher-stakes decisions, educators may also add another writing assessment or two to improve confidence in the hand-scoring. Of course, this is only one possible approach, and it still requires research support. Indeed, an area of future research would be to conduct generalizability and decision studies to examine the extent to which reliability coefficients improve for samples of struggling writers when adding successive numbers of human ratings to a baseline reliability coefficient achieved through AES.
Limitations and Directions for Future Research
First, due to resource limitations, students hand-wrote their responses, which were later transcribed into Word documents for scoring by PEG. In practice, to maximize the efficiency of AES, students must use word-processing software or AWE systems when writing.
Second, we examined the PEG Overall Score, which was formed as the sum of the six individual trait scores. We did this for conceptual and psychometric reasons. However, future research using other automated trait scoring systems should explore the generalizability of automated trait scores for supporting instructional decision making.
Third, the large correlations between genres were consistent with results of a prior generalizability study of PEG (Wilson et al., 2019) but inconsistent with findings from other studies exploring the effects of genre on composing. Natural prompts developed by teachers may not always be as tightly controlled and, therefore, may lead to different results. Future research should examine how much varying prompts might affect student responses and scores.
Fourth, PEG scores via grade bands (e.g., Grades 3–4 and 5–6), whereas humans scored based on the PEG range finders for each grade level individually. This approach may have contributed to the superior reliability of human scoring to PEG for scoring the writing of struggling writers in Grade 5, specifically. In PEG, Grade 5 is the beginning of a new scoring band that spans Grades 5–6, and Grade 5 students generally score on the lower range of that band (see Table 10). Range restriction decreases reliability, which may have doubly affected Grade 5 struggling writers when scored by PEG.
Similarly, differences in reliability of the rating methods for struggling and non-struggling writers may have been due, in part, to range restriction in the struggling writer sample. Nevertheless, the population of struggling writers is, by definition, one with a restricted range of performance, so the issue is not specific to this study; it is ubiquitous. However, the degree to which a cutpoint, such as the 25th percentile, is lowered (10th percentile) or raised (35th percentile) may exacerbate or ameliorate the impact of range restriction on score reliability, particularly for PEG scoring. Indeed, even though the non-struggling writer sample did not score above 21, far lower than ceiling of 30 in PEG, the degree of range restriction was not so severe as to negatively affect reliability for that subgroup of students.
Finally, the present study examined whether human and automated scoring were differentially reliable for struggling writers, a vulnerable subgroup. However, we did not explore the potential for differential reliability, or bias, in scoring for other subgroups, such as students from racial, ethnic, and language minorities. It is important that future research extend the model of investigation we applied to ensure that decisions made from writing assessment data about students from these groups are equitable.
Supplemental Material
sj-pdf-1-aer-10.3102_00028312221106773 – Supplemental material for Examining Human and Automated Ratings of Elementary Students’ Writing Quality: A Multivariate Generalizability Theory Application
Supplemental material, sj-pdf-1-aer-10.3102_00028312221106773 for Examining Human and Automated Ratings of Elementary Students’ Writing Quality: A Multivariate Generalizability Theory Application by Dandan Chen, Michael Hebert and Joshua Wilson in American Educational Research Journal
Footnotes
Notes
D
M
J
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
