Abstract
Previous research has shown the effectiveness of peer reviewing on the improvement of writing quality. However, the fact that students themselves, arguably novices, judged the improvement leads to concerns about the validity of peer reviewing. We measured writing quality before and after peer reviewing using Coh-Metrix, which computationally evaluates the linguistic properties of writings. Participants also evaluated the quality of their peer writings using the SWoRD system. Both measurements, particularly the computational measurement, confirmed the effectiveness of peer reviewing. In addition, the computational measures found that awareness of cohesion, including the clarity, explicitness, and concreteness of writing, improved over the course of peer reviewing. The results are discussed, along with their possible implications for the complementary roles of peer reviewing and computational writing measurements.
Introduction
Despite the importance of writing ability with respect to students’ academic success and their participation in the nation’s economic success, there has not been significant progress in writing education (The National Commission on Writing, 2003). This is partly due to the heavy workload of instructors: Writing assessment is intensively time and effort consuming. Although research has established the value of learning writing by doing writing (Bangert-Drowns, Hurley, & Wilkinson, 2004; Couzijn, 1999; DiPardo & Freedman, 1988; Haswell, 2005, 2008), instructors are generally reluctant to give writing assignments due to the workload involved in commenting on and grading them (The National Commission on Writing, 2003). Consequently, students might not receive feedback even if they produce a considerable number of writings. Given this situation, there is an educational need for a system by which students can obtain feedback to improve their writing abilities.
Peer review has been regarded as an important alternative to the learning writing by doing writing approach. Building on the learning writing by reviewing approach (Cho & MacArthur, 2011), students learn writing by identifying both problems and solutions in the course of reviewing and commenting on their peers’ writings (Cho & MacArthur, 2010). In a typical peer review interaction, student writers receive feedback from peer reviewers and in turn serve as reviewers to give feedback to peer writers (Y. Cho & K. Cho, 2011). A substantial body of research has demonstrated the effectiveness of bilateral feedback on the learning and performance of writers (Althauser & Darnall, 2001; Beach & Friedrich, 2006; Cho & MacArthur, 2010, 2011; Cho, Schunn, & Charney, 2006; Couzijn, 1999). In addition, peer reviewing is practically valuable in that it allows students to receive feedback from multiple peers instead of from just one instructor, as well as dramatically reduces the instructor’s workload.
In spite of the substantial benefits of the peer review system, the reliability and validity of feedback generated by students have been questioned due to potential concerns that students might be novice and bias-prone as reviewers (Y. Cho & K. Cho, 2011; Hamer, Purchase, Luxton-Reilly, & Denny, 2015; Lu & Law, 2012) and might lack motivation to review peer students’ work “faithfully and fairly” (Wang, Ai, Liang, & Liu, 2015, p. 180). In addition to theoretical debates about the reliability and validity of feedback (see Cho, Schunn, & Wilson, 2006, for review), Rushton, Ramsey, and Rada (1993) have pointed out that students themselves tend to be suspicious of peer assessments, although there is little qualitative difference between the comments made by students and those made by instructors.
In the same vein, the reliability and validity of grading writing have been questioned (Cho et al., 2006; Evans, 2013). In the previous literature on peer reviewing, student reviewers graded the quality of peer writings. Thus, the improvement of writing as judged by those novices could not be guaranteed. Because of the subjective nature of feedback and grading given by novices, the practice of peer reviewing is limited or often abandoned until its effectiveness becomes measurable (Rushton et al., 1993). To ensure the value of peer reviewing, it is necessary to investigate its effect using objective measurements.
In this research, therefore, we examined the effect of the learning writing by reviewing approach using unbiased computational measurements. To achieve this goal, two computational systems were selected: (a) SWoRD (Cho & Schunn, 2007), an automatized and computerized writing support system that allows students to review their peers’ writings; and (b) Coh-Metrix (Graesser, McNamara, Louwerse, & Cai, 2004; http://cohmetrix.com), a computational linguistic tool to measure the quality of writings based on a number of linguistic properties. To be specific, we investigated the correlations between grades given by peer reviewers in SWoRD and scores of various indices in Coh-Metrix before and after the peer reviewing. We expect that the improvement of writing by peer reviewing, if any, will be captured by the pattern of correlations with the linguistic indices that Coh-Metrix computes.
The results of our study contribute to the following issues. First, the computational measurement will improve the validity and reliability of comments and grades generated by peer reviewers. Second, this computational approach will improve understanding of what specific aspects students learn as a result of peer reviewing. For example, linguistic properties (e.g., the degree of coherence and cohesiveness of writing) reflect writers’ strategies regarding how to express their thoughts in the most efficient way without ambiguity. Exploring what and how linguistic properties reflected in the indices change over the course of the peer review helps to appreciate changes in the mental representation of good writing.
As such, we also conducted a regression analysis to identify the relevant linguistic features responsible for the improvement of writing quality. In addition to the theoretical contribution, our work has an important practical implication. We will ascertain whether a more fruitful approach is to view these two approaches—the peer reviewing and computation-based approaches—as complementary. Integration of the two approaches will shed light on the development of a useful computational tool along with peer reviewing in writing education.
Peer Review of Writing
Peer review allows students to receive feedback from multiple peers and thus various points of view (Brindley & Scoffield, 1998; Y. Cho & K. Cho, 2011; Cho & MacArthur, 2010, 2011; Cho & Schunn, 2007; Ertmer et al., 2010; Holliway & McCutchen, 2004; Lundstrom & Baker, 2009; Patchan, Hawk, Stevens, & Schunn, 2013; Topping, 1998; Warren & Cheng, 1997). Cho and Schunn (2007) demonstrated that feedback from multiple peers led to more effective learning than feedback from a single expert reviewer (see also Cho & MacArthur, 2010). Multiple reviewers are likely to diagnose more problems. When reviewers are not in agreement on some specific problems, students are exposed to rich and potential perspectives (Miyake, 1986). When this occurs, students tend to view the comments as more reliable, and hence, take them more seriously.
It is generally accepted that feedback from instructors improves the quality of writing of feedback recipients (Beach & Friedrich, 2006; Berg, 1999; Nelson & Schunn, 2009; Paulus, 1999; Topping, 1998; see Kluger & DeNisi, 1996, for general review). Contrary to our common belief that feedback from novice students is not helpful, previous work has suggested that peer feedback could be as effective as instructors’ feedback. Students share knowledge, problems, and language. These commonalities lead them to be more effective in diagnosing and solving writing problems from their common perspective. It also helps them to convey their ideas using understandable comments. Cho et al. (2006) demonstrated that the effectiveness of feedback given by student peers with guidance on peer assessment and carefully constructed rubrics is comparable to that given by experts.
In addition, peer reviewing enhances reviewers’ writing abilities by providing feedback. First, to be a good writer, it is critical to understand the reader’s perspective. Reviewing drafts allows reviewers the opportunity to take an audience or reader’s perspective (Butterfield, Hacker, & Albertson, 1996; Holliway, 2004; Traxler & Gernsbacher, 1992, 1993; Wyngaard & Gehrke, 1996). Moreover, to provide comments on peers’ writings, student reviewers must exert cognitive effort to evaluate the strengths and weaknesses of the writing that they are reviewing and to advise solutions (Bereiter & Scardamalia, 1987; Fitzgerald, 1987; Flower, Hayes, Carey, Schriver, & Stratman, 1986). Consequently, student reviewers are likely to learn implicitly about what qualities good writing should and should not have. Finally, by reviewing peers’ writings, student reviewers are required to offer reasonable explanations to justify their evaluations and to suggest solutions to their peer writers. In other words, the review process involves developing problem-solving skills on their own (Liu & Carless, 2006), enhancing metacognitive awareness of how good or bad their writing is (e.g., Cho & MacArthur, 2011; Cho & Schunn, 2007; Liu & Carless, 2006; Topping, 1998), and improving students’ motivation and self-confidence (Brindley & Scoffield, 1998; Dochy, Segers, & Buehl, 1999; Warren & Cheng, 1997).
The active development of reviewers’ metacognition significantly emerges through the activity of explicitly writing comments. Cho and MacArthur (2011) asked participants in one condition to provide explicit comments on their peers’ writing samples, whereas participants in the other condition were asked to simply review the writing samples without providing comments. The researchers observed significant writing improvement only for those who explicitly produced comments. Furthermore, Y. Cho and K. Cho (2011) demonstrated that providing comments does not always elicit an equivalent amount of improvement in writing skills. They found that students’ writing performance improved most when students provided comments that criticized the weakness of texts at a micro level and that expressed the strength of texts at a macro level. In short, previous studies suggest that peer reviewing develops metacognition of good writing from the reader’s perspective by asking reviewers to generate explicit comments.
Although the research on peer review is generally positive, in that the give-and-take of comments leads students to develop evaluation ability, expand responsibility for their own learning, and learn how to be better writers, there remain doubts about the efficacy of peer review-based writing education (Boud, 1989; Cho & Schunn, 2007; Lynch & Golen, 1992; Magin, 2001; Rushton et al., 1993; Stefani, 1994; Swanson, Case, & van der Vlueten, 1991). First, one of the main concerns is the possibility of low validity and reliability of the comments offered by peer reviewers. Since students are novices with respect to assessing the quality of writing, they are vulnerable to biases caused by race, friendship, or temptation to inflate grades (W. T. Dancer & J. Dancer, 1992; Mathews, 1994). However, note that Cho and MacArthur (2010) demonstrated that with a little guidance on how to assess the writing, students were able to write comments with sufficient reliability and validity comparable to experts’ assessments. Nonetheless, the researchers also pointed out that the validity and reliability of students’ comments and grading of others’ drafts remain open to question (Cho & MacArthur, 2011). Therefore, it is necessary to validate the benefits of peer reviewing using an objective and quantitative method to support its use.
In addition, little is known regarding what aspects of the mental representation of good writing student reviewers have in mind change as a result of reviewing their peers’ writings. This is somewhat due to the fact that previous research on peer reviews has tended to adopt a pragmatic approach. The research has tried to establish the validity of peer review education by showing that the writing quality increases by virtue of peer reviewing. The majority of studies explained the effectiveness of peer review at a metacognitive-level (Cho & MacArthur, 2010, 2011; Cho & Schunn, 2007; Holliway & McCutchen, 2004; Liu & Carless, 2006; Topping, 1998). Although metacognitive-level analysis provides useful insight into why the peer review experience makes students better writers, it is also necessary to elaborate the locus of why and how the use of peer review leads to the improvement of the reviewer’s own writing.
In summary, previous work has indicated that the reciprocal nature of peer review helps student writers to improve writing quality. However, open questions remain regarding the validity of comments generated by student reviewers and what linguistic aspects of writing develop during peer review. To address the issues just raised, a number of computational linguistic methods have been suggested, which is the topic of the next section.
Computational Approach to Writing Evaluation
There have been objective and computational attempts to evaluate the quality of writings by measuring the linguistic properties of writings, as found in tools such as Coh-Metrix (Graesser et al., 2004), RANGE (Nation & Heatley, 2002), Latent Semantic Analysis (LSA; Foltz, Kintsch, & Landauer, 1998), and e-rater (Attali & Burstein, 2004; Chodorow & Burstein, 2004). Linguistic properties reflect a writer’s mental representation. Generally, in order to write a good text, writers should know “what” to write. At the same time, they also need to know “how” to write. In other words, writers should be strategic about how to express their thoughts in the most efficient way without ambiguity. Good writing should be outspoken, clear, coherent, and natural in flow. Computational methods use linguistic properties including coreferential cohesion and lexical diversity to capture such mental representations. Hence, it is not surprising that the computational tools use linguistic properties as an objective measurement to assess the quality of writing. For example, early attempts calculated indices of syntactic or grammatical complexity (Flahive & Snow, 1980), text length (Homburg, 1984), and indices of lexical complexity (Engber, 1995), and then combined them to estimate the writing quality. LSA computes the similarity among words, sentences, and texts, using cosine measures in multidimensional spaces (Landauer, Foltz, & Laham, 1998). Studies using LSA similarity measures have successfully predicted the qualities of essays that were produced on several occasions (Foltz et al., 1998; León, Olmos, Escudero, Cañas, & Salmerón, 2006).
In this study, we estimated writing quality using Coh-Metrix (Graesser et al., 2004), which provides a large number of linguistic measures at the word, sentence, and discourse levels. In addition, our motivation was to investigate changes in mental representation about good writing by measuring the linguistic properties of writing. To be specific, we are interested in the degree of cohesion of a text in which the reader mentally connects and follows ideas in the writing (Graesser, McNamara, & Louwerse, 2003). In addition to measurements associated with cohesion, we were also interested in other measurements showing lexical characteristics (e.g., word frequency and word concreteness) and the readability of texts.
More importantly, we noticed that Coh-Metrix could be used to test the effectiveness of peer reviewing objectively as the validity of Coh-Metrix has been demonstrated in various studies (Crossley, Louwerse, McCarthy, & McNamara, 2007; Crossley & McNamara, 2008; Crossley, McNamara, Weston, & McLain Sullivan, 2011). Crossley et al. (2011) used Coh-Metrix measures in discriminating among the differences in the linguistic features of texts written by 9th graders, 11th graders, and college students. They found that high school students produced texts with less complex and more cohesive sentences and used fewer sophisticated words than college students. Coh-Metrix measures were also used in the evaluation of writing quality in English as Second Language (ESL) texts (Crossley et al., 2007) and in writing produced by ESL students (Crossley & McNamara, 2008).
Of interest to us, Coh-Metrix measurements could be used to objectively understand how the linguistic qualities of writing change over the course of writing and rewriting. In this experiment, we would like to examine how student reviewers’ evaluation of writing quality compares to the linguistic indices measured by Coh-Metrix, and how students’ evaluation based on linguistic qualities can develop as they increasingly review their peers’ writing. We expect that the improvement of writing quality after peer review will be reflected in Coh-Metrix’s measurements.
Overview of the Study
We collected writing samples using SWoRD (Cho & Schunn, 2007), a computerized writing support system for students to review their peers’ writings easily. Participants first wrote and logged onto the system to upload their own texts (i.e., first drafts). They then received and downloaded the texts of peers and the fixed rubric to instruct evaluation. Their job was to review those texts by providing explicit comments and scoring their quality. After finishing the first round of peer review, participants logged onto the system again and submitted their comments and scores. After that, participants received their own draft back with peer comments and scores. Participants were asked to revise and upload their revised texts (i.e., second drafts) based on the comments that peers had produced. Finally, participants again received the revised drafts of peers and rated the quality of those drafts as in the first round of peer review.
Meanwhile, we computed the linguistic properties of the first and second drafts using Coh-Metrix. The relationship between objective linguistic properties measured by Coh-Metrix and the subjective quality of writing evaluated by peer reviewers was tested using a series of correlation and regression tests for the first and second drafts, respectively. A brief layout of the procedure is depicted in Figure 1.
The layout of the procedure.
We hypothesize that student reviewers’ linguistic sensitivity while evaluating the quality of writing will be enhanced as they increase their experience in reviewing peer texts. If our hypothesis is correct, we expect to observe evidence of improved linguistic sensitivity about what constitutes good or bad writing. Improved linguistic sensitivity, according to our meaning, refers to satisfying the following hypotheses: Hypothesis 1: Significant correlations between the linguistic properties of texts and writing quality that peer reviewers evaluate will occur more extensively for the revised drafts than for the first drafts. Hypothesis 2: The regression power of how well a set of linguistic properties predicts writing quality will be stronger for the revised drafts than for the first drafts.
Method
Participants
Forty-one undergraduates in an introductory physics course at a Tier 1 research university in the United States participated in this study as part of their course activities. They were all native English writers. Most students were in their first or second year at the university.
Procedure
Peer reviews were collected with the SWoRD system by following all SWoRD procedures (Cho & Schunn, 2007). Participants were asked to write a laboratory report about the velocity of sound. They were also told that they would receive four of their peers’ drafts, which they were required to review. In the laboratory reports, participants were asked to account for their scientific findings using sound theories. In writing, they were asked to follow the general format of laboratory reports; that is, the reports consisted of an abstract, an introduction and theory, an experimental setup, data analyses and results, and a conclusion. After submitting the draft, participants went through reciprocal peer reviews for 4 weeks. They were assigned to review and evaluate four first drafts written by different peers. In particular, they were asked to explicitly comment on the strengths and weaknesses of their peers’ drafts. The scoring criteria and rating ranges were available in the SWoRD system. The scoring ratings ranged from 1 (Disastrous) to 7 (Excellent) in three domains: logic, flow, and insight. Flow concerns how easily the main points of the argument could be followed (Flower et al., 1986). Logic addresses the degree to which the writing is logically coherent. Insight involved the extent to which the writing shows information inferences contributed beyond the content of assigned course activities. All three dimensions are important not only for a scientific report but also for general writing. After rating the quality of writing, reviewers also submitted written comments.
Participants then got their first draft back, which had been reviewed by four peers. They were asked to revise their first drafts with respect to the comments they had received. After submitting their second drafts, they evaluated and commented on four revised drafts; student reviewers evaluated first and revised drafts from the same student writers. Finally, participants rated the helpfulness of the review comments that they had received on a 7-point scale (i.e., from 1 = least helpful to 7 = most helpful ). All steps for this procedure are displayed in Figure 2. Individual students used the evaluation rubrics at least 18 times for writing, reviewing, and rating helpfulness.
Steps for writing and peer reviewing conducted in the SWoRD system.
All writers and reviewers were blind to each other to prevent them from being unnecessarily critical in evaluating and commenting on peer drafts when they knew who the authors or reviewers would be (Crampton, 2001). To ensure the anonymity of writers and reviewers, participants used pseudonyms when they wrote their drafts, and reviewers were identified by numbers that an experimenter had randomly assigned to them. After finishing the SWoRD procedure, Coh-Metrix was used to analyze the abstract, introduction, and conclusion sections for each draft.
Measures
We computed two types of measures for each draft. One was the subjective quality scores assigned by peers. These scores were automatically computed in the SWoRD system. The score was computed as the mean of the writing scores given by four peer reviewers. The writing quality ranged from 3.58 to 6.67 in the first drafts (M = 5.39, SE = 0.11) and from 4.67 to 6.92 in the second drafts (M = 6.09, SE = 0.09). The other measure was the set of linguistic indices, measured by the Coh-Metrix tool. We computed 54 linguistic measures across seven dimensions that were available in the public version of Coh-Metrix 2.0. Because all drafts were scientific reports, the main body of the drafts mostly consisted of graphs, math functions, and mathematical algorithms with little text. We decided that it would not be meaningful to assess the linguistic properties of the main body. As a result, the indices were calculated based on the abstract, introduction, and conclusion of each draft.
Statistical Analyses
With the measurements of the subjective quality score and Coh-Metrix, we conducted a series of statistical analyses in the following order:
Testing whether the two measurements were significantly correlated in the first and second drafts. Testing whether Coh-Metrix linguistic indices contributed to predicting the subjective quality score in the first and second drafts.
Results
We first checked whether peer reviewers judged that writing quality improved over the course of peer review. The subjective quality scores of the second drafts were significantly higher than those of the first drafts, t(20) = 7.86, p < .001. Our results are consistent with the previous findings in that writing quality improved after peer review (Cho & MacArthur, 2010, 2011; Cho & Schunn, 2007). Because our goals were to investigate how computational linguistic indices were associated with the degree of writing quality, we conducted two sets of correlation analyses. We first mapped the correlations between the subjective writing scores and Coh-Metrix indices using the first drafts, and then conducted the same procedure using the second drafts.
Correlations Between the Coh-Metrix Indices and the Ratings of First Drafts
Results From Correlation Analyses Between Rating of Writing Quality in the SWoRD System and Linguistic Indexes Obtained From Coh-Metrix.
Dimension is determined following the general guideline of Coh-Metrix indices at http://cohmetrix.memphis.edu.
p < .05; **p < .01.
Our results indicate that the more cohesive texts are not evaluated as being as good. Our result seems contradictory to some previous studies that suggest lexical diversity to be a property that good writing should have (McNamara, Crossley, & McCarthy, 2009). However, we noted that for our study, students wrote science reports in which we think the clarity of writing might be more important than the diversity. Alternatively, lexical diversity might be a characteristic of texts written by advanced writers or experts rather than beginners (see also Crossley et al., 2011). For the first drafts, the other indices were not significantly correlated to the subjective writing scores.
Correlations Between the Coh-Metrix Indices and the Ratings of Second Drafts
The number of significant correlations between the quality of the second drafts and the Coh-Metrix indices greatly increased. Nine Coh-Metrix indices that were significantly related to the subjective writing scores are presented in Table 1. In contrast to the correlation pattern observed in the first drafts, the extended patterns of correlations in the second drafts are exactly as we hypothesized (Hypothesis 1). These correlation results supported our hypothesis.
The details of the nine significant correlations are as follows. First, the drafts with a higher number of sentences were determined to be better written than those with a lower number of sentences. The subjective writing scores were positively correlated with the number of sentences, corroborating the previous finding that text length is often used as a determiner as to whether texts are good or bad (McNamara et al., 2009).
The second finding is that drafts with concrete words were considered lower quality. There was a negative relation between the scores and number of hypernyms. Given the nature of the analyzed sections such as the abstract, introduction, and conclusion per draft, in which writers usually tend to generalize what they have found, it makes sense that drafts using more concrete or specific words were judged to be less good. In fact, we found some comments in which student reviewers recommended their peers use general rather than specific words with respect to the nature of an abstract.
Third, drafts with more pronouns were evaluated to be less good than those in which pronouns were used less. The number of personal pronouns (PPs) and the pronoun ratio (PR) were negatively correlated with the subjective writing scores. This negative pattern made sense with regard to ambiguity. That is, in reading texts, if readers do not know which pronoun refers to which person or which antecedent, it is difficult to judge the texts to be of good quality. Namely, a high density of pronouns in comparison to noun phrases often yields referential cohesion problems, resulting in unclear texts that might contain significant ambiguity.
Fourth, referential indices such as adjacent anaphoric reference (AAR) and anaphoric reference (AR) were also negatively related to writing quality. These correlations indicate that the ratings of writing quality decreased as the proportion of ARs and AARs referring back to a constituent up to five sentences previously increased. Observing these results was not surprising given that the other indices (i.e., PP and PR) associated with coreferential cohesion have already shown significant negative correlations with writing quality. Altogether, more frequent occurrences of anaphoric usage in texts tend to elicit higher ambiguity and thus decrease the clarity of the texts.
Finally, the degree of writing quality increased as the use of expressions representing intentional cohesion (e.g., in order to, so that, or for the purpose of) increased. The subjective writing scores had a negative relation with the occurrence of intentional content (e.g., the use of intentional verbs) and had positive relationships with the ratio of intentional cohesion. Given that the drafts were scientific reports, it seems sensible that drafts that clearly expressed the writers’ purposes (i.e., intentional cohesion) were judged to be good, whereas those that injected writers’ subjective voices (i.e., intentional content) were judged to be less good.
Sign of Self-Learning Based on Comparisons Between First and Second Drafts
To explore what linguistic properties changed over the course of the peer review process, we compared the correlation coefficients between the first and second drafts among the 54 indices. Figure 3 presented correlation coefficients that were significant in the second drafts.
The comparison of correlations in the first and second drafts. NH = noun hypernym; NS = number of sentences; PP = personal pronouns; PR = pronoun ratio; TTR = type-token ratio; AAR = adjacent anaphoric reference; AR = anaphoric reference; Ico = intentional content; Ich = intentional cohesion.
Overall, the changes in correlation patterns indicate that students’ awareness of good writing improves, suggesting that student self-learning emerges. Note that only the TTR measure was significant in both drafts. For the first drafts, indices for coreferential cohesion such as PP, PR, AAR, and AR were weakly associated with the ratings of writing quality, but the patterns of these negative correlations were the same as those found in the second drafts. The same patterns of correlation found in the first and second drafts means that students did have some understanding of what constitutes good writing when they evaluated the quality of the first drafts. The enormous improvement in significant correlations in the second drafts suggests that student reviewers’ understanding of what constitutes good or bad writing improved while they were reviewing the second drafts.
The significant correlation patterns observed in the second drafts revealed that student reviewers seemed to consider clarity and cohesiveness as the most important criteria to judge the degree to which writing is good. Our speculation is based on the fact that linguistic indices referring to the clarity of writing were significantly correlated with writing quality in the revised draft. We reason that student reviewers became aware of the role of clarity and cohesiveness in writing, and they enhanced their understanding of what constituted clear and cohesive writing. As we hypothesized earlier, students’ linguistic sensitivity to clarity and cohesiveness improved as they gained more experience reviewing their peers’ writing. If linguistic sensitivity, in this context, captures what student reviewers learn during reviewing, the explanatory power of linguistic properties with regard to the variance in writing quality would increase as students’ linguistic sensitivity increases. For this further examination, we conducted two sets of regressions. The results of these regressions are reported later.
Regression Analysis
On the basis of the results from correlation analyses associated with the first and second drafts, we further aimed to identify which linguistic indices would significantly contribute to explaining the degree of writing quality. In the same vein, we explored whether the performance of regression models would be improved when the model ran for the second drafts in comparison to when the model ran for the first drafts.
Regression of Writing Quality Explained by Coh-Metrix Indices
For the first drafts, only TTR was significantly correlated with writing quality. Thus, we only submitted the TTR measure to a regression analysis as shown in Figure 4. The model’s adjusted R2 is .12, meaning that the model can explain 12% of the data. Although the power of this model is not strong, its performance significantly accounted for the variance of writing quality (F(1,39) = 6.28, p < .05).
The regression model using TTR and the rating of writing quality for the first drafts.
For the second drafts, we submitted a set of factors that were significantly correlated with the ratings of the drafts—that is, number of sentences, noun hypernym, TTR, PP, PR, AAR, AR, intentional content, and intentional cohesion. Because we had no prior belief about which factors should be entered into a model, we employed a backward stepwise regression in which the section of significant predictors was determined by sequentially removing variables from a full regression model. The results of the final model are displayed in Figure 5. Overall, the model’s adjusted R2 is .35, meaning the model can explain 35% of the data with the given set of significant predictors (i.e., TTR, PR, and noun hypernym; F(3,37) = 8.01, p < .001).
The regression model using pronoun ratio, noun hypernym, and TTR, and the rating of writing quality for the second drafts.
The results of the regression analysis suggest that clarity (TTR), explicitness (PR), and concreteness (noun hypernym) are the relevant linguistic properties to measure and predict the quality of scientific writings. As hypothesized (Hypothesis 2), the model’s performance using the second drafts dramatically improved in comparison to the model’s performance using the first drafts. The bigger effect size of the second model than the first model supports our hypothesis by showing that reviewers’ linguistic sensitivity in judging peers’ writing changed through the review process.
Discussion
The current study aimed to enhance the learning writing by reviewing approach with an objective computational linguistic measurement. To achieve these goals, we asked participants to write a laboratory report and to review and evaluate the drafts of four peers. Hence, a draft was reviewed and commented on by four different peers. Participants had to revise their initial draft based on comments offered by these four peers. Finally, they reevaluated the peers’ drafts. Critically, the quality of writing was measured with respect to the first and second drafts using Coh-Metrix (Graesser et al., 2004), which measures the linguistic properties of writing.
Our results indicate that the second draft significantly improved in comparison to the first draft, thus demonstrating the effectiveness of peer reviewing. First, the subjective ratings given by participants before and after the peer review were significantly different. Second, the correlation between the subjective ratings and scores from each of the Coh-Metrix indices corroborates the subjective ratings. While only one index revealed a significant correlation with subjective ratings in the first draft, nine indices did so in the revised draft, reflecting what students learned from peer reviewing.
Our results also indicate that awareness of clarity and cohesiveness is attributed to this improvement. The newly emerged indices with respect to the second draft converge with the clarity and cohesiveness of the writing. In addition, the regression analysis confirms the importance of clarity and cohesiveness of writing. The three indices representing clarity, explicitness, and concreteness are the only relevant predictors of the subjective quality of writing, explaining 35% of the data.
Our study might help establish the validity of peer reviewing built on the learning writing by reviewing hypothesis. Although prior empirical findings have demonstrated that peer reviewing improves writing (Cho & MacArthur, 2011), the effectiveness of peer reviewing has been challenged from the theoretical and practical perspectives. The main concern is the fact that novice students generate the comments and grading in peer reviewing. Unlike experts such as instructors, it is believed that students lack subject matter knowledge and writing skills and are also vulnerable to cognitive or social bias. Consequently, the quality of their comments might not be reliable. Contrary to this belief, our results suggest that the comments and grades reliably lead to the improvement of writing.
Previous work has suggested that peers can effectively generate beneficial comments. Peers share problems and languages. They often experience and observe joint difficulties. Thus, they may be more effective at pointing out and diagnosing problems according to their issues and offering solutions using the same level of language. Y. Cho and K. Cho (2011) demonstrated that students learn writing from reviewing peer writing.
At the same time, it should be noted that our participants were not trained to evaluate peers’ drafts. They only received the rubric and evaluated drafts accordingly. In contrast to previous findings that training how to evaluate is essential to ensure the validity of comments and grades (Berg, 1999; MacArthur, Graham, & Schwartz, 1991; Van Steendam, Rijlaarsdam, Sercu, & Van den Bergh, 2010), our study demonstrates that students can generate beneficial comments and give reliable grades.
The other benefit of peer reviewing is that students receive feedback from multiple peers. First, student writers can encounter multiple perspectives with respect to the same problem, or multiple reviewers can agree on the same problems. Agreement or disagreement on specific issues might be beneficial or salient for the writers to revise their draft. Second, multiple peers are more likely to cover all of the content than a single evaluator (Miyake, 1986). As a result, comments or grades given by multiple peer reviewers could be reliable and valid by reducing blind spots, providing various perspectives, and minimizing the negative impact of negative feedback (Cho & MacArthur, 2010).
Our study takes a further step to investigate what aspects of writing change as a result of peer reviewing. Previous studies have tended to focus on the effectiveness of activities within peer reviewing that are responsible for writing improvement. These activities include adopting a reader’s perspective, exerting cognitive effort to detect problems and suggest solutions, and enhancing metalevel knowledge. Specifically, previous studies were interested in the effect of the amount of feedback, the difference between comments written by students and those written by experts, the difference between reading a draft and commenting on a draft, and the type of comment, that is, whether it concerns micro or macro level structure. In these studies, an expert or novice student measured the overall quality of writing. Hence, little is known about what specific property of writing is developed as a result of learning. The current study focuses on the linguistic properties of writing and reveals that students learned the importance of clarity and cohesiveness. Furthermore, we demonstrate that clarity and cohesiveness can be used to predict the quality of writing.
In addition, our study sheds light on the possible integration of two different approaches, the peer reviewing approach and the computational approach. Although the two approaches come from very different theoretical backgrounds, they share an ultimate practical implication: improvement of writing ability. In this study, we demonstrate that the measurements based on the two approaches are significantly related. That being said, they could be used complementarily. The computational approach gives the peer review approach additional validity and reliability, whereas the peer reviewing approach offers evaluation that is about more than just linguistic properties. In practice, the combination of the two approaches could improve writing because they can play complementary roles. Students can benefit from peer reviewing as well as Coh-Metrix analysis.
Limitation and Future Directions
The current findings should be interpreted cautiously due to various limitations. First, it should be noted that our study did not consider differences among students, which could have an influence in generating and responding to feedback. Rucker and Thompson (2003) reported a gender difference in accepting feedback from an instructor. The level of knowledge or skills is also a factor that selectively influences the peer reviewing process. McCutchen, Hull, and Smith (1987) indicated there is a difference in revision strategies, which leads to varying revision quality between high-knowledge students and low-knowledge students. Future studies should incorporate analysis of such individual differences. The results hardly generalize to other age groups, particularly elementary and secondary students. It is reported that these groups produce less reliable and valid comments and ratings and are vulnerable to the length of writings (Cho et al., 2006).
Second, in the current study the main body of the student writings was excluded from Coh-Metrix measurements because it contained materials that Coh-Metrix cannot analyze. Since most of the main body of the drafts is nonverbal, such as graphs, mathematical functions, and figures, Coh-Metrix cannot determine the degree of linguistic properties. We asked participants to follow the rubrics that emphasize general consistency and insight rather than the content itself (see Cho et al., 2006, for details). We assume that the quality of writing measured by peers matches that generated by Coh-Metrix. However, this assumption should be validated. For example, different topics that do not require the use of nonverbal features should be considered for future research.
Lastly, our study is correlational rather than experimental. Because there are few studies in this field, we conducted an exploratory study using two different reliable measurements. While the procedure and results are consistent with previous research, our findings nonetheless seem tentative. Future studies using experimental methods are required to validate our results further. Whether using evaluation rubrics or not, peer reviewing can change its nature. Therefore, it would be interesting to examine how computational approaches can play a role in those contexts. In addition, computational approaches can be used to analyze good reviews and poor reviews from peer students. These studies could be strengthened by performing qualitative investigation on the associations between peer reviewing and computational analysis of writing.
Footnotes
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by the National Research Foundation of Korea (NRF-2010-32A-H00011).
