Abstract
Sequential evaluation is the hallmark of fair review: The same raters assess the merits of applicants, athletes, art, and more using standard criteria. We investigated one important potential contaminant in such ubiquitous decisions: Evaluations become more positive when conducted later in a sequence. In four studies, (a) judges’ ratings of professional dance competitors rose across 20 seasons of a popular television series, (b) university professors gave higher grades when the same course was offered multiple times, and (c) in an experimental test of our hypotheses, evaluations of randomly ordered short stories became more positive over a 2-week sequence. As judges completed repeated evaluations, they experienced more fluent decision making, producing more positive judgments (Study 4 mediation). This seemingly simple bias has widespread and impactful consequences for evaluations of all kinds. We also report four supplementary studies to bolster our findings and address alternative explanations.
Keywords
We rely on experienced human decision makers to accurately determine which students enter top universities and medical schools, which athletes are the world’s best, and even which scientific papers reach publication in peer-reviewed journals, such as this one. To make such determinations, evaluators judge targets in sequential evaluation by applying standard criteria. Yet this ubiquitous procedure, by virtue of its very process, allows for possible contamination in even experienced judges’ decisions. We propose that the process of repeated evaluations leads to increased experiences of processing fluency and, ultimately, an inflation of evaluations that come later in a sequence.
For decades, researchers have identified systematic biases among novices and experienced raters alike (Kahneman & Klein, 2009; Kahneman & Tversky, 1974; Tversky & Kahneman, 1981). While experience may lead to better decisions through accumulated knowledge (Dreyfus & Dreyfus, 1986), some researchers argue that experience is not necessarily associated with better decisions (e.g., physicians; Choudhry, Fletcher, & Soumerai, 2005; Ericsson, 2007).
Repeated evaluations, by definition, increase process-specific experience and knowledge. Yet even when stakes are high, judges resort to intuitive judgments, becoming cognitive misers (Fiske & Taylor, 2013) who use heuristic processing in lieu of more effortful deliberation (Harteis & Billett, 2013; Jiang & Hong, 2014; Kahneman & Frederick, 2005).
We posit that sequential evaluation causes a repeated decision process to feel more fluent (i.e., the perceived ease of a cognitive task; Alter & Oppenheimer, 2009), producing a systematic inflation of ratings that come later in a series. The more evaluations they make, the more evaluators perceive the process itself as fluent. As the objects of evaluation are perceived with more fluency, individuals’ judgments become more positive (Reber, Schwarz, & Winkielman, 2004). Similarly, individuals can also experience processing fluency, leading them to become more confident in their attributions (Petty, Briñol, & Tormala, 2002), to believe statements are more true (McGlone & Tofighbakhsh, 2000; Werth & Strack, 2003), and to prefer more familiar options (Bornstein, 1989). Yet these examples and related work on fluency have focused primarily either on objects, such as the familiarity of stable evaluation targets (Zajonc, 1968), or on external features that influence the evaluation process, such as the readability of text font, syntax complexity, and other perceptual cues that sway metacognitive perceptions (Alter & Oppenheimer, 2009).
To push this understanding further, we hypothesized about and investigated a pervasive but previously unexplored influence on cognitive judgment that is embedded in sequential review itself. We propose that the process of repeated evaluations is akin to patterns of misattribution. Throughout a sequence, as the evaluation process becomes more fluent, experiences of cognitive ease about the evaluation process will be misattributed to the content of the target itself. Similar to misattribution, cognitive fluency or ease of processing serves as a peripheral cue that can influence judgments in virtually any domain or situation when it is used as an inference about a judgment target (Oppenheimer, 2008). We extended these findings by testing whether sequential evaluations produce cognitive fluency, which in turn increases evaluations.
As the metacognitive perceptions of fluency increase, individuals tend to rely more on heuristic processing (Alter, Oppenheimer, Epley, & Eyre, 2007), leading to a convergence on default responses (Levav, Heitmann, Herrmann, & Iyengar, 2010; Litt, Reich, Maymin, & Shiv, 2011). Often a judgment task itself elicits a specific “naive theory,” or inference rule, linking experiences of fluency with default judgments of targets in that domain (see Schwarz, 2004). For example, when evaluating the intelligence of a writer, individuals who experience written text more fluently evaluate the writer’s intelligence more positively (Oppenheimer, 2006). We argue that cognitive fluency from repeated evaluations will produce similar inferences when one evaluates, for example, dancers in a dance competition, academic performance in a class setting, or stories in a writing contest.
Fluency should elevate judgments consistent with a domain-specific naive theory, but only if judges do not spontaneously discount those experiences, such as when an obvious alternative attribution is available (e.g., difficult-to-read text printed with an empty printer toner; Oppenheimer, 2006, Study 5; see also Oppenheimer, 2004) or when a manipulation is more overt (Bornstein & D’Agostino, 1994). In short, we predicted that individuals would report increased fluency in repeated evaluations because the experience feels easier to process, while at the same time they would remain unaware of the effects that fluency has on their own judgments. Thus, repeated evaluations will increase processing fluency, which will be used as a misattributed inference to evaluate novel targets—both people and objects—more positively over time.
Overview of Studies
In four studies, across more than 12,000 sequential evaluation events, we examined judgments spanning a range of time frames and broadly diverse contexts. In Study 1, we examined professional evaluations of contestants in 20 seasons of a popular television show and, in Study 2, we examined undergraduate university students’ grades over a 10-year period when instructors taught the same course multiple times. In such real-world contexts, other plausible influences likely co-occur with our proposed mechanism to influence focal outcomes. To address these, we included a number of robustness checks to rule out potential alternative explanations. In Study 3, we conducted a controlled experiment in which we randomized short stories over multiple weeks, and in Study 4, we directly tested our proposed theoretical mechanism of perceived process fluency. Finally, we report four additional studies in the Supplemental Material available online.
Study 1
Method
Dancing With the Stars is a TV reality show where celebrities partner with professional dance partners and compete in choreographed dances. The show has a set of three judges who have served through the past 20 seasons; other judges typically serve for a single season before being replaced. Each team performs multiple dances in each episode and is rated on a 10-point scale. Sometimes judges will not rate a team that they have coached. Teams are eliminated as the season progresses, and a set of finalists remain for the top three spots.
We asked a research assistant, who was blind to the study hypothesis, to record the judges’ scores of the teams in each episode for all available seasons. Each contestant dyad performed multiple dances per episode and received a rating from 1 to 10. Across the 20 seasons available, we collected a total of 5,856 scores, of which 5,511 (94%) were provided by the three permanent judges. We analyzed whether these 5,511 scores changed as the judges made sequential evaluations over time.
Results
Evaluation
To test our primary hypothesis that evaluations become more positive with experience, we tested for the linear relationship of evaluations over seasons and found that scores increased significantly as a function of season (1–20) across the entire series, b = 0.02, F(1, 5509) = 45.11, p < .0001, η p 2 = .001 (Table 1). To ensure that these results were not driven by outliers, we conducted a robustness check in which we split the data into two halves: Seasons 1 through 10 and Seasons 11 through 20. Analyzing scores across the two halves, we found that scores in the second half of the seasons (M = 8.18, SD = 1.36) were significantly higher than those in the first half (M = 7.87, SD = 1.50), F(1, 5509) = 62.65, p < .0001. Next, we address several potential alternative explanations.
Study 1: Permanent Judges’ Scores of Dancing With the Stars Contestants Across All Seasons
Professional partners improve over time
One poten-tial explanation is that the professional partner who is paired with the celebrity is improving over time. We addressed this explanation by controlling for the professional partner’s experience on the show. We found that scores increased across episodes even after we controlled for number of previous episodes that the professional partner has been in the show, b = 0.28, t(5508) = 57.05, p < .0001. Importantly, the correlation between the episode number (range = 1–12) and the number of previous episodes the professional partner has appeared in (0–147) was low (R2 = .005). We also conducted a similar analysis on the effect of seasons on evaluations, while controlling for the number of previous seasons that the professional partner has appeared in. In this analysis, the correlation between the number of seasons (1–20) and the professional partner’s season experience (0–18) was moderately high (R2 = .32). In spite of this correlation, the effect of seasons on scores remained significant after controlling for the professional partners’ season experience, b = 0.01, t(5508) = 2.23, p = .0256. We note that the covariate (the professional partner’s season experience) also has other issues, specifically in terms of direction of causality—professionals who score well in a season may be more likely to “survive” to subsequent seasons.
More skilled dancers appear in later seasons
The second alternative explanation is that the show is attracting higher-quality dancers (professionals or celebrities) across successive seasons. To address the issue that better professional dancers hired in later seasons may have accounted for inflated evaluations, we restricted our analysis to the 13 (out of 42) professional dancers who appeared in at least one of the first 10 seasons as well as in at least one of the last 10 seasons. These dancers accounted for approximately 69% of the data (3,821 of 5,511 observations) and appeared, on average, in 11.5 seasons. For these repeating dance partners, we found the same positive effect of season on evaluations, b = 0.018, t(3817) = 3.96, p < .0001.
Furthermore, according to U.S. Nielsen ratings, the show’s popularity has declined, with the last five seasons having the lowest viewership on record. 1 Consequently, if celebrities are attracted to popular programs, the shine of Dancing With the Stars appears to be waning. These observations run counter to the argument of increasing quality of celebrity participants.
Last, we analyzed the effect of seasons on scores provided by the transient guest judges (~6% of all observations). We found a directionally negative (but nonsignificant) effect of season on evaluations for temporary judges, b = −0.05, F(1, 343) = 1.95, p = .1638. We coded an indicator as 1 for temporary judges and as 0 for permanent judges. Note that the effect of season on evaluations was significant for each permanent judge analyzed individually: b1 = 0.02, t(1856) = 4.18, p < .0001; b2 = 0.02, t(1793) = 2.94, p < .001; b3 = 0.03, t(1856) = 4.38, p < .0001. We estimated a regression model with score as the dependent measure and the main effects of season and judge type (temporary or permanent), and the Season × Judge Type interaction as predictors. The model revealed a significant main effect of season, as before, b = 0.02, t(5852) = 6.74, p < .0001, and a significant main effect of judge type, b = 1.41, t(5852) = 2.13, p = .0330. Most importantly, the Season × Judge Type interaction was negative and marginally significant, suggesting that the effect of season was more negative for temporary (vs. permanent) judges, b = −0.07, t(5852) = −1.96, p = .05, although it is important to note that these two samples were not balanced in size. If anything, this evidence is suggestive that nonsequential evaluations showed no rising bias and that participant quality may be consistent or declining.
Discussion
Overall, these data suggest that as judges on Dancing With the Stars grow more accustomed to judging contestants, they evaluate them more favorably. Next, we considered a different evaluation context—grades provided by instructors—and investigated whether the average grades awarded in a class increase over successive course offerings. By exploring this context, we are also able to further address several possible alternative explanations raised in Study 1.
Study 2
Method
We acquired all available archival grades for all courses offered in a specific school (one of the top five U.S. schools specific to the discipline of study) at a large U.S. university (one of the top public universities in the United States), covering all successive spring and fall semesters from Spring 2000 through Spring 2009 (e.g., Spring 2000, Fall 2000, Spring 2001). We converted the archival data (from printouts, Word documents, and Excel files) into a data set that included 1,854 courses offered over this 10-year period. We then excluded 34 courses that did not list any instructor (independent studies, individual instruction courses, and three courses with no instructor of record) and another 24 courses where no grades were assigned (noncredit or pass/fail courses), leaving 1,796 graded courses in the data set for which we had an instructor. We next excluded a specific set of courses that were team taught and jointly graded, where only one of the two instructors was listed. These courses presented a challenge because we could not attribute the grade to a specific instructor, and one of the two instructors was not on the records. There were 438 sections of these courses in the data set. Excluding these courses left us with 1,358 course sections.
We started with these 1,358 sections and identified those courses where at least four sections were offered by a specific instructor. We used this cutoff because up to three sections of the course could be simultaneously offered by an instructor in a semester, in which case we would not be able to test the order effect from one semester to the next. We included only those courses where at least some of the sections exceeded 15 students; our data therefore have a few instances where a specific section size drops below 15 students. Using these criteria led us to identify 991 sections (73% of the data) where instructors offered the course over several semesters. For these courses, we had information about the class grade point average (GPA), class size, instructor, semester (fall, spring), and year (2000–2009). The cumulative class size was 45,292 students (~46 students per section).
For each of these 991 course sections, we coded a course-order variable that was set at 1 the first time a course was offered in our data set by a specific instructor, at 2 the next time the instructor offered that course, and so on. When multiple sections of a course were offered by an instructor in the same semester, they had the same course-order value. A few times an instructor would offer one section of a course to start and would then offer two sections the next time around. This led to a slightly smaller number of sections with an order of 1 than of sections with an order of 2. We used order as the primary independent variable in our analysis. Table S1 in the Supplemental Material shows the average class GPA for successive offerings of courses in our data set.
Results
GPA over successive offerings
As shown in Figure 1 and consistent with findings from our first study, a regression with course GPA as the dependent measure and course order (the successive offering of the course by an instructor) as the predictor revealed a significant effect of order, b = 0.016, F(1, 989) = 45.22, p < .0001, η p 2 = .04. Next, we added covariates that were expected to predict GPA differences. We conducted an analysis of covariance that included GPA as the dependent measure and course order, class size, class level (1 = freshman, 2 = sophomore, 3 = junior, 4 = senior, 5 = graduate, treated as a categorical variable), and semester (fall, spring) as predictors. The main effect of the order of successive offerings on GPA was positive and significant, b = 0.02, F(1, 983) = 108.37, p < .0001, consistent with the effect seen in the simple regression. Class level was also a significant predictor of GPA, F(4, 983) = 98.15, p < .0001, with Level 4 classes having the highest GPA (see Table S2 in the Supplemental Material for the cell means). Spring semester classes (M = 3.48, SD = 0.28) had higher GPAs than fall semester classes (M = 3.42, SD = 0.30), F(1, 983) = 18.19, p < .0001. The effect of class size on GPA was also significant, with larger classes having a lower GPA, b = −0.0006, F(1, 983) = 5.65, p = .0177.

Study 2: increase in class grade point average (GPA) over successive course offerings, drawn from a sample of 1,854 courses at a top public U.S. university and controlling for class size, class level, and semester.
Excluding the long tail
Table S1 reveals that the number of sections dropped steeply as order increased, with the second half of the series (order > 10) including only 9% (90/991) of all the sections. To ensure that our results were not being skewed by this small subset of the data, we restricted our analysis to sections where order was less than 11, 10, 9, and so on (see Table S3 in the Supplemental Material). The effect of course order on GPA remained consistent with that observed for the full data set.
Are courses that award higher grades more likely to be offered again?
Another alternative explanation is that because we were analyzing the data across courses, the results were driven by the surviving courses (which are offered across a greater number of semesters). If more generous courses are in greater demand and survive longer, this will lead to a positive effect of course order on GPA. We controlled for this explanation by restricting our analysis to only those courses that are offered at least 10 times and testing the effect of course order on the first 10 offerings. The regression revealed that the effect of course order on GPA for this smaller sample remained positive and significant, b = 0.018, F(1, 288) = 10.54, p = .0013. Table S4 in the Supplemental Material presents similar analyses for different cutoffs. Furthermore, a repeated measures analysis on the complete data that controls for the course/instructor combination also revealed a positive effect of class order, b = 0.010, F(1, 902) = 37.23, p < .0001.
Does instructor improvement lead to better student performance?
Perhaps our results are a consequence of instructors learning over time and providing a better quality of instruction in successive sections. If GPA is an objective measure of student learning, this should lead to an increase in GPA over successive course offerings from the same instructor. While we acknowledge this as a potential alternative explanation, we would expect this learning effect to be more pronounced when an instructor first offers a course (say, in the first three offerings) rather than in later sections. Instead, our data suggested that the effect persisted even in later sections of the course offerings (see Table S1), consistent with other findings that teaching effectiveness, if anything, tends to decline with age and years of experience without systematic intervention (for a review, see Marsh, 2007). Furthermore, the conditional analysis in Table S4 indicated that GPA increases are more pronounced for courses that are offered at least four times. These patterns suggest a process different from that of increasing instructor quality.
Does student quality increase each year?
We ruled out the alternative explanation that student quality increases each year by controlling for calendar year. The correlation between course order and calendar year was moderate (R2 = .27); even after controlling for the calendar year (2000–2009) we found that the effect of course order on GPA remained significant, b = 0.016, F(1, 988) = 22.61, p < .0001. Notably, the effect of calendar year on GPA was not significant, p = .8341; if student quality increased over successive years, the effect of calendar year on GPA should be positive and significant.
Discussion
Tracking grades across a 10-year period, we found that successive offerings of a course by an instructor have higher average grades. These rich data allowed us to address and rule out several alternative explanations. Next, we tested our predictions in a controlled setting in which we could randomly present different stimuli across a period of several weeks. Naturally, this setting is also less susceptible to potential alternative explanations posed by field and archival data.
Study 3
Method
We investigated these effects for sequential evaluations in a more controlled context while randomizing the order of evaluation targets. Student participants evaluated a series of 10 short stories, at the rate of 1 story per day over a 2-week period, as part of a research requirement. We expected that there would be some drop-off among participants across the 2-week period. Given the results of prior studies, we estimated a sample size of approximately 200 participants to complete the study, consistent with a power analysis using estimates for small effect sizes (e.g., η p 2 ≤ .05). To account for some drop-off, we recruited 210 participants to complete an initial survey opting into the full 2-week study.
First, undergraduate student participants signed up for the study and evaluated a sample story for the initial survey. The deadline to complete this story was the Sunday evening before the start of the 2-week period. Beginning the next day, on each subsequent working day of the next 2 weeks (Monday–Friday) they received 1 story to evaluate, for a total of 10 stories. We randomized the order in which each participant received the 10 stories over the 10 days, such that each participant evaluated a story only once. Each short story was evaluated on a 7-point scale (1 = very unfavorably, 7 = very favorably). On the last day of the study, we asked a set of retrospective trend perceptions to capture participants’ lay beliefs about their evaluations throughout the study (see Results) and their self-reported expertise.
To analyze the results, we included all data from participants who had completed at least the first day (to sign up) and the last day (for demographic and trend information). Some participants completed an evaluation twice on the same day for the same story; in these 24 instances, we examined the time stamps on the surveys and retained the first evaluation while discarding the second evaluation. Finally, a programming error led to 1 participant rating the same story on 2 days; in this case, we kept the first evaluation but excluded the second. In sum, we had 1,572 observations from 168 participants (average age = 19.45 years, SD = 1.90; ~53% male). On average, each participant evaluated approximately 9.4 stories. We calculated an order variable, which ranged from 1 to 10, that tracked the sequence of short stories evaluated by each individual. Table S5 in the Supplemental Material shows the frequency counts for this variable. Consistent with the preceding studies, Study 3 tested whether story evaluation varied as a function of evaluation order.
Results
Evaluation
A repeated measures analysis controlling for multiple responses from each participant and also for the differences between stories revealed that the main effect of (mean-centered) order was positive and significant, b = 0.032, F(1, 1394) = 5.44, p = .0198, η p 2 = .003. Table S6 in the Supplemental Material lists the story evaluations at each level of order. This pattern is plotted in Figure 2. The scores for each story are listed in Table S7 in the Supplemental Material. The effect of participants’ self-reported expertise on story evaluations (M = 4.20, SD = 1.42), was significant and positive, b = 0.095, t(1570) = 3.34, p = .001. The effect of order remained positive and significant, p = .025. The Order × Expertise interaction was not significant, p = .428.

Study 3: average raw evaluations of short stories (1 = very unfavorable, 7 = very favorable) for each order of evaluation.
Retrospective trend perceptions
On the last day of the 2-week study, we also captured participants’ retrospective lay beliefs about their evaluation processes. In particular, we asked them to report, “Throughout the study, the more stories I rated . . .” “. . . the easier it was to rate each story,” “. . . the quicker I evaluated each story,” “. . . the more I enjoyed evaluating each story,” and “. . . the more positively I rated each story” (1 = strongly disagree, 7 = strongly agree). These items were important because we predicted that individuals would experience metacognitive fluency (i.e., perceived ease of evaluations) as order increased but that individuals would still be unaware that order or fluency would lead to increases in their evaluations (a perception that might lead them to spontaneously discount their ratings; Oppenheimer, 2004). Consistent with our predictions, participants agreed that, across all days, the process of evaluating stories became easier (M = 4.90, SD = 1.34), quicker (M = 4.65, SD = 1.47), and more enjoyable (M = 4.36, SD = 1.61). Each of these measures was significantly above the scale midpoint of 4, all ps < .01. Importantly, however, participants disagreed with the lay belief that their evaluations became more positive over successive stories (M = 3.50, SD = 1.45), which was significantly below the scale midpoint, t(167) = −4.46, p < .001. In other words, participants in hindsight reported increased experiences of ease as they rated more stories, but they remained unaware that increases in order or felt ease would lead to more positive evaluations.
Discussion
In Study 3, we explored the effect of order on evaluations in a lab setting in which we manipulated sequential reviews directly over a period of multiple weeks. Even when participants evaluated a unique randomly selected story each day, participants as a group confirmed our hypotheses by evaluating stories later in the sequence as more favorable. These findings replicated the pattern of results from Studies 1 and 2 and provided important clarity in a more controlled setting, which helped to rule out alternatives such as quality of targets increasing over time. Next, we replicated this study and directly tested our proposed theoretical mechanism that drives such a positivity bias.
Study 4
Thus far, we had not tried to measure processing fluency directly after each evaluation or to interrupt it through manipulation. In Study 4, we tested whether fluency accounts for rises in sequential evaluations in these two ways. First, we measured fluency directly after each evaluation experience to test whether it mediated the relationship between order and evaluations, something we had not attempted previously. We did this by assessing metacognitive fluency (Alter & Oppenheimer, 2009) after each day’s story evaluation. Second, we attempted to interrupt fluency by manipulating it. One way to do this is to interrupt processing fluency by asking participants to elaborate on their story evaluation (vs. no elaboration). However, one attempt at this manipulation in a pilot study (see Study S4 in the Supplemental Material) failed to moderate the relationship between order and evaluations, although in that study we did find significant evidence for the same positive main effect of order on evaluations as in all other studies. As a result, we explored a different approach to disrupt fluency by asking all participants to elaborate on their evaluations, but to shift the evaluation criteria daily for some participants: simple elaboration (fluent) versus complex elaboration of a randomly selected specific evaluation criterion each day (disfluent).
Recall our prediction that repeated evaluations would increase the experience of processing fluency, which would serve as an inferential cue that leads to more positive evaluations (and mediates the relationship between order and more positive evaluations). As one potential alternative mechanism—that participants experience greater flow the more they evaluate stories—we also directly measured perceptions of flow after each evaluation (Csikszentmihalyi & Rathunde, 1993).
Method
We used TurkPrime to recruit panelists from Amazon’s Mechanical Turk (Litman, Robinson, & Abberbock, 2017) to evaluate a series of 10 short stories at the rate of 1 story per day over a 2-week period (Monday–Friday). These short stories were drawn from the same pool used in previous studies (see Table S8 in the Supplemental Material). To account for potential drop-off across the 2 weeks and to explore the moderating effect of elaboration on evaluations, we increased the sample size from prior studies and recruited 706 responses to a screener questionnaire describing the study and asking interested individuals to opt into it. A total of 642 individuals (~91%) opted into the study. Beginning the next day, on each subsequent working day of the next 2 weeks (Monday–Friday), these 642 people received 1 story to evaluate, for a total of 10 stories. As in Study 3, we randomized the order in which each participant received the 10 stories over the 10 days so that each participant evaluated a story only once.
Each participant was randomly assigned to one of two elaboration conditions, between subjects, at the beginning of the study. In the simple elaboration condition, participants were asked to write a few sentences about what they liked or did not like about each story on each of the 10 days. In the complex elaboration condition, participants were asked to write what they liked or did not like about one specific dimension of the story (the characters, descriptions, emotions, ending, message, or plot). Each day, one of these dimensions was presented at random so that individuals in the complex elaboration condition would be evaluating stories on different dimensions as they progressed through the 10-day task. The dimensions were chosen on the basis of a text analysis of open-ended elaborations provided by participants in a preceding pilot study (Study S4). This served as our primary attempt to manipulate fluency directly. After the elaboration, participants evaluated each story on a 7-point scale (1 = very unfavorably, 7 = very favorably).
To directly measure increases in fluency over time, we asked participants to provide responses to process measures for fluency and flow after the story evaluation each day, allowing us to test for mediation of order on evaluations through both fluency and flow. At the end of the last day’s survey, we asked participants to guess the purpose of the study (as an open-ended response). Finally, participants completed the same set of retrospective trend-perception measures over the 10-day period as in Study 3, in addition to a 10-item Positive and Negative Affect Schedule (Watson, Clark, & Tellegen, 1988), their self-reported expertise at evaluating stories, and demographic measures.
Of the 642 invited individuals, only 518 people completed at least 1 of the 10 days’ evaluations. For our analysis, we included all data from participants who had completed at least the last day (for demographic and retrospective measures); there were 362 such participants (our reported results carry through if we use the larger sample of 518 people). Some participants completed an evaluation twice on the same day for the same story; in these six instances, we examined the time stamps on the surveys and retained the first evaluation while discarding the second evaluation. In sum, we had 3,414 observations from 362 participants (average age = 40.88 years, SD = 13.53; ~56% male). We calculated an order variable, which ranged from 1 to 10, that tracked the sequence of short stories evaluated by each individual. Table S9 in the Supplemental Material shows the frequency counts for this order variable.
Consistent with the preceding studies, we tested whether story evaluation varied as a function of the evaluation order. In addition, this study allowed us to investigate whether the type of elaboration (simple vs. complex) moderates the effect of order. Finally, the fluency and flow process measures each day enabled us to test whether these variables mediate the effect of order on evaluations.
Results
Evaluation
First, we adjusted the story evaluations to control for the effect of individual stories (Table S8 lists mean differences in evaluations across stories). A repeated measures analysis, controlling for multiple responses from each participant, revealed that the main effect of (mean-centered) order on the adjusted evaluations was positive and significant, b = 0.054, F(1, 3050) = 28.44, p < .0001, η p 2 = .01. The main effect of elaboration type was not significant, F(1, 360) = 0.25, p = .6207. The Order × Elaboration Type interaction was also not significant, F(1, 3050) = 0.04, p = .8354. Thus, the type of elaboration (simple vs. complex) did not moderate the effect of order on evaluation. This pattern of effects remained unchanged if we used unadjusted evaluations as the dependent measure and controlled for the main effect of story. Table S10 in the Supplemental Material shows the mean unadjusted evaluation at each level of order.
Fluency
We calculated a fluency score for each individual’s perception of each day’s evaluation process as an average of three items (easy, quickly, enjoyed; see Oppenheimer, 2006; Cronbach’s α = .66; M = 5.34, SD = 1.31). We conducted a repeated measures analysis with fluency as the dependent measure. The analysis revealed that the main effect of order on fluency was positive and significant, b = 0.051, F(1, 3050) = 66.06, p < .0001. The main effect of elaboration type was not significant, F(1, 360) = 2.06, p = .1524. The Order × Elaboration Type interaction was also not significant, F(1, 3050) = 1.85, p = .1742.
Flow
The flow measure for each individual was an average of 10 items (adapted from Fu, Su, & Yu, 2009; Cronbach’s α = .87; M = 6.35, SD = 0.74). A repeated measures analysis with flow as the dependent measure revealed that the main effect of order on flow was positive and significant, b = 0.028, F(1, 3050) = 138.46, p < .0001. The main effect of elaboration type (contrast coded; simple = −1, complex = 1) was also significant, indicating a lower flow score for complex elaborations, b = −0.073, F(1, 360) = 4.67, p = .0313. The Order × Elaboration Type interaction was not significant, F(1, 3050) = 1.22, p = .2699. Table S11 in the Supplemental Material lists fluency and flow means for each level of order.
Mediation
We tested whether fluency and flow mediated the effect of order on evaluations. To do so, we used the adjusted evaluation score that controlled for the fixed effect of each story. Because elaboration type did not moderate the effect of order on evaluation or moderate any of the mediation results below, we collapsed across the two elaboration conditions. We tested whether fluency and flow mediated the effect of order on evaluation by computing the bias-corrected bootstrapped 95% confidence interval (CI) around the indirect effects with 5,000 resamples in a multiple mediator model (Hayes, 2013; Preacher & Hayes, 2004). Consistent with our prediction, we found a significant and positive indirect effect through perceptions of fluency, b = 0.022, 95% CI = [0.020, 0.037], whereas the indirect effect through perceived flow was not significant, b = 0.002, 95% CI = [−0.001, 0.002]. As individuals evaluated more stories, they perceived the evaluation process as becoming more fluent, which increased evaluations across the sequence.
Open-ended coding mediation
Three research assistants blind to hypotheses coded the open-ended responses from participant evaluations of each short story across all 10 days. Each coder reported the number of positive and negative comments written by participants about each story. We created a percentage score of positive comments by subtracting the number of negative comments from the number of positive comments for each evaluation, then dividing that difference score by the total number of comments for that evaluation. Scores ranged from −1 to 1, where higher numbers indicate a greater percentage of positive comments.
We calculated this score for all story evaluations coded by each of the three research assistants. There was high agreement (Cronbach’s α = .98), so we averaged the three scores to create a single measure indicating the percentage of positive comments made for each story evaluation.
We explored whether evaluation order would influence the percentage of positive comments made about a story. Consistent with findings reported above, order had a significant and positive effect on the percentage of positive comments written about each story, b = 0.0190, t(3403) = 3.759, p < .001. In other words, the more stories that participants rated, the greater percentage of positive comments they wrote about each story. Furthermore, controlling for the effect of order, the percentage of positive comments made about a story had a significant and positive effect on evaluations, b = 1.533, t(3400) = 58.47, p < .0001. The effect of order was still positive and significant, p = .0018. Last, we tested a mediation model by computing the bias-corrected bootstrapped 95% CI around the indirect effect with 5,000 resamples. We found a significant and positive indirect effect of order on evaluation through percentage of positive comments, b = 0.0291, 95% CI = [0.0129, 0.0440].
Next, we ran a serial mediation model by including fluency in the model above to test whether order increases fluency, fluency increases the percentage of positive comments written about a story, which increases the story’s evaluation score. Results supported this model. As noted previously, the effect of order on fluency was positive and significant. Controlling for order, the effect of fluency on percentage of positive comments was also positive and significant, b = 0.2580, t(3400) = 25.75, p < .0001. Last, controlling for both order and fluency, the effect of positive comments on story evaluation was also positive and significant, b = 1.469, t(3399) = 49.96, p < .0001. We tested the serial mediation model by computing the bias-corrected bootstrapped 95% CI around these indirect effects with 5,000 resamples. The indirect effect of order on evaluation through fluency, then through percentage of positive comments, was positive and significant, b = 0.0182, 95% CI = [0.0124, 0.0241]. The indirect effect through fluency alone was also positive and significant, albeit weaker, b = 0.010, 95% CI = [0.0067, 0.0138], and the indirect effect through positive comments alone was not significant, b = 0.0085, 95% CI = [−0.0048, 0.0218]. In summary, fluency mediated the effect of order on evaluations, both directly and through the percentage of positive comments.
Retrospective trend perceptions
As in Study 3, we asked participants their lay beliefs about their experiences at the end of the study. On average, participants at the end of 10 days agreed that as they rated more stories, the evaluation process became easier (M = 5.96, SD = 1.37), quicker (M = 5.07, SD = 1.84), and more enjoyable (M = 5.61, SD = 1.72). Each of these measures was significantly above the scale midpoint of 4, all ps < .001. Importantly, however, and as in Study 3, they disagreed with the belief that their evaluations became more positive over time (M = 3.30, SD = 1.65), which was significantly below the scale midpoint, t(360) = −8.08, p < .001.
Perceptions of study purpose
In addition, the same three research assistants who coded the open-ended responses also coded the 362 participants’ responses about the study’s purpose, which we collected at the end of the last day. They manually coded two scores—(a) whether participants mentioned something about our focal mechanism (i.e., predictions about ease or speed of evaluations) and (b) questions about the study’s overall purpose—using five categories of responses, including two focal categories: studying something about evaluating stories over time and about the effect of order on more positive evaluations (e.g., guessing our hypothesis). Coders agreed on 99% of responses for the first coding about guessing the study’s focal mechanism; the three responses where one coder disagreed were resolved by using the score agreed on by the other two coders. Across all responses, only 2 participants (0.6%) mentioned ease or speed, and neither mentioned it in relation to order or more positive evaluations.
For the second measure (the study’s purpose), all three coders agreed on 74% of responses. We resolved the 25% of cases where one coder disagreed by selecting the score given by the other two coders. On the two instances (0.6%) where all three coders disagreed, the authors selected the most appropriate response of one coder. Among all responses, 29.6% were unsure about the study’s purpose; 51.4% reported that our aim was to test something about authors, short-story content, or writing quality (e.g., what makes a good writer or short story); 13.3% guessed incorrectly that we were interested in studying the effect of specific stories on some outcome (e.g., the effect of different kinds of stories and themes on participants’ mood); 5.5% reported something about evaluations changing over time (e.g., if attention would diminish if more stories were evaluated: “Possibly to see how throughout a given week people’s attention span and opinion of the stories shifts”); and 1 participant wrote about evaluations becoming more positive over time (“I think it was to study how the reviews of the short stories improved over time”). In sum, across all open-ended responses, few, if any, participants described either our focal mechanism or our predicted relationship between order and evaluations.
Affect and self-reported expertise
Last, we explored self-reported affect and expertise. We first reverse-scored the five negative-affect items and averaged these with the five positive-affect items to create a 10-item positive-affect scale (Cronbach’s α = .78; M = 5.24, SD = 0.92). The effect of positive affect on story evaluation was positive and significant, b = 0.265, t(3412) = 7.86, p < .001; the main effect of order remained positive and significant, p < .0001; and the Order × Positive Affect interaction was also positive and significant, b = 0.028, t(3410) = 2.386, p < .017. The effect of self-reported level of expertise (M = 3.97, SD = 1.69) was also positive and significant, b = 0.018, t(3404) = 2.40, p = .017; the effect of order remained positive and significant, p < .0001; and the Order × Expertise interaction was not significant, t(3402) = 0.221, p = .825.
Discussion
In Study 4, we continued to find support for the hypothesis that sequential review, even with randomized targets, increases evaluations. Importantly, we tested for our proposed mechanism in two ways. Despite finding no interaction of elaboration complexity to disrupt fluency experiences, we did find significant evidence of mediation for our predicted process measure of experienced fluency and not flow. By capturing direct measures after each day’s study over a 2-week period, we showed that participants’ metacognitive fluency experiences mediate the effect of order on evaluations.
Meta-Analysis of Studies 1 to 4
To explore the robustness of findings across all studies, we conducted a meta-analysis of our results. Because of variation in the evaluation measures (Study 1 dancer rating: 1–10; Study 2 course GPA: 0–4; Studies 3 and 4 story evaluation: 1–7), we z-scored the dependent measure for each study (evaluation). The dependent measures for Studies 3 and 4 were adjusted to control for the story effect before they were z-scored. Similarly, because of variation in the predictor variables (Study 1: 1–20 seasons; Study 2: 1–19 courses; Studies 3 and 4: Days 1–10), we z-scored the predictor variable (order). We calculated the effect size two ways: a generalized linear model, controlling for participant ID, η p 2 = .009, and a mixed model that allowed repeated measures, η p 2 = .008. Both models revealed similar results with a positive and significant effect of order. See Table S12 in the Supplemental Material for effect-size estimates for all studies.
General Discussion
These studies call attention to people’s reliance on experienced evaluators across a range of contexts. Sequential evaluation is assumed to be a benchmark of fair review because it holds the rater and judgment standards stable for all targets of evaluation. As individuals become more familiar with a decision-making process, they are likely to perceive the process as more fluent. This metacognitive experience then serves as an inference that increases the likelihood of a positive response specific to the domain in question (Levav et al., 2010; Schwarz, 2004). These findings present an important new bias that is likely to be widely pervasive given the prevalence of comparable sequential evaluation systems.
While we cannot be certain of all external influences on such evaluations, we took particular care to rule out or control for several plausible alternative explanations. Even when we controlled for these and other factors, the effect of order persisted in all studies, and importantly, we documented that perceived fluency, directly assessed after each evaluation, mediates this effect.
We acknowledge that the effect sizes found here, while remarkably robust and consistent across studies, contexts, and settings, are indeed relatively small. Nonetheless, these effects are for each sequential evaluation and can still play an important role in decisions (Prentice & Miller, 1992), especially over several evaluations. In Study 2, for example, the course GPA over successive offerings rose approximately from a B+ to an A–.
We note that some researchers have found evaluations becoming either more negative (Danziger, Levav, & Avnaim-Pesso, 2011) or more positive (de Bruin, 2005) similarly as a result of heuristic inferences of evaluations in a single sitting, yet without influencing later evaluations. Importantly, and by contrast, we found that the effect of sequential evaluations may span many years and implicates evaluators’ increasing experience with the evaluation process. One factor determining the direction of fluency effects is the domain-specific naive theory that individuals bring to mind (Schwarz, 2004). Consequently, not just being sensitive to experiences of fluency but understanding one’s own naive theories about how metacognitive processes influence outcomes may be enough to diminish or even reverse these effects (Oppenheimer, 2004).
In all examples, the extraneous influence of order was remarkably inextricable from the decision process itself. Even hallmarks of objectivity and fairness such as a coin flip or randomized order may not restore objectivity for targets where chance landed them in the final performance slot or toward the bottom of the pile of exams to be graded. Our findings unmask an unsuspected culprit embedded in every corner of daily life from the trivial to the life altering: that the decision process itself might contaminate people’s evaluations as they become more experienced with it.
Supplemental Material
OConnorOpenPracticesDisclosure – Supplemental material for Do Evaluations Rise With Experience?
Supplemental material, OConnorOpenPracticesDisclosure for Do Evaluations Rise With Experience? by Kieran O’Connor and Amar Cheema in Psychological Science
Footnotes
Acknowledgements
The authors thank Michael Atchison, Thomas Bateman, James Burroughs, Benjamin Converse, Steven Johnson, Timothy Wilson, and Carl Zeithaml for helpful comments.
Action Editor
Marc J. Buehner served as action editor for this article.
Author Contributions
Both authors contributed equally to this article. Coding of contestant scores for Study 1 was supervised by K. O’Connor. Both authors transcribed the historical grades data for Study 2. K. O’Connor conducted Studies 3 and 4. A. Cheema analyzed data for all studies. Study 4 mediation analyses were conducted by K. O’Connor. Both authors wrote the manuscript and approved the final manuscript for publication.
Declaration of Conflicting Interests
The author(s) declared that there were no conflicts of interest with respect to the authorship or the publication of this article.
Funding
Research support from the McIntire School of Commerce is gratefully acknowledged.
Open Practices
Data and materials for Studies 1, 3, 4 and S1 to S4 have been made publicly available via the Open Science Framework and can be accessed at https://osf.io/ugdnv/ (for each study, data are on the “data” tab, and materials are on the “variables” tab). Grade data from Study 2 are confidential and not available for dissemination. The complete Open Practices Disclosure for this article can be found at http://journals.sagepub.com/doi/suppl/10.1177/0956797617744517. This article has received the badges for Open Data and Open Materials. More information about the Open Practices badges can be found at
.
Notes
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
