Abstract
As the evidence linking test scores to long-run student outcomes has grown, standardized assessments have become a widely used management tool in education including addressing the racial education gap. One of the concerns with the use of standards tests is the perverse incentive for teachers to alter test scores. The consequences for educators found cheating can be substantial, although little is known about how students are impacted by cheating. Using an 11-year panel of individual-level data on students and teachers from a predominantly Black urban school district where widespread test-score manipulation occurred, we investigated the impact of teacher cheating on subsequent student test scores. To access the impact, we used school-grade and classroom fixed effects as well as measure potential omitted variable bias (OVB). We found that for each additional wrong-to-right altered test question it is associated with a reduced future achievement of between 0.003 and 0.014 standard deviations depending on the specification. Although the evidence from OVB analysis does not suggest that test-score manipulation itself harmed or benefited students. Our evidence also contributes to a growing literature on the importance of sensitivity tests for OVB. We show how failure to conduct such tests could lead to erroneous findings.
Keywords
Introduction
Standardized assessments have become an important management tool in elementary and secondary education especially to address the long-standing Black-White education test score gap. While test scores are not a socially-relevant end to themselves, robust evidence suggests that high-quality teachers and smaller class sizes raise student test scores in a way that transmits to a range of long-run outcomes society has an interest in improving: teen pregnancy, incarceration rates, educational attainment, and labor market earnings (Chetty et al., 2011, 2014; Dobbie and Fryer Jr, 2015; Angrist et al., 2016), all which are important outcomes to address racial inequality. 1 Recognizing their predictive value, administrators now use test scores to evaluate teachers, determine salaries, identify targets for state takeover, measure racial/ethnic academic gap, and close underperforming schools. With test scores driving such consequential management decisions, perverse incentives have arisen to raise test scores by any means possible, even in ways that may not transmit to improved long-run outcomes for students. In the United States we have seen an increase of cheating cases, in extreme, teachers and school leaders have even orchestrated test-score manipulation schemes to create an illusion of success (Judd, 2012; Perry et al., 2012). 2 Understanding the impact of cheating is a crucial question not only because they reduce the usefulness of test scores, but also because students are potentially harmed when teachers cheat.
In this paper, we seek to measure the effects of widespread test-score manipulation in a predominately Black district on future student outcomes. In particular, we ask what is the impact of teachers altering students test responses (and hence their scores) on students future test scores. Conventional wisdom has held that students are indeed hurt by cheating, and judges have cited this harm when sentencing educators to jail (Sarrio, 2015). While extant research has developed methods for detecting test-score manipulation by teachers or administrators (Jacob and Levitt, 2003a,b), few studies have attempted to measure how students are impacted by the falsification of scores. To measure cheating’s impact, we draw upon an 11-year panel of student-level data from an anonymous predominantly Black urban school district where documented teacher cheating occurred at a large scale. We use school-grade and classroom fixed effects to disentangle the consequences of test-score manipulation from differences in school and teacher quality potentially associated with the choice to cheat. We also implement sensitivity tests for omitted variable bias (Altonji et al., 2005; Oster, 2019).
State and national policymakers have an interest in ensuring the accuracy of exams whether or not cheating harms students. Unless exams are accurate, they will not be useful as oversight tools. Therefore, even if test-score manipulation has little impact on students, we doubt the finding would lead to a reduction in monitoring. However, if cheating does cause long-term harm to students, the evidence may motivate the allocation of additional resources to its prevention. Knowledge about cheating’s effect would also be helpful for district policymakers unsure what to do in the aftermath of cheating’s discovery. Jacob and Levitt (2003b) find that cheating occurred in approximately four to five percent of Chicago’s classrooms. A more recent investigation found evidence of cheating in 15 central cities of large U.S. metro areas (Perry et al., 2012). 3 If students are harmed by test score manipulation, once district leaders discover its existence, they may want to direct resources to those impacted by a teacher’s nefarious actions. Alternatively, if cheating is more a symptom of low-functioning schools than a direct cause of harm, districts will be better served to direct their energies toward broad improvements of school leadership and teacher quality that will flow to all students rather than targeting efforts at specific students whose answers were changed.
Our results show statistically significant associations between the number of wrong-to-right erasures and future test scores even when school and teacher quality is held constant. We find that for each additional wrong-to-right altered test question this is associated with a reduced future achievement of between 0.003 and 0.014 standard deviations depending on the model specified. However, measures of omitted variable bias persuade us that the act of changing students’ test responses had little effect on the students themselves. Notably, our model without fixed effects is robust to tests for omitted variable bias. We interpret this as suggestive evidence that being in a school environment where cheating was more prevalent reduces students’ future scores. However, it is unclear whether cheating is the root cause or if schools unsuccessful at improving achievement in more constructive ways were simply more likely to turn to cheat. In the latter case, we would have observed a reduction in these students’ future scores even if a student’s answers were never changed. In addition to the evidence on how students are impacted by cheating, our work also contributes to a growing literature on the importance of tests for omitted variable bias in applied research. A researcher simply comparing the stability of coefficients across our models may conclude that the stable coefficients represent strong evidence that test-score manipulation causally reduced future student achievement. Sensitivity tests for selection bias (Altonji et al., 2005; Oster, 2019) persuade us otherwise. While much of the research relying on such tests have been conducted with survey data, our setting demonstrates the usefulness of such tests even when rich administrative data are available. We show how researchers can use knowledge about likely omitted variables to tailor such tests and the required assumptions to a given context.
The paper proceeds as follows: section 2. provides background information and our data, section 3. is a literature review, section 4. explains our research design and econometric models, section 5. presents our results, and the final section concludes.
Context and Data
Our source of data was an urban school district where documented test-score manipulation occurred on a large scale. Allegations of widespread cheating first became public in 2009, and in early 2010 the state conducted an analysis of erasures on the state-mandated criterion-referenced exams (CREs). Classes were “flagged” based on high numbers of wrong-to-right (WTR) erasures and schools were categorized based on the proportion of flagged classrooms in the school.
Schools where cheating was most prevalent were identified for detailed investigation (henceforth “investigated schools”), which included interviews with school personnel. Just over 60 percent of the district’s elementary and middle schools received a detailed investigation. In over half of these schools, educators confessed to cheating and investigators concluded that systemic misconduct occurred in over three-fourths of the schools that were investigated in detail. The investigation also revealed that cheating had been going on for some time, perhaps as far back as 2001 in some schools. 4
Figure 1 illustrates the magnitude of cheating in the district; students in flagged classrooms in 2009 had substantially higher normalized scores (on the order of 0.5 standard deviations) than their own scores in 2010. This stands in stark contrast to students in investigated schools but not in flagged classrooms who achieved largely consistent scores in 2009 and 2010.

Distribution of CRE Scores in 2009 and 2010 by Flagged Status.
Using data provided by the district, we constructed a longitudinal data set covering the 2005 to 2016 school years. The district files included individual-level information on enrollment, attendance, disciplinary incidents, withdrawal from school, high school diploma receipt, student demographics and program participation. In addition, we obtained test results from state-mandated CREs administered annually in grades 1–8 in five different subjects. Finally, we received scores from a low stakes national normed-reference test (NRT) that was administered by the district in grades 3, 5 and 8 in each of the years 2006–2008.
The CRE testing data included the individual-level number of correct answers as well as the number of WTR erasures for the years 2009 through 2013. 5 The NRT data included not only scaled scores and national curve equivalent scores; they also included individual item-level responses on each subject-area exam. 6
It is worth discussing which students were targeted by the cheating scheme. Policymakers may be particularly concerned when cheating is targeted at non-white, low-income, and low-achieving populations. Prior research has found non-white and less-empowered students to be particularly responsive to variation in education and neighborhood inputs (Krueger and Whitmore, 2001; Abdulkadiroğlu et al., 2016; Chyn, 2016). In the district we studied, cheating was more prevalent for black, low-income, and low-achieving students. Schools in the district with majority-white populations and those with higher income profiles did not see abnormal WTR erasures. Therefore, we believe our evidence came from the subset of students for whom, if cheating has an effect, its effect would likely be most pronounced.
Table 1 presents a set of descriptive statistics of our sample. The full sample represented all students in investigated schools. In our analysis sample, we did not include students in non-investigated schools, which serve substantially different student populations, both racially and economically. In the table the full sample or the investigated schools were about 96 percent Black and 78 percent received free or reduced lunch, an indicator of low-income status, hence our sample was composed predominately by low-income Black children. The table shows that students in flagged classrooms had approximately three times the number of WTR erasures than non-flagged classrooms. The flagged and non-flagged classrooms were broadly similar across other measures including ability (measured by the NRT), student demographics, and receipt of special education services.
Descriptive Statistics.
Note: The full sample is the universe of students in grades 3-8 attending investigated schools in 2009.
Related Literature
The Effect of Cheating on Students
Relatively little research has attempted to evaluate the impact of teacher cheating on student outcomes. Most of the cheating literature focused on the identification of cheating (Jacob and Levitt, 2003a,b; van der Linden and Jeon, 2012) or how to prevent it (Bertoni et al., 2013) rather than its consequences.
Three studies conducted at the same time as our own did attempt to measure how students were affected by test-score manipulation. Dee et al. (2019) found that inflating a score on New York Regents examinations led to a 22 percentage point increase in the probability of a student graduating, but reduced the likelihood that students took advanced courses. Diamond and Persson (2016) studied the manipulation of high-school math exams in Sweden. They found that inflated scores raised students’ achievement in future classes and increased earnings at age 23. Borcan et al. (2017) found that an effort to reduce cheating in Romania led to an increase in score gaps between poor and non-poor students. In turn, poor students were less likely to gain admission to an elite university.
This research is distinct from ours in that it all focused on high-school students. While the exams that we studied were used for state accountability and did have some role in determining eligibility for remediation services, the fact that we studied elementary and middle school students meant that the scores on the manipulated exams did not directly play a role in signaling ability to colleges or determining graduation. Instead, they were primarily known to the student, the students’ family, and his or her future educators. In addition, the New York and Sweden studies relied on teacher discretion in grading. Our context differed in that teachers were not intended to have any discretion over the grading of the exams, but nonetheless changed student responses to multiple-choice questions in order to inflate scores.
Mechanisms Through Which Cheating Could Impact Students
There are a number of mechanisms through which one may expect students to be impacted by test-score manipulation.
First, students may not have been directed to services they would have otherwise received. In addition to the possibility of summer school and retention, the district we studied offered an “early intervention program” (EIP) to students who did not meet proficiency benchmarks. The program is characterized by either “pull-out” small-group instruction or a reduction in class size. While the efficacy of this specific program has not been studied rigorously, other research suggested that small-group tutoring (Fryer, 2014) and reduced class sizes (Angrist and Lavy, 1999; Krueger and Whitmore, 2001) can have significant positive impacts on students.
Second, in an environment where cheating was prevalent, teacher effort may have been altered, reducing effective instruction. While early experimental studies found little evidence that teachers adjusted their effort based on incentives (Springer et al., 2012; Fryer, 2013), more recent analysis by Dee and Wyckoff (2015) of a district-wide scheme operated at scale (Washington DC’s IMPACT teacher accountability system) provided strong evidence that existing teachers will in fact adjust their teaching performance in response to incentives. If the culture of the district we studied rewarded the illusion of success rather than high-quality instruction, teachers may have reduced their effort in the classroom, knowing they could inflate student scores after the exams were administered. In addition to changes in effort, the cheating environment may have shifted the composition of the workforce away from a focus on staff quality to a focus on compliance with the cheating regime. 7 Given the importance of teachers as an education input (Chetty et al., 2014), any reduction in teacher quality would likely impact students in lasting ways.
Finally, students may have exhibited behavioral responses to the signals sent by the false scores. If a student is cheated and then later learns that their true performance is lower, it is possible this could damage their self-esteem. Prior studies have shown that self-esteem may have had positive effects on both educational attainment and labor market outcomes (Waddell, 2006; de Araujo and Lagos, 2013). Alternatively, false test scores may lead students to overestimate their ability and put less effort into their studies (Babcock, 2010).
Identification Strategy and Methodology
The task of determining which students were cheated was relatively straightforward in our context. We relied on the number of WTR erasures for students in flagged classrooms as an indicator of the presence and magnitude of test score manipulation. 8 Without plausible experimental variation, choosing an approach that allowed us to accurately measure the effect cheating had on students was more challenging.
Evaluation of Available Strategies
We began by considering what we know about selection from the state investigation. We knew that the degree of cheating varied a great deal from one school to the next and in some cases from one class to the next within the same school

Distribution of Students by Percentage of Schoolmates in Flagged Classes.
Once a staff member chose to cheat, we knew that they did not always cheat at random. Instead, some staff members identified students they knew to be weaker and changed their responses, leaving test scores in tact for students they were less concerned about. Figure 3 shows the distribution of WTR erasure within flagged and non-flagged classrooms. 9 Even in cases where the cheating was doled out evenly, weaker students would have been more likely to have answers changed from wrong to right. 10 After all, an answer must have been wrong in the first place before it could be corrected. Thus, variation in cheating within classrooms was likely correlated with true student ability and the teacher’s perception of each student’s ability.

Distribution of WTR Erasures by Flagged Status,
We also explored whether teachers were more likely to cheat students who prior to cheating were slightly below important thresholds used for state accountability. 11 Prior research on the consequences of high-stakes accountability has found that teachers may focus their efforts on students near a proficiency threshold used for accountability (Neal and Schanzenbach, 2010; Reback, 2008). Figure 4 presents the results of this analysis for Math, ELA, and Reading. While the figure illustrates the selection on ability noted above, there was no evidence that teachers corrected more answers for students just below an accountability threshold than they did for those just above. 12

Threshold Analysis of WTR Erasures
This complex selection process and the absence of a source of exogenous variation meant that unbiased measures of cheating’s effect relied to some extent on our capacity to control for student ability and other characteristics that influenced the selection process. We did have access to an important ability control from a low-stakes exam, the NRT discussed earlier. Therefore, we estimated the following equation.
We present three specifications that begin without fixed effects, then add school-grade fixed effects and finally classroom fixed effects to account for differences in instruction quality that may be correlated with the choice to cheat. Our models including the fixed effects do not measure impacts of the cheating scandal that worked through the mechanism of teacher quality. Instead, they isolated effects arising from the answer changing itself. The model without fixed effects provided some suggestive evidence keeping the teacher-quality mechanism open; however, we urge caution against over-interpreting the results as it is unclear to what extent the cheating regime caused changes in instructor quality as opposed to less able instructors choosing to cheat. 15
Measuring Potential Omitted Variable Bias
The identification strategies above effectively dealt with differences in student ability from one class to the next and differences in instructor quality that may be correlated with cheating. However, we were only able to address within-class selection of whom to cheat by controlling for observable characteristics of each student, including a measure of ability. Given that the NRT measured ability with error and that teachers likely knew much more about students’ cognitive and non-cognitive skills, we anticipated some degree of selection occurred that we were unable to control for.
Altonji et al. (2005), henceforth AET, presented a formal framework for bounding the potential omitted variable bias using assumptions about the relationship between observed and unobserved confounders. Oster (2019) extended this framework and found that in the years following the publication of AET, few empirical works in top economics journals employed the method. She suggested one reason for lack of adoption may be that researchers were unwilling to assume all of the variation in the outcome could ever be explained, even with perfect knowledge. She offered a method of relaxing this assumption by selecting maximum values of R-squared. 16 We found this a useful extension and adopted it in some of our tests.
While we were not willing to make the assumption that selection on unobservables was exactly equal to the magnitude of selection on all observables available to us, we knew that some amount of selection occurred using information about how students performed on the 2009 CRE prior to test manipulation. We, therefore, employed sensitivity tests that measured the magnitude of selection on unobservables which would be required to nullify our findings relative to the selection that occurred on the untainted ability measure we observed, the NRT. 17 This evidence played an important role in informing our conclusions.
Results
Results Assuming No Omitted Variable Bias
We begin our presentation of results with evidence of baseline equivalence, showing the “effects” of treatment on exogenous student characteristics cheating could not have impacted. Pei et al. (2019) argued that this form of analysis is particularly important when measurement error existed in a control. We knew that our ability control, NRT, was measured with error, and in addition to highlighting the magnitude of selection on observables, the analysis helped us understand how the degree of selection varied across different sources of variation in the data (across schools, across classes within schools, and within classes).
Table 2 shows how the number of WTR erasures varied with the NRT, free lunch eligibility, gender, and race. It was clear that selection occurred across each of our three model specifications. Importantly, the magnitude of selection on the NRT was approximately three times the magnitude within classes as it was before fixed effects were added. This is not particularly surprising given what we knew from the state investigation. Not all teachers participated in the cheating and some of the teachers who refused to cheat likely had weak students. However, once a teacher chose to cheat, the weakest students received substantially more WTR erasures. It is important to consider what this implies for selection on unobserved student characteristics. If selection on omitted variables was proportional to selection on observed variables, we expected that selection on unobserved characteristics was more meaningful within classes than across the full data.
WTR Erasures and Exogenous Characteristics.
Note: Robust standard errors in parentheses are clustered at the 2009 classroom level. * p<0.10, ** p<0.05, *** p<0.01.
Panel A of Table 3 shows the association between WTR erasures and future test scores when no control for ability was included. 18 The way to interpret these results was that each additional WTR erasure in a flagged classroom was associated with reduced future achievement of between 0.008 and 0.014 standard deviations. We saw that the magnitude of the measured association strengthened as we moved from the model without fixed effects to the model with classroom fixed effects. This was consistent with the greater degree of selection on ability within classrooms observed above.
The Effect of Cheating on Future Test Scores.
Note: Robust standard errors in parentheses are clustered at the 2009 classroom level. Individual controls include race, gender, free lunch status, and special education. All regressions are weighted by the inverse of the number of times a student is observed in the sample.
* p<0.10, ** p<0.05, *** p<0.01.
Panel B of the same table, shows similar models but with the NRT included as a control for ability. Across each of the models, the measured association fell substantially, but not all the way to zero. In the classroom fixed effects model, the magnitude of the association was one third what it was before the NRT control was added. In Figure 5, we present visual evidence of the attenuation without parameterizing the relationship between WTR erasures and future test scores. The binned scatterplots compare residualized WTR erasures and residualized future test scores after controlling for observable student characteristics and classroom fixed effects. The NRT was first omitted, then included. The degree of attenuation was very clear; however, we did not yet know whether the relationship would disappear entirely if we were able to control for all relevant information rather than only the incomplete measure of ability available to us.

Wrong-to-Right Erasures vs. Future Test Scores.
Returning to Table 3, if researchers did not understand that the magnitude of selection within classes was larger than the magnitude across the full data, they might conclude that coefficients attenuate when controls were first added, but then stabilized as fixed effects were introduced. This could lead to an interpretation that the results in Panel B suggested a meaningful causal effect. However, we reached a different conclusion.
Evaluation of Omitted Variable Bias
As explained in Section 4., we acknowledged that there was likely some degree of selection on omitted variables, and if we could have controlled for all relevant information, we would have observed attenuation of our results. However, without additional tests, it was not clear whether we should expect the existence of omitted variables to nullify our findings or simply reduce their magnitude. We therefore present tests for omitted variable bias first proposed by AET and later extended by Oster (2019). Following the terminology used by Oster, consider the regression model:
We followed an extension to this framework proposed by Oster (2019) which improves its applicability to our context. It would be somewhat challenging to develop intuition for omitted variable bias relative to all of the observed controls in our model. For example, it was tough to conceptualize unobserved selection relative to several hundred classroom fixed effects. However, we were much more comfortable evaluating selection arising from omitted variables relative to a specific observable: performance on the low-stakes assessment, the NRT. Therefore, when calculating the delta that would be required for
Even if we observed all of the relevant information above and could incorporate it into our regressions, there would likely remain unexplained variation in future test scores. Things like measurement error in the future tests, idiosyncratic events in the student’s life, and randomly being assigned better or worse teachers in the future would have caused variation in the future test scores unrelated to the selection process. In the regression equation above, AET assumed
Table 4 showed a value for delta across our different model specifications and different assumptions for the maximum value of R-squared. The results showed that in models holding school and teacher quality constant, if selection on omitted variables were approximately one half the magnitude of selection on the NRT, we would measure no causal effect of test-score manipulation. As expected, the smaller the distance between the explained variation and the maximum R-squared, the larger delta must be in a given model.
Test for Omitted Variable Bias.
Note: Individual controls include race, gender, free lunch status, and special education. All regressions are weighted by the inverse of the number of times a student is observed in the sample. The R-squared Uncontrolled represents the R-squared from a regression excluding the NRT as a control. R-squared Controlled represents the R-squared from a regression including the NRT as a control. The R-squared max represents the share of variation in the outcome that could be explained if all relevant information were observed.
It certainly seemed reasonable that unobserved selection of this magnitude may have existed, but it is somewhat difficult to say conclusively since our expectation of delta arising from the three omitted variables explored qualitatively above ranged from less than 1 to above 1, and we did not know what share of the variation in future tests each omitted variable might explain. We also did not know with certainty that an R-squared maximum of 0.900 is sufficiently low. It is possible that variation in future test scores unrelated to the selection process exceeded 10 percent.
Approaching this test in another way strengthened our conclusions. We abstracted from assumptions about selection on all unobservables, and instead considered selection on a single unobservable: performance on the 2009 CRE prior to cheating. From the state investigation and our knowledge of how cheating occurred, we were convinced that selection on the 2009 CRE was at least equal to selection on the NRT, potentially even stronger since it played a direct role in determining WTR erasures. So it would be helpful to test whether a delta of one results from an AET/Oster analysis if the only relevant unobservable were the 2009 CRE. To implement such a test, we needed to set the R-squared maximum to the level R-squared would reach if the 2009 CRE were added. The data allowed us to develop a reasonable expectation for what that level should be.
We relied on data for students who were not cheated to determine the appropriate increase in R-squared to expect. 21 For this subset of students, their 2009 test score was an accurate measure of their ability. Table 5 shows that when the 2009 CRT was added as a control to regressions of future test scores already including the NRT for non-cheated students, the R-squared rose substantially in each of the models. It reached approximately 0.73, 0.78, and 0.80 in the no fixed effects, school-grade fixed effects, and classroom fixed effects models, respectively. This confirmed that the 2009 CRE contained important information about student ability not captured by the NRT.
Evaluation of Explainable Variation for Non-Cheated Students.
Note: Individual controls include race, gender, free lunch status, and special education. All regressions are weighted by the inverse of the number of times a student is observed in the sample.
We returned to our AET/Oster analysis and set the maximum R-squared at these levels, reflecting the additional variation in future test scores we expected would be explained if the 2009 CRE prior to cheating were observed and could be used as a control. Table 6 shows that in both the school-grade and classroom fixed effects models, the resulting delta was approximately one. We therefore concluded that any effects we measured in models holding teacher and school quality constant were entirely explained by the fact that the 2009 CRE was directly used in determining WTR erasures, the fact that it contained important information about student ability not reflected in the NRT, and the fact that we were unable to include it as a control because it is tainted for some students. If all relevant information were observed, our models would measure no causal effect of changing student test responses on the students’ future achievement.
Test for Bias from Omitted 2009 CRE Only.
Note: Individual controls include race, gender, free lunch status, and special education. All regressions are weighted by the inverse of the number of times a student is observed in the sample. The R-squared Uncontrolled represents the R-squared from a regression excluding the NRT as a control. R-squared Controlled represents the R-squared from a regression including the NRT as a control. The R-squared max represents the share of variation in the outcome that could be explained if all relevant information were observed.
Our tests for omitted variable bias lent support to a causal interpretation in the specification without fixed effects. For example in Table 6, delta would have to be 3.373 in order for the treatment effect to reach zero. We interpreted this as evidence that students attending schools where cheating was prevalent fared worse in the future than what could be explained by their own characteristics and abilities. Recall that we were not holding school or teacher quality constant in this specification. It is unclear to what extent the cheating regime led schools to provide a worse education. Accounts from the state investigation suggested a focus on compliance took precedent over instruction quality; however, we were unable to disentangle these effects econometrically. To the extent that weaker teachers chose to cheat because they were not successful raising student achievement through more constructive means, we would have expected students to receive a worse education even if cheating had never occurred.
Summary and Conclusion
The rise of standardized tests as a management tool in elementary and secondary education has introduced an incentive for educators to falsify test scores. Several instances of cheating have been uncovered around the country in recent years, but very little prior evidence existed on whether students were harmed by test-score manipulation. Three studies conducted at the same time as our own have looked at how high school students were affected, where performance on the assessments played an important role in determining graduation and signaling ability for college admissions. Ours was the first study to analyze how elementary and middle school students were affected by widespread cheating in a predominately black school district. We found no evidence that test score manipulation itself hurt or benefited students, though we did find evidence that students attending schools engaging in cheating fared worse in the future, not holding school or teacher quality constant.
These findings suggested that district leaders looking to respond to cheating’s discovery should direct their energies toward general school improvement rather than targeting specific students whose answers were changed. To the extent that cheating had an effect, suggestive evidence pointed to the effect operating at the school level, perhaps as a focus on compliance with the cheating regime taking precedence over a focus on staff quality. Our results may also prove informative for those responsible for determining the consequences teachers face when cheating is discovered. Consequences are warranted for the role that cheating plays in making test scores less useful as management tools, and punishments should also consider whether school or district leadership built a culture that prioritized false signals of success over quality instruction. However, the evidence did not suggest that sentences be set based on a belief that answer-changing itself harmed students. In this predominately Black school district, Black teachers were disproportionately penalized where our evidence did not point towards harm of answer changing in and of itself.
In addition, our research showed the usefulness of tests for omitted variable bias even when rich administrative data were available. Much of the prior work employing such tests has been in settings where researchers used survey data. We demonstrated how researchers unwilling to make some of the AET/Oster assumptions can nonetheless gain important insight by making less arduous assumptions appropriate in a given context where researchers have knowledge about the likely sources of omitted variable bias. Rather than measuring the degree of selection on unobservables relative to all observed characteristics, researchers can measure such bias relative to selection on the subset of observables the researcher believes is most similar to the omitted variables.
Despite the fact that we did not find consequences of test score manipulation on students’ future outcomes, districts and states may still benefit from further investment in ensuring the accuracy of results. Annual exams are increasingly used in teacher compensation schemes, charter school renewal decisions, and the choice to close under-performing schools. As management tools, test scores have little value when there is not sufficient oversight to ensure their reliability. The prevalence of cheating across the country suggests oversight has not risen proportionally with the role tests play in supervision. To the extent that coordinated cheating schemes may reduce instructor quality, all students – not just those whose tests were changed – would benefit from its quick discovery and eradication.
Supplemental Material
sj-pdf-1-rbp-10.1177_00346446211068163 - Supplemental material for Teacher cheating: Are students affected by teachers who cheat? Evidence from a predominately Black district
Supplemental material, sj-pdf-1-rbp-10.1177_00346446211068163 for Teacher cheating: Are students affected by teachers who cheat? Evidence from a predominately Black district by Carycruz Bueno and Jarod Apperson in Review of Black Political Economy
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship and/or publication of this article.
Supplemental Materials
The online Supplemental Materials include some computational details, additional numerical results and technical proofs.
Notes
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
