Abstract
Despite the concerns regarding value-added models (VAMs), advocates hold strong to VAMs’ theoretical strengths and potentials, while adopting a set of agreed-upon albeit “heroic” set of assumptions, without research in support. These assumptions transcend promotional, policy, media, and research-based pieces, but they have never been made explicit as analyzed as a set or whole. The purpose of this study was to make unambiguous the assumptions that have been made within the VAM narrative that have often been accepted without challenge, and situate these assumptions within the research. Sources for analyses included 470 pieces, published within traditional and nontraditional outlets, from which we derived 27 prevailing assumptions.
Keywords
Introduction
Value-added models (VAMs) are intended to objectively measure the amount of “value” that a teacher “adds” to (or detracts from) student learning and achievement from one school year to the next. More specifically, VAMs are statistical tools meant to measure the observable relationships between a teacher’s instruction and how that instruction contributes to student learning and achievement over time. This is done by statistically measuring growth in achievement from one year to the next using students’ large-scale test scores, while controlling for students’ prior testing histories, and sometimes controlling for student-level demographics (e.g., race, poverty, English language proficiency, special education status), other student-level variables (e.g., attendance, suspension, retention), and other classroom- and school-level variables (e.g., class size, average prior achievement) as available and as specified per model.
Prior to the recent January 2016 passage of the Every Student Succeeds Act (ESSA; 2016), which helped to curb the stronger teacher and teacher education accountability reforms ongoing across the United States since President Obama’s Race to the Top (2011) initiative (by which states were incentivized/required to adopt VAM-based teacher evaluation policies), 44 states and D.C. (88%; n = 45/51) had adopted and had at least begun implementing VAMs to evaluate and, in many cases, make important decisions about teachers in effect (e.g., teacher tenure, merit pay, teacher termination). In addition, 30 states and D.C. (61%; n = 31/51) had state legislative acts and regulations that mandated such teacher evaluation initiatives (Collins & Amrein-Beardsley, 2014). Given the recency with which ESSA was passed, however, it is now somewhat uncertain what states are doing post ESSA. Although some states seem to still be moving forward with their teacher evaluation and accountability plans regardless of ESSA (e.g., Florida, New Mexico, New York, Ohio, Tennessee, Texas), others have at least paused or begun to reframe their stronger accountability efforts (e.g., Alabama, Arizona, Georgia, Oklahoma), although these states still seem to have some measure of value-added included within what still appear to be similar teacher evaluation systems, albeit with less consequential stakes attached.
Perhaps the fact that states appear to still be moving forward in these regards, albeit in varying ways, is because the idea behind the development and implementation of VAMs is not necessarily wrongheaded; hence, and especially conceptually, these reform ideas continue to resonate among many state-level educational policymakers and educators. However, it can also be argued that educational policies in support of such strategies are nonetheless misguided given the research surrounding VAMs, including VAM adoption, implementation, and use, and especially when said uses are tied to high-stakes consequences (e.g., teacher tenure, merit pay, teacher termination). Specifically, contemporary concerns about using VAMs for teacher evaluation and accountability purposes, as per the current research, primarily pertain to (a) methodological concerns (e.g., reliability, validity, bias), (b) logistical concerns (fairness, transparency, usability), and (c) consequential concerns (e.g., high-stakes attached to outcomes, effects on school culture and the teaching profession). The purpose of this article is to explore these concerns, while attempting to understand the persisting debate that surrounds them.
There is a growing divide between academic scholars who have taken to either side of the VAM debate, with many educational economists often (although not always; see, for example, Rothstein, 2009, 2010, 2014) on one side, promoting the strengths and potentials of these models (see, for example, Chetty, Friedman, & Rockoff, 2014a, 2014b; Hanushek, 1970, 1971, 1979, 2009, 2011; Kane & Staiger, 2008, 2012), and educational researchers often on the other side, more often representing VAM critics (see, for example, B. D. Baker, Oluwole, & Green, 2013; E. L. Baker et al., 2010; Briggs & Domingue, 2011; Corcoran, 2010; Eckert & Dabrowski, 2010; Gabriel & Lester, 2013; Graue, Delaney, & Karch, 2013; Hill, Kapitula, & Umlan, 2011; Newton, Darling-Hammond, Haertel, & Thomas, 2010; Papay, 2011; Rothstein, 2009, 2010, 2014). Given the still current momentum of VAM adoption and use, then, it is reasonable to posit that economists have more or less taken the lead in influencing now state-level policies in this area, as well as international initiatives and policies in this area (Araujo, Carneiro, Cruz-Aguayo, & Schady, 2016; Sørensen, 2016), which is not surprising in light of their rising influence in the public policy arena at large (Fourcade, Ollion, & Algan, 2015; Lazear, 1999, 2001)
The Rise of the Economist in Education Policy
Of all the social science disciplines, economics has risen to the top in terms of its political and social influence (Lazear, 1999, 2001). Fourcade et al. (2015) describe this rise in clout, arguing that despite economists’ inability to predict the financial crisis of 2007-2009, their analytic tools to make predictions about social processes are trusted more now than they were 30 years ago. Fourcade et al. also point out that economists see their discipline as the most rigorous, and accordingly rely almost exclusively on what they view as the most rigorous statistical models and measures to conduct their research, to often yield economics-based (e.g., “bottom-line”) answers. Related, the epistemological underpinnings of the economics discipline, that is about what is known and can be known using such methods and models, often contradict those of other social science disciplines. This creates an increasingly insular network of economists that rarely incorporate, and often marginalize other interdisciplinary traditions (e.g., sociology, political science, psychology, education). Lazear (1999) further explained,
It is the ability to abstract that allows us [economists] to answer questions about a complicated world . . . I have argued elsewhere that the strength of economic theory is that it is rigorous and analytic. But the weakness of economics is that to be rigorous, simplifying assumptions must be made that constrain the analysis and narrow the focus of the researcher. It is for this reason that the broader thinking sociologist, anthropologist and perhaps psychologist may be better at identifying issues, but worse at providing answers. Our narrowness allows us to provide concrete solutions, but sometimes prevents us from thinking about the larger features of the problem. This specialization is not a flaw [however, as] much can be learned from other social scientists who observe phenomena that we often overlook. (pp. 5-6)
Given that economists largely rely on parsimonious frameworks, it is in the nature of the economist’s work to make sense of complex issues using sophisticated tools, in or to simplify the issue. However, from an epistemological standpoint, statistical tools that intentionally disregard uncontrollable factors cannot always help social scientists get at the complex social processes that they are tasked to investigate. Thus, while econometrics might appeal to policymakers for the tangible, “simple” outcomes and solutions they can yield, other analytical tools might better serve our understandings of complex social domains, such as those resident within education.
Related, factions of educational policy have come to be based on such economic models and measures, especially given the policy-popular accountability initiatives related to teacher quality (e.g., VAMs). For example, after the sociologist, Coleman and colleagues (1966) reported that teacher quality was the most influential in-school variable when explaining differences in student achievement scores, the economist, Eric Hanushek, began to conceptualize teacher effects in economic terms as based on the relationship between inputs (e.g., education status, years of experience) and outputs (i.e., student achievement scores). After applying an economic-based analytical framework to a set of data in one California school district, Hanushek (1970) argued that (a) years of experience and graduate education were not related to higher student achievement, (b) teacher effects did not explain Mexican American students’ achievement outcomes, and (c) teacher effects did have an impact on White students’ achievement outcomes, regardless of socioeconomic status. Hanushek (1979), encouraged by these findings, proceeded to call upon additional econometric instruments to develop a model that could measure the amount of value that a teacher added or detracted from student learning. Thus, the VAM—which was most commonly used in business and agriculture prior—made its way into the education scene. Soon thereafter, a different set of econometricians further established the VAM in more applied and pragmatic terms, for example, in the state of Tennessee (Sanders, 2003; Sanders, Wright, Rivers, & Leandro, 2009).
VAMs have since taken a progressively significant role in teacher accountability policies and practices, regardless of the extent to which the methods might be thwarted by theoretical, methodological, and practical problems, some of which we mentioned prior (see, for example, B. D. Baker et al., 2013; E. L. Baker et al., 2010; Briggs & Domingue, 2011; Corcoran, 2010; Eckert & Dabrowski, 2010; Gabriel & Lester, 2013; Graue et al., 2013; Hill et al., 2011; Newton et al., 2010; Papay, 2011; Rothstein, 2009, 2010, 2014; Schochet & Chiang, 2012). Accordingly, VAMs continue to be the source of concern surrounding discussions of educational reform. This tension might also be due, at least in part, to the theoretical and epistemological differences of the economists who have had the most influence on education policy/practice, and other social scientists who have criticized VAM-use on the grounds of simplicity.
In this conceptual piece, we argue this might just be the case, as we attempt to disentangle these discrepancies by reviewing a large set of the VAM literature coming from multiple disciplines, as well as nontraditional sources, such as op-ed pieces, blogs, and newspaper articles. We do this while tracking and coding the implicit and explicit assumptions that have, over time, been accepted as “truth” despite interdisciplinary disagreement and interdisciplinary research evidence countering the assumptions being made. Ultimately, we argue that a multidisciplinary approach should be used to understand VAMs and VAM-use.
Conceptual Framework
As mentioned, there is a group of economists subscribing to a particular ideological camp who accept and therefore work from a set of theoretical and epistemological assumptions that underpin VAMs and VAM-use. These particular economists work from a discipline-specific assumption that teachers are rational beings who make rational decisions. Logically, VAM methodology is seen as a legitimate means for measuring teacher effects and using VAM outcomes to incentivize teachers to perform better. This is in stark contrast to many other social scientists (including some economists) who conceptualize education and teaching as complex social domains that cannot possibly be understood in such rational and simplistic ways. This former set of economists who, again, subscribe to this particular ideological camp, tend to work from the presupposition that “the properties of ‘collectivities’—groups, institutions, societies—can be reduced to statements about the properties of individuals” (Ingham, 1996, pp. 245-246).
In contrast, for example, sociologists tend to work from the presupposition that individuals are part of larger social structures that influence the makeup of the individuals (Ingham, 1996). Accordingly, the methodological approaches to understanding human behavior and society are quite different. Other social science disciplines have entirely different sets of epistemological assumptions within which they ground their work. While the field of economics is unquestionably made up of far more than a monolithic group of scholars and analytical approaches, there is a particular faction of the discipline that has embraced VAMs by anchoring their arguments in assumptions that fail to hold up in other disciplines’ approaches.
Going back to the previous example, the sociologist who typically sees individuals as parts of larger social structures might seek to understand teacher quality as a complex mixture of social factors that cannot, entirely, be statistically observed or measured, or much less statistically controlled. Sociologist Raudenbush (2004) suggests something similar in his statement: “[T]he estimates from VAM, when combined with other information, have potential to stimulate useful discussion about how to improve practice. But they should not be taken as direct evidence of the effects of instructional practice” (p. 128). Hence, the differences in epistemological positions create challenges when scholars representing multiple disciplines attempt to understand a single phenomenon or subject, such as teacher quality or how teacher accountability mechanisms can improve it. Likewise, this calls for an multidisciplinary approach to understanding VAMs and VAM-use.
To this end, our purpose for writing this article was to address the conflict that specifically relates to VAMs and VAM-based policies and practices. We sought to understand the core of the debate by investigating two primary research questions: (a) to identify the assumptions, or conditions, that have not been made explicit when calculating teachers’ value-added, but have been implicitly assumed as “true” conditions upon which VAM-based output can be used; and (b) to map these assumptions onto the greater VAM literature to interpret the feasibility of such assumptions as being “true” given a multidisciplinary lens.
For the purposes of this article, we define “assumption” broadly to include room for all social science disciplines’ definitions, for we find that such inclusion is not only important but also critical for understanding the complexities involved in the process for measuring teacher value-added. Specifically, we define an “assumption” to mean a condition which has been accepted as “truth,” with the understanding that violations of said assumptions can ultimately result in invalid inferences, spurious conclusions, and sometimes unintended consequences at some level of VAM development, implementation, or use. As such, assumptions can include traditional, statistical-based assumptions that must be met for VAM estimates to be considered reliable and valid, or assumptions can include more publicly accepted beliefs about, for example, teachers’ attributional or causal effects on student achievement over time (i.e., value-added).
Consequently, and regardless of the validity or truth of the assumptions made, or whether the assumptions made are done so in implicit or explicit terms, such assumptions are inextricably linked to the rationales behind VAMs, as well as the justifications used to perpetuate the VAM narrative. Accordingly, once assumptions are made explicit, violations of assumptions regardless of assumption type must be acknowledged and then questioned as to “whether plausible departures from those assumptions would lead to substantially distorted inferences” (Reardon & Raudenbush, 2009, p. 493).
As educational policy researchers, what sparked our interest in this study was the ongoing acceptance of VAMs and VAM-based policies at high levels of policy decision-making (e.g., federal and state governments), despite mounting concerns coming from educational scholars, but also given the mounting influence of many economists in the area of VAMs. As we argued prior, econometric methods are increasingly being regarded above all other disciplinary methods of inquiry (see also Fourcade et al., 2015), as are economists themselves as the sage protectors of the public good; hence, we argue that making explicit and examining these assumptions is a necessary precondition to better understanding VAMs and VAM-use, with implications for educational policy.
Specifically, we were interested in looking at the assumptions surrounding all sides of the debate in an effort to understand the epistemological conditions that might be obscuring the possibility for reaching similar “truths” regarding VAMs. As such, we conducted a systematic review of the VAM literature, including both traditional (i.e., peer-reviewed academic articles) and nontraditional (e.g., newspaper articles, blog entries) sources, while focusing on the various implicit and explicit assumptions that contributed to the sources’ conditions, methods, findings, and conclusions.
Research Methods
As duly noted, we worked from the position that a group of economists have led the work, as well as the narrative, on VAMs (Chetty et al., 2014a, 2014b), especially as it relates to the asserted need for VAMs (Weisberg, Sexton, Mulhern, & Keeling, 2009), the development of VAMs (Hanushek, 1970, 1971, 1979, 2009, 2011), and the use of VAMs for teacher evaluation purposes (Harris, 2011; Kane & Staiger, 2008, 2012; Sanders, 2003; Sanders et al., 2009). Accordingly, VAMs and VAM-based policies and practices have been built upon a set of assumptions that are appropriate to the discipline of economics, such as the notion that teacher effects can be measured in isolation of or controlling for outside factors, and that teachers can (or will) make better instructional decisions given the right conditions and incentives (e.g., the promise of merit pay).
However, given the complexities that are inherent in educational systems, it might serve our understandings to apply multidisciplinary approaches to think about not only the capabilities of VAMs but also the assertions regarding the need for VAMs, the practical application of VAMs, and the potential consequences related to VAM-use. We argue that by looking at the VAM literature in terms of these disciplinary assumptions, we might better understand the nature of the debate, while unpacking the problems that can arise by depending almost entirely on one disciplinary approach to any problem.
To this end, we conducted a systematic review of the VAM literature, covering a total of 470 unique sources, both in the traditional and nontraditional academic senses. Our goal was to locate and identify the implicit and explicit assumptions that were most commonly made across pieces, while situating these assumptions within a multidisciplinary framework. In other words, and as mentioned prior, we did not limit our definition of “assumption” to the strict statistical sense of the word. Rather, we took “assumption” as more broadly conceived, to encompass that which is accepted as “truth,” devoid of empirical proof. For example, while some economists might assume that factors affecting student learning can be statistically controlled for, other social scientists might not work from the assumption that this can truly be achieved. These epistemological differences can and should be unpacked, which is what we attempted to do.
Data Collection
Sources for this study included 470 distinctly different pieces that we collected via multiple sources. These sources included traditional journal outlets, for example, via journals’ official publication announcements and emails, multiple series of Education Resources Information Center (ERIC) searches (e.g., using search terms such as value-added, value-added models, VAMs, teacher growth, teacher accountability, teacher evaluation), and via a snowball sample approach, whereby the references in one article also lead us to others. As for our nontraditional sources (e.g., newspaper articles, blog entries), we used multiple Google Alerts (e.g., using the same search terms used for the ERIC searches) to collect articles on a daily basis. Although we also employed a similar snowball approach here if, for example, a blog post leads us to a news story or journal article.
Employing these systematic albeit imperfect data collection methods on both the traditional and nontraditional ends, we attempted to safeguard against our own biases that may have otherwise subconsciously informed our selections, and also ensure that we included counterfactuals as also key. While we do not have evidence that the proportions of articles that we pulled in favor or in opposition to VAMs indeed match the true proportions resident within the entire population of traditional and nontraditional pieces out there on this topic, given this would likely be impossible to gauge, via the data collection methods that we used we believe that our article sample was appropriate, rich, and representative of the general arguments surrounding VAMs and VAM-use.
It must be mentioned again, however, that in no way is this list complete or comprehensive given the controversial nature of this policy trend and the rapidity with which pieces were being published at the time of this study, in some cases multiple times per day. To read everything that has ever been written on this topic would be nearly if not entirely impossible. Regardless, we made extensive efforts to stay as current, comprehensive, and unbiased and objective as possible, all the while remaining open to both traditional and nontraditional sources.
In order by volume, resources that we read and analyzed included peer-reviewed research studies (29%, n = 138); articles published in media outlets (21%, n = 98); organization, foundation, business, and think tank research studies (20%, n = 95); editor- and self- or author-reviewed studies (14%, n = 66); federal, local, and other promotional materials (10%, n = 46); and blog posts (6%, n = 27) (see Figure 1).

Types of VAM-based articles read and analyzed for this study.
More specifically, we read and analyzed peer-reviewed research sources (29%, n = 138) including most prominently articles published in American Educational Research Association (AERA) journals (e.g., Journal of Educational and Behavioral Statistics), open-access journals (e.g., Education Policy Analysis Archives), economics journals (e.g., American Economic Review), National Council on Measurement in Education (NCME) journals (e.g., Journal of Educational Measurement), and other peer-reviewed academic journals (e.g., Education Finance and Policy).
Next, we read and analyzed articles published in media outlets (21%, n = 98) including national news outlets (e.g., the New York Times), professional press outlets (e.g., Education Week), and state and other local news outlets (e.g., the Chicago Tribune). In addition, we read and analyzed non–peer-reviewed research articles and technical reports (20%, n = 95) published by research organizations (e.g., the National Bureau of Economic Research [NBER]); foundations (e.g., the Bill and Melinda Gates Foundation); businesses, companies, and research firms (e.g., Mathematica Policy Research); and think tanks (e.g., Education Sector).
Finally, and least in terms of volume, we reviewed articles published in editorially and author- or self-reviewed outlets (14%, n = 66) including journals (e.g., Phi Delta Kappan), books and book chapters (e.g., Doug Harris’s, 2011, Value-Added Measures in Education), doctoral dissertations, and papers presented at national and other relevant conferences (e.g., the National Conference on Value-Added Modeling). We reviewed promotional and other advocacy materials (10%, n = 46) produced and released by federal (e.g., the U.S. Department of Education), local (e.g., Tennessee Office of Education Accountability), and some VAM-relevant businesses, companies, and research firms (e.g., SAS Institute Inc.). We read and reviewed blog posts (6%, n = 27) as well (e.g., the Huffington Post).
Data Analyses
To gain an understanding of how the VAM narrative and VAM-based practices have been founded on a set of conditions that may or may not be feasible given a multidisciplinary lens, we carefully read through each of the aforementioned sources, noting when the authors made or advanced assumptions, or implicitly accepted as “truth” potentially necessary preconditions, without acknowledging research or references in support. In other words, we attempted to make sense of the way in which VAMs have been politically and socially accepted, despite the academic contention, by tracking the narrative at the base of the discrepancies—the assumptions or conditions upon which VAMs are possible, necessary, and successful at measuring teacher quality as conceptualized.
As we read, we marked each explicit or implicit assumption and ultimately found and coded 1226 instances. Then we collapsed the assumptions into grander categories, using a set of bins (e.g., assumptions about the tests used to calculated VAM estimates). We then reduced and refined the assumptions to better explain the assumptions resident within each bin for conclusion-drawing purposes (Miles & Huberman, 1994; see also Leech & Onwuegbuzie, 2008). Finally, we mapped these assumptions onto the greater VAM literature to determine the feasibility, practicality, and appropriateness of VAM-use given a multidisciplinary lens. Simply put, we compared each assumption against the general research to see whether the assumption, or condition, held when situated within other disciplinary approaches.
This approach, while not a traditional type of a meta-analysis (Glass, 1976), is currently the most comprehensive review of the traditional and nontraditional sources surrounding VAMs conducted thus far. It is also the only of its kind to disentangle the assumptions present in both the traditional and nontraditional artifacts surrounding the VAM-based narrative as compared against the general VAM research consensus (see also Reardon & Raudenbush, 2009; Rubin, Stuart, & Zanutto, 2004; Scherrer, 2011).
This approach yielded a total of 27 assumptions related to, and positioned within the (a) assumptions used as rationales to justify VAM adoption, (b) major statistical and methodological assumptions about VAMs, and, related, (c) assumptions specifically made about the large-scale standardized tests used for value-added calculations. We discuss and disentangle these assumptions as per these four grander categories next. Again, we present each assumption as situated within the greater, multidisciplinary VAM literature so as to demonstrate the need for considering VAMs and VAM-use as part of complex schooling systems that can best be understood via multiple approaches instead of a single, and primarily economics-based disciplinary approach that dominates the current VAM narrative and policy context.
Findings
In this section, we identify each assumption (in italics) and then provide the multidisciplinary-based consensus regarding the assumption to demonstrate not only the complexities related to the assumption but also the implications for either systematically disregarding the assumption or accepting the assumption as “truth.” To be clear, we treated the 470 pieces as qualitative artifacts and, using a grounded approach, allowed the assumptions and multidisciplinary research consensuses to emerge accordingly, after which we investigated each one within the external research. Again, we are using “assumption” in the broadest sense of the word to include that which has been accepted as truth regardless of empirical support. The following assumptions have been accepted as truth, as evidenced in policy discussions and policy initiatives, despite the empirically supported concerns coming from disciplines outside of economics.
Assumptions Undergirding the Rationales of VAM Development, Adoption, and Use
The first set of findings represents seven common assumptions related to the rationalization of VAM development, adoption, and use. The justification for VAM-based practices and policies rests on a set of assumptions related to teachers, teacher quality, and the relationship between teacher effects and students. While these assumptions may sound legitimate on the surface, and have been presented as such in the media and policy narrative, various scholars have studied these assumptions explicitly and found fallacies in their core premises.
Teachers are the most important factors that impact student learning and achievement. Teachers are strong school-level factors that influence student learning and achievement, but teachers operate alongside many other school-level factors that also impact student learning and achievement, and these other factors largely conflate determinations about the authentic effects of teachers. The best and most recent estimates put teachers as responsible for between 1% and 14% of all effects on student achievement, as measured by large-scale standardized test scores, whereas the other 86% to 99% of effects on student achievement are caused by factors beyond the teacher (see, for example, American Statistical Association [ASA], 2014; Berliner, 2013, 2014; Good, 2014). This is also likely why for years, and across nations, we have observed that investing in education has not necessarily resulted in increased achievement scores or performance outcomes, not to mention decreased achievement gaps, fewer educational inequities, and the like (see, for example, Checchi, 2006; see also Coleman, 1966).
Good teaching comes from enduring qualities that teachers possess and carry with them from one year to the next, regardless of context. The teacher effect (i.e., 1%-14% of the variance in test scores) is not strong enough to supersede the powers of the aforementioned student-level and out-of-school influences and effects (i.e., 86%-99% of the variance in test scores) from one year to the next. Different students, with different backgrounds, in different years, and in different contexts require different instructional approaches. Success with one group of students does not guarantee success with an entirely different group of students (see, for example, Berliner, 2013; Newton et al., 2010).
Too many of America’s public school teachers are ineffective, unqualified, unskilled, lazy, or uninspired, and they are the ones hindering educational progress. This assumption is often used to justify the need for personnel decisions to be based on VAM as well as general teacher evaluation outputs (e.g., more objective observational system). However, various researchers have found this reasoning to be flawed. For one, not nearly as many teachers as is often assumed are in fact ineffectual, particularly when “teacher effectiveness” is quantified in normative terms (i.e., below average effectiveness via comparison with the mean). Thus, by statistical design, there will always be some teachers who appear relatively less effective simply because they all on the wrong side of the bell curve (more on this forthcoming). Otherwise, we really have no idea how many teachers are (or are not) effective because how one subjectively or objectively defines teacher effectiveness varies widely, and often varies by discipline and approach as well (see also Graue et al., 2013). While we all might have prior experiences with whom we considered to be ineffectual teachers, and this might be the reason we see simple statements like this in a lot of the popular literature surrounding VAMs, teacher reform generally, we cannot really support this statement or its opposite—that all U.S. teachers are great.
VAMs will improve upon, if not solve, that which is wrong with the faulty teacher evaluation systems traditionally used. There is no strong research evidence, from any discipline (including economics) to support the accurate identification of teachers using VAMs. Also, using VAMs in isolation or as the dominant indicator in evaluative systems based on multiple measures is not yet warranted, and perhaps should not be encouraged in that VAM-based evidence often correlates lowly or contradicts other indicators derived via multiple measures (i.e., a serious validity issue; see, for example, Braun, 2005; Eckert & Dabrowski, 2010; Hill et al., 2011; Newton et al., 2010; Papay, 2011; Rothstein, 2009, 2010; Schochet & Chiang, 2012).
To reform America’s public schools, we must treat educational systems as we would corporations. Educational corporations produce knowledge, the quality of which can be controlled by objectively measuring knowledge “products” via VAMs. Schools are social institutions that do not operate like mechanistic corporations. Teaching is one of the most complex occupations in existence, and thinking about teacher effects in economic terms is misleading, unsophisticated, and obtuse (see, for example, Braun, 2008; Harris, 2011; Linn, 2008). can be used for merit pay, which will motivate teachers to perform better. This, again, relies on the assumption that teachers are rational and they make rational decisions when provided the “right” information and the right dis/incentives. But, according to the extant research, there is no evidence to support this assumption. Teachers are not altogether motivated to improve student achievement by the potential to be monetarily rewarded for gains in student test scores (see, for example, Goldhaber, 2009; Gratz, 2010; Johnson, 1984). Also, pay-for-performance bonuses are not typically large enough to motivate much of anything, especially when considering the cuts in pay and halted levels of salary growth that teachers have often simultaneously realized. Furthermore, maximum levels of teacher effort have likely been reached more often than assumed, which also inhibits teachers’ capacities to continuously demonstrate VAM-based growth over time (see, for example, Yuan et al., 2013). What may be worse is that implementing pay-for-performance plans may trigger adverse consequences that well outweigh the intended, illusive, benefits of the pay-for-performance plans implemented (see, for example, Gratz, 2010; Johnson, 2015; Layard & Dunn, 2009).
Teachers can both understand VAM-based information and use VAM-based information to diagnose needs and improve upon their instruction, after which increased levels of student achievement and growth will result. There is no evidence to suggest that the research about using general assessment information generalizes when using VAM-based information in the same ways. There is no evidence that demonstrates that if teachers are provided increased access to VAM-based information, this will enhance teachers’ abilities to understand or use this information in instructionally meaningful or relevant ways. There is also no evidence to suggest that doing so will increase levels of student achievement or growth, either (see, for example, Gratz, 2010; Layard & Dunn, 2009). See also Dragoset et al. (2015) which also directly supports this counterassertion, importantly as recently released by the U.S. Department of Education.
Major Statistical and Methodological Assumptions About VAMs
The next set of findings represents the 10 assumptions often made about the methodological capabilities of VAMs. These include the major statistical and methodological assumptions about VAMs, after which we present the assumptions specifically made about the large-scale standardized tests most often used for value-added calculations.
Once students’ prior learning and other background variables are statistically controlled, or factored out, the only thing left in terms of the gains students make on large-scale standardized achievement tests over time can be directly attributed to the teacher’s causal effects. The variables typically available in the databases used to conduct VAM analyses (e.g., binary variables that capture a small although measureable set of student-level demographics) are very limited and, accordingly, cannot be used to conduct the statistical work nearly as well as often assumed. Even the richest of quantified variables, however (e.g., students’ learning opportunities at home or during the summer months) may not liberate value-added analyses from the confounding effects that students’ backgrounds and other risk factors have on achievement growth over time (see, for example, Ishii & Rivkin, 2009; Scherrer, 2011).
The placement of students into teachers’ classrooms occurs more or less at random, so really any effects that might be observed among teachers might be reasonably attributed to students’ teachers provided advanced statistical tools and techniques are used. The assignment of students to classrooms (and teachers to classrooms, as well as students and teachers to schools) is much farther from random than is often assumed, and this biases VAM estimates and weakens arguments that teachers, not students, were responsible for the effects observed (see, for example, Bausell, 2013; Braun, 2005; Paufler & Amrein-Beardsley, 2014; Rothstein, 2009, 2010, 2014).
The issues and errors caused by the nonrandom placements of students cancel each other out or can be canceled out using complex statistics to make nonrandom effects tolerable, if not ignorable. Even the most complicated statistics cannot, and may not ever be able to effectively counter for the deleterious effects caused by such nonrandom assignment practices (see, for example, Newton et al., 2010; Reardon & Raudenbush, 2009; Rothstein, 2009, 2010, 2014; Scherrer, 2011).
Bright students learn no faster than their less intellectually abled classmates, and students with lower aptitudes for learning learn as fast as their peers. Learning is neither linear nor consistent. Learning is generally uneven, unstable, and discontinuous, as well as often detected in punctuated spurts. The learning trajectories of group of students over time are not essentially the same, and they often deviate from the linear, constant, unwavering forms statistically predicted. For this reason, economic models that make predictions about expected student growth are not appropriate for measuring learning and/or teacher effectiveness in that students more likely grow in stints predicted by punctuated equilibrium.
VAMs use large-scale standardized achievement test scores that capture student achievement while students are under the direct tutelage of the same teacher, within the same academic year, from the pre- to posttest occasions being fall-to-spring. In almost every case, the tests used to estimate value-added are administered annually, from spring of year X to the spring of year Y, always if not almost always covering the summer months. Summer losses and gains, correspondingly, account for between one half and one third of the achievement gap that is persistently prevalent between high- and low-income students (see, for example, Harris, 2011). To assume that this growth is also linear over such extended periods of time is also false.
The influence that students’ prior teachers have on student learning (i.e., teacher persistence or residual effects) is negligible, and if not negligible, they can be statistically controlled for under the assumption that these effects decay quickly and, again, at the same rate over time. Because the pretest scores used to calculate teacher Y’s value-added are collected in year X, and the value measured from point X to Y includes teacher X’s prior effects, as also noted prior, it is highly unlikely that statements can be made that the value teacher Y added or detracted to his or her students was entirely due to teacher Y’s efforts (see, for example, Papay, 2011; see also B. D. Baker et al., 2013; Briggs & Weeks, 2009).
Even though multiple teachers teach and interact with students in multiple ways, value-added effects can be attributed solely to single teachers under examination. Teachers do not work in isolation of other teachers, and student test scores cannot be attributable to individual teachers in isolation of others’ efforts and effects (see, for example, E. L. Baker et al., 2010; Bausell, 2013; Harris, 2011; Scherrer, 2011). Parsimonious models that attempt to isolate the effects of one teacher (e.g., by permitting teachers to estimate what percentage of the year he or she spent with each student to be held accountable for that fraction or percentage) are not capable of capturing and controlling the effects of others’ contributions.
Peer effects are negligible. The extent to which students academically or socially influence other students in their classrooms or attending their schools is trivial. Students interact inside and outside of their individual classrooms in academic and social ways, both positively and negatively. As well, they likely boost or take away from one another’s achievement, respectively (see, for example, E. L. Baker et al., 2010; Braun, 2005; Checchi, 2006; Corcoran, 2010; Ishii & Rivkin, 2009; Lazear, 1999, 2001; Linn, 2008; Newton et al., 2010; Reardon & Raudenbush, 2009; Rothstein, 2009).
VAM analyses and estimates are based on sufficient amounts of data, and any data that are missing are missing at random. Data are very frequently missing, data are missing more often for high-needs students, and this is further complicated when more years of data are needed to conduct value-added analyses. Missing data are not missing at random, which is an assumption with its own title (i.e., missing at random [MAR]). Likewise, treating missing data as random further biases VAM-based estimates (see, for example, Bausell, 2013; Braun, 2005; Briggs & Domingue, 2011; Harris, 2011; Linn & Haug, 2002; Raudenbush, 2004).
VAM accuracy is defensible across teachers’ classrooms regardless of the numbers of students taught any given year. Even if no data are missing, the accuracy of all VAM estimates still depends on the size of the samples used to generate teachers’ value-added estimates. Estimates for teachers with fewer students (e.g., n = 6, n = 10) are less accurate than their colleagues to whom they might be compared (see, for example, Sanders, 2003; Sanders et al., 2009; see also E. L. Baker et al., 2010; Goldhaber & Hansen, 2010).
Assumptions Related to the Large-Scale Standardized Tests Used for VAMs
For VAMs to be considered appropriate, the tests upon which VAM-based estimates are calculated must also meet the conditions necessary for reliability and validity. VAMs, as currently conceived and used, are correspondingly built on conditions that the large-scale standardized tests used are appropriate measures of student achievement. The following set of 10 assumptions is specifically related to such tests.
Because test-based measurements typically yield a mathematical and assuredly highly scientific value, an appropriate level of certainty or exactness comes along with the numerical scores that result. However, VAM-based estimates are fraught with both random and systematic errors, and both types of errors pull the “true” values away from their individual “truths.” Random errors are often associated with measurement noise, for example, when idiosyncratic factors (e.g., a student’s mood) distort test scores and falsely inflate or deflate scores away from the “true” value that is to be observed and captured (Linn & Haug, 2002; Miller & Modigliani, 1961). Systematic errors are often equated with measurement bias or construct-irrelevant variance (CIV) when outside factors consistently, albeit falsely, inflate or deflate the measurement of a variable and therefore distort its interpretation and validity (see, for example, Haladyna & Downing, 2004). All of this, of course, results in extremely large or wide confidence intervals, inversely illustrating largely decreased confidence with really all teacher-level VAM estimates.
Large-scale standardized achievement tests serve as precise measures of what students know and are able to do. Tests offer very narrow measures of what students have achieved, and they do not often effectively assess students’ depth of knowledge and understanding and their ability to think critically, analytically, or creatively; solve contextual problems; or even accomplish authentic, performance-based tasks (see, for example, E. L. Baker et al., 2010; Corcoran, 2010; Harris, 2011; Toch & Rothman, 2008). However, one must assume that tests are capable of this before they can aggregate students’ test scores up to the teacher value-added level.
Regardless of whether student-level consequences are attached to tests, students inherently want to perform well on large-scale standardized achievement tests. Students often believe that tests are of no consequence. Inversely, when serious consequences are attached to tests, students might perceive the tests to matter, possibly more than they really do. Either scenario typically yields underestimates or overestimates, respectively, pulling test scores away from the “true” scores of interest.
Large-scale standardized achievement tests yield objective numbers based on complex statistics and therefore yield “true” scores about student learning. Statistics are being used to convey authority and intimidate others into accepting VAMs simply because VAMs are based on statistics and their mystique, despite statistics’ gross limitations when used to make presumably highly scientific inferences derived via tests (see, for example, Gabriel & Lester, 2013). Just because they are based on sophisticated mathematics and statistics, however, does not mean they are accurate.
The subject areas that are tested using large-scale standardized achievement tests matter more than the subject areas not tested. Sometimes science and, more often, social studies, music, physical education, foreign languages, and the arts are marginalized by the fact that it is very difficult to construct tests to assess student learning in these subject areas. Hence, what is tested often dictates what is taught, implying that the tested subject areas matter more than the others and implying that teachers of the tested subject areas matter more than the others and they should be, naturally but not de facto, held more accountable (see, for example, Corcoran, 2010; Toch & Rothman, 2008).
The items used to construct large-scale standardized achievement tests are included to measure things that students should know and should be able to do. However, the test items most often included on tests are those that “discriminate” well, and help test developers fit test scores into normal curves (i.e., with increased variance). Likewise, they are not representative of the test items we might value most or that are most often taught by teachers as written into state standards or as valued in the real world (see, for example, Braun, 2005; Corcoran, 2010).
The large-scale standardized achievement tests being used to measure growth are set on equal interval scales. This would mean that a 10-point difference between a score of 50 and 60 is the same thing as a 10-point difference between a score of 80 and 90. The tests being used are typically not set on equal interval scales, though, and “even the psychometricians who are responsible for test scaling shy away from making this assumption [about equal intervals]” (Harris, 2009, p. 329; see also Ballou, 2009).
A normal distribution, or a bell curve, of teacher effectiveness exists and can and should be used to norm large-scale standardized achievement test scores for VAM analyses. Half of all of America’s public school teachers is above average and the other half is below, so naturally (i.e., in terms of Social Darwinism), VAM estimates do not indicate whether a teacher is good or highly effective, but rather whether a teacher is better or worse than other “similar” teachers as reflective of a natural order (see, for example, B. D. Baker et al., 2013).
Scores derived via large-scale standardized achievement tests yield strong indicators about what students learn in schools. Test scores provide weak signals about educational quality and often much stronger signals about what students bring with them every year to the schoolhouse door, including out-of-school variables that are related to students’ backgrounds, peers, parents, families, neighborhoods, communities, and the like (see, for example, ASA, 2014; Berliner, 2013; Good, 2014).
Large-scale standardized achievement tests are the best and most objective measure we have, and because they are readily available and accessible, but despite the fact that they are designed to measure student achievement and not teacher effects, we should use them to make causal distinctions about teachers regardless. The tests on which VAM estimates are based are inherently flawed for measuring student achievement in and of themselves. Hence, aggregating and then using such already faulty measures to make causal statements about teacher effects simply exacerbate the issues (see, for example, B. D. Baker et al., 2013; E. L. Baker et al., 2010).
Conclusion
Not only are value-added analyses convoluted and technically intricate, the contexts within which they are to work—classrooms—are also highly complex, or rather complicated settings that vary by place, time, context, the individuals involved, and so forth. Appropriately, scholars from varied theoretical and methodological traditions have attempted to make sense of such complexities. However, as of the past few decades, the education narrative, generally, and the VAM narrative, specifically, have been more or less dominated by a single academic discipline—economics—which has largely shaped the conversation, as well as many large-scale policies surrounding this narrative. In alignment with economics-based principles, then, there seems to be a dominant view that complex matters, such as education, can be understood in parsimonious and quantifiable and, subsequently, measurably objective ways.
Teachers and teacher quality, specifically, have been reconceptualized in terms of a production function model where teachers produce a product that is consumed by students and is measurable in terms of how much “value” a teacher “adds” to the production. Given this formulation, and the economics-based epistemological position that society is made up of individuals who make rational decisions given the right conditions (Lazear, 1999), teachers are positioned as people who make good/bad decisions that result in added/detracted value to student learning, respectively. When approached from a different epistemological position, such as that of a sociologist, this notion fails to stand because teachers are presumed to be parts of larger social structures that influence not only teacher behavior but also the “outcomes” of teaching (e.g., student achievement). For this reason, the study of VAMs and VAM-based policies and practices should be approached from multiple disciplines to better capture the complexities involved in understanding the relationship between teaching and learning.
In this study, we have attempted to situate VAM practices and the VAM narrative within a multidisciplinary context by unpacking the conditions upon which VAMs and VAM-use are built. We attempted to do this by also mapping these conditions from within and onto the larger value-added literature. When taking a multidisciplinary approach, it becomes rather clear that the conditions that must stand as “true” for VAMs and VAM-use to be possible, feasible, and appropriate for that which they are intended and tasked to do, is much more difficult than is currently represented in the public narrative. The many assumptions that must be accepted and met to “truly” measure the isolated impact of teacher effects on student learning are, therefore, nearly (and most likely) impossible. Nonetheless, policies across the country are requiring school administrators not only to apply such methods to their teacher evaluation systems but also to often attach high-stakes personnel decisions to such evaluative outcomes.
This is where the true policy implications from this work dwell. If only policymakers, and perhaps more importantly those who inform policymakers, could begin to better understand the assumptions with which the educational community writ large must accept in order to accept teacher-level value-added estimates, they might be more critically aware of not only whether the assumptions are supported by research but also reasonable given the research. To critically consume and understand this information, perhaps educational policies that are more informed about versus bent on these models would help others move forward with the general teacher evaluation enterprise in much wiser and research-based ways.
Similarly, if there is any “truth” that can be gleaned from this work, it is that teachers work in highly complex institutions that are constantly affected by uncontrollable factors. Assuming that it is possible (or appropriate) to reduce teaching to a single numerical outcome is indefensible when taking into account the entire breadth of the literature on the topic. While VAM-based policies and their associated narratives would have the public assume that teaching can be simplified to a value-added estimate, other disciplines warn that doing so is impossible at best, and unethical at worst (Ravitch, 2014). Whereas a high level of methodological precision is almost always implied by those advocating VAM-use, the actual precision with which VAM estimations are made is almost never made transparent. This is particularly true in terms of transparency with data errors (e.g., missing data and data limitations), statistical errors (e.g., confidence intervals and standard errors of measurement), levels of reliability (e.g., consistency over time), and related types of evidence of validity (e.g., correlations with other similar measures; see, for example, B. D. Baker et al., 2013; Braun, 2005; Kane & Staiger, 2012). VAM estimations are, by definition, rough calculations that are largely imperfect, so “users of VAM[s] must resist the temptation to interpret estimates as pure, stable,” or even “true” measures of teacher effectiveness (Lockwood et al., 2007, p. 61; see also Rubin et al., 2004). Regardless, many state-level policymakers and other promoters still continue to endorse VAMs.
Lazear (1999) reminds us that it is the intention of the economist to make “simplifying assumptions” (p. 5) to sustain a level of statistical rigor. He also argued that economists are better at offering solutions than their noneconomics colleagues. It is feasible to assume, then, that such work has a certain appeal that might lure politicians and the public into trusting their analytical instruments and analyses. On the contrary, policy analysts, sociologists, psychologists, and other social scientists (e.g., educational researchers) tend to focus more on issues than solutions, which might be more challenging to directly translate into policy. This is not to say, however, that policy should depend solely on economics, but that policy should incorporate multidisciplinary approaches and analyses into policy consideration, development, implementation, and assessment. VAM research is but one example that demonstrates the problems that can arise with a reliance on a single discipline.
Footnotes
Authors’ Note
Dr. Jessica Holloway is now affiliated to Deakin University, Melbourne, AUS.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
