Abstract
Increasing job demands and continuing struggles to improve teacher evaluation practice raise the question of how peers might assist principals with teacher evaluation. Using a robust international sample (TALIS2013) of 36,411 teachers from 2,759 schools in 11 countries, we tested the hypothesis that teacher-led evaluation practices are associated with more teacher-reported positive changes in classroom practice, confidence, and motivation than principal-led evaluation practices in three areas evaluation: (1) classroom observations, (2) assessments of teacher content knowledge, and (3) analysis of student test score data. We found that teacher-led evaluation is associated with more positive feelings of motivation and change in practice for all three evaluation areas, but particularly for assessments of teacher content knowledge and test score data analysis. Further, principals’ reported use of extrinsic motivational tools to reward or punish teachers based upon their evaluation was also negatively associated with teachers’ motivation and reports of positive change in practice.
Keywords
Introduction
In their role as supervisors, principals have traditionally conducted teacher evaluations. During Race to the Top (RttT), which had an explicit goal of increasing the effectiveness of teacher evaluations, principals continued to shoulder the primary responsibility of formative and summative teacher evaluations, including: pre- and post-observation conferences, conducting formal and informal observations, providing feedback to teachers, and writing up summative evaluations. In some cases, such demands required that principals spend even more time observing teachers and documenting their effectiveness (Donaldson & Woulfin, 2018; Lavigne & Good, 2015). According to principals, accommodating these time demands has been an enormous challenge (Donaldson & Woulfin, 2018; Goldring et al., 2015; Kraft & Gilmour, 2016; Lavigne & Chamberlain, 2017). To cope, principals have resorted to cutting observations short, double-dipping, or reducing the number of observations they conduct (Donaldson & Woulfin, 2018; Stecher et al., 2018), calling into question the quality and/or utility of such feedback (Kraft & Christian, 2022). Even with the increased flexibility in teacher evaluations, principals have experienced little change in these instructional leadership demands since the passing of the Every Student Succeeds Act. In addition to the demands, the costs to schools and districts are substantial. It is estimated that it costs close to $700 million for principals to evaluate all 3.1 million U.S. K-12 public school teachers twice in a given year (Dynarski, 2016). Given these substantial investments, understanding if and how the information generated from teacher evaluation systems is translating into improved teaching and learning is critical.
The principalship has become, in recent years, increasingly challenging and complex work (Lavigne & Good, 2019; Ridge & Lavigne, 2020). With all of the pressures placed on principals, teacher supervision and evaluation frequently take a back seat. This is exacerbated by the increased evaluation loads and documentation demands from recent reform efforts that have resulted in principals spending more time writing their evaluations and recording data from observations instead of observing teachers and providing teachers with rich feedback (Flores & Derrington, 2017; Kraft & Christian, 2022). Furthermore, principals struggle to provide feedback that is content-based as they often lack critical pedagogical knowledge but often observe teachers in content areas outside of their expertise (Donaldson & Firestone, 2021; Kraft & Gilmour, 2016). As such, almost half of teachers report that the feedback they receive from their principal is not useful (Cherasaro et al., 2016). Together these findings point to many barriers to using principal-led teacher evaluations to improve teaching and learning. This is coupled with evidence that new teacher evaluation reform efforts supported by principal-led evaluations have generally failed to improve teaching and learning, particularly in the United States (Kim & Sun, 2021; Lavigne & Good, 2019).
This raises the question of how the underutilized practice of peer-based evaluation approaches (e.g., coaching and/or peer assistance and review; Goldstein, 2007; Harvey et al., 2019; Johnson, 2019; Johnson et al., 2010; Ridge & Lavigne, 2020) might be useful in supporting (or replacing) principal-led teacher evaluations as the primary means by which teachers are evaluated, thus potentially reducing the burden currently shouldered by administrators (Galey-Horn & Woulfin, 2021). Although peer-based evaluation approaches offer substantial promise and are commonly practiced outside of the U.S., it has continued to be a relatively uncommon practice in the U.S. (Johnson, 2019). As such, these approaches have been understudied and largely absent from reviews of effective teacher evaluation systems in the U.S. (National Council on Teacher Quality, 2018).
However, programs that utilize expert or more experienced peers as evaluators have some demonstrable benefits, according to the literature (Johnson, 2019; Johnson et al., 2010; Papay & Johnson, 2012). While not without their own set of organizational and financial challenges, there is some evidence that peer-coaching/evaluation models might reduce teacher turnover as well as improve teacher quality, instruction, and with it, teacher confidence (Johnson, 2019; Papay & Johnson, 2012). Likewise, teacher-led observation provides a mechanism to help schools and districts consider different possibilities for separating teacher supervision and evaluation. Oftentimes the principal is the coach that supervises teachers throughout the school year, providing formative feedback, and also serves as the judge in the summative task of teacher evaluation. This structure results in supervision traveling under the guise of evaluation (Glanz & Hazi, 2019; Hazi, 1994), which is a problem that continues to plague teacher evaluation even under the Every Student Succeeds Act. Twenty-five states continue to embed formative feedback within the summative portion of their teacher evaluation models (Mette et al., 2020).
Yet a transition away from administrator-led evaluation toward more peer-led models should be based upon evidence that teachers and their practice would benefit from the change, and, in this case, more evidence is needed. The literature does seem clear that teachers are more likely to improve in their craft when they have experienced evaluators that deeply understand teaching and know what teachers need/want to learn (Donaldson & Firestone, 2021; Ford, 2018; Kraft & Christian, 2022). An important part of understanding the potential benefits of teacher-led evaluation concerns teacher cognitive sensemaking about their evaluation experience and how it does (or does not) motivate change in practice (Donaldson, 2021; Ford et al., 2017, 2018). There is growing evidence to suggest that fellow teachers have strong potential to be effective evaluators of other teachers’ practice, particularly when the evaluation emphasizes formative approaches (Donaldson, 2021; Johnson, 2019). There are more obvious reasons why this might be the case: experienced teachers are more likely to have intimate knowledge of practice to help propel novice teachers forward and they can be tapped based upon needed content area expertise (Johnson, 2019). Yet there is also evidence that how evaluation is framed including how teachers experience the evaluation process can influence teachers’ perceptions of the credibility, utility, and/or validity of the feedback and thus their likelihood of using it for instructional improvement (Cherasaro et al., 2016; Ford, 2018; Ford et al., 2017; Jiang et al., 2015; Pallas, 2021; Reddy et al., 2018). There is arguably no one in the school building better equipped to understand how teachers experience evaluation and thus no one more likely to prioritize teachers’ needs throughout the evaluation process than a fellow teacher. However, direct comparisons of the outcomes of administrator versus teacher-led evaluations are largely absent within the literature.
Therefore, the purpose of this paper was to test the hypothesis that teacher-led evaluation practices are related to more teacher-reported positive changes in classroom practice, increased confidence, and motivation than principal-led evaluation practices. To test this hypothesis, we used a robust international sample of schools from the TALIS 2013 to acquire the kind of variation in teacher evaluation practice needed to draw meaningful conclusions. Evidence supporting our hypothesis would be seen in the form of a positive relationship between teacher/mentor-led evaluation and teacher positive change in motivation/practice and/or a negative relationship between principal/management-led evaluation and teacher positive change in motivation/practice. The three research questions below which guided the study derived from a central, overarching question: Does who is doing the evaluation matter?
In their review of effective teacher evaluation systems, the National Council on Teacher Quality (2018) outlines seven principles that were common across all models: multiple measures of teacher effectiveness, student surveys, objective measures of student growth, at least three rating systems, annual evaluations and observations of all teachers, professional development that is tied to teacher evaluation, and written feedback after each observation. Recognizing this, in this study, we focus on the most critical as well as common and enduring components of teacher evaluation models: the use of observational data—also one of the most commonly used components in TALIS-participating sites (Organization of Economic CoOperation and Development [OECD], 2013)—and analysis of test score data (Lavigne & Good, 2020). However, with the flexibility of the Every Student Succeeds Act in mind, we also consider more fine-grained measures and practices such as the use of teacher content knowledge assessments. Thus our specific research questions each address one of the formal evaluation components we have chosen:
Is there a relationship between who is leading the formal observations of classroom teaching and teacher perceived positive change in motivation/practice?
Is there a relationship between who is leading the assessment of teacher content knowledge and teacher perceived positive change in motivation/practice?
Is there a relationship between who is leading the analysis of student test score data and teacher perceived positive change in motivation/practice?
Teacher Evaluation: Then and Now
Glanz and Hazi (2019) trace the development of teacher evaluation back to the origins of supervision, a tool that emerged to provide efficiency for superintendents during the rapid economic growth that transformed schools during and following the industrial revolution (Tyack, 1974). At that time, supervision was established as a role that fell under the purview of school administration (Payne, 1875). In the 1900s, scientific management and efficiency pushed for supervision as a means of coordinating and controlling labor (Bobbitt, 1913; Taylor, 1911). According to Glanz and Hazi, the 1920s marked a shift toward a more democratic and middle-manager as opposed to a bureaucratic approach under the responsibility of school administration (Dewey, 1929; Hosic, 1920; Newlon, 1923). Supervision began to establish its professional autonomy and then merged with the field of curriculum. In the 1990s, the accountability movement elevated teacher evaluation and supervision began to function incognito. This shift coupled with the emergence of instructional leadership meant that teacher supervision and evaluation primarily resided under the role of school leaders (Glanz & Hazi, 2019).
During the last decade, reform efforts have placed teacher evaluation at the center of accountability and school improvement efforts and have focused on the effectiveness of principals in that role. Launched in 2009, Race to the Top allocated funds to states to develop comprehensive systems of improvement in four areas: college and career readiness, teaching and leadership, data systems and technology to improve instruction, and turnaround efforts for low-performing schools. Reforms to teacher evaluation fell under teaching and leadership and included the use of standardized tests as a significant measure of teaching effectiveness, multiple rating categories, and high-stakes teacher evaluations (Lavigne & Good, 2019). While the 19 states that won Race to the Top dollars were significantly more likely to implement Race to the Top policies and offer less flexibility (Gagnon et al., 2017), its effects on teacher evaluation were widespread with changes noted in even 56% of the states that never applied for Race to the Top (Howell, 2015).
Today, many of the components of teacher evaluation systems that emerged in Race to the Top continue to sustain under the Every Student Succeeds Act (ESSA) despite greater flexibility (Lavigne & Good, 2019; Walsh et al., 2017). One important change under ESSA is that student achievement is no longer required as a significant measure in a teacher evaluation model (U.S. Department of Education, 2015). This has been the area that has changed the most as states have made modifications. In particular, states have decreased their use of value-added models (VAMs) within their teacher evaluation systems and have made steps to elevate and emphasize formative feedback practices, reversing some of the changes to teacher evaluation models from the Race to the Top era (Close et al., 2020).
Principal-Led Evaluations
One aspect of teacher evaluations that has remained relatively constant over the years is the primary role of the principal. It also seems clear that, since RttT, teachers are receiving, on average, more feedback than in times past (Donaldson, 2021). However, the documentation demands of recent reform efforts have been so burdensome that principals spend more time writing their evaluations and recording data from observations than they do observing teachers and providing them with rich feedback (Flores & Derrington, 2017). This has also hindered principals’ ability to provide frequent feedback (Kraft & Gilmour, 2016), document deficiencies (Range et al., 2011), and rate teachers ineffective (Kraft & Gilmour, 2017). Principals have cut observations short, double-dipped or reduced the number of observations, or have been unavailable to address teacher concerns (Donaldson & Woulfin, 2018; Stecher et al., 2018).
Given the overwhelming emphasis on training principals on new teacher evaluation policies, models, and observation instruments, principals still receive limited training on how to provide feedback despite plentiful science on cognition and performance appraisal to inform such training (see Donaldson, 2021 for a review). As such, principal evaluations tend toward positive affirmations of teachers’ practices rather than suggestions for improvement (Kraft & Gilmour, 2017). This is compounded by the fact that principals oftentimes observe and provide feedback to teachers in content areas outside of their area of expertise (Cherasaro et al., 2016) which results in principals providing feedback almost entirely focused on pedagogy (Kraft & Gilmour, 2016). Likewise, principals and districts often have different responses and agency in relation to teacher evaluation that prompts them to tinker, modify, and adapt teacher evaluation systems to align with their beliefs and to meet local needs and concerns (J. Cohen et al., 2020; Donaldson & Woulfin, 2018; Marsh et al., 2017; Woulfin et al., 2016). To a certain degree, these behaviors belie a recognition, however implicit, that there are aspects of evaluation policy and practice in need of adjustment or reconsideration.
While, in one study, teachers indicated that they received timely, frequent, and accurate feedback that included specific suggestions for improvement, just a little more than half (55%) felt that the feedback they received was useful. Further, only 62% indicated they had access to professional development, 61% had planning time to implement new strategies, and 60% had support from an instructional leader. Notably, only 33% had access to observing an expert teacher model feedback (Cherasaro et al., 2016). Other studies have found similar teacher perceptions and rates of support for evaluation (see Firestone & Donaldson, 2019 for a review). Teachers who do not receive specific feedback tend to report lower self-efficacy (E. C. Smith et al., 2020); furthermore, failing to address content knowledge translates to missed opportunities to increase teacher effectiveness (Hill et al., 2008). Together these findings point to many barriers to using principal-led teacher evaluations to improve teaching and learning and indicate a general paucity of opportunities to learn from their peers in meaningful ways, including through feedback on instruction.
The Potential of Teacher-Led Evaluation and Enhancing Motivation for Improvement
Fortunately, greater flexibility under the Every Student Succeeds Act offers the opportunity to think about teacher evaluation in different ways, something that has not yet been realized in most states (Kim & Sun, 2021; Mette et al., 2020). Some have advocated for a more facilitative rather than evaluative role of the school leader that focuses primarily on supporting and developing teachers in the refinement of their skills and practices (Ford, 2018; Ford & Hewitt, 2020; Guenther, 2021; Murphy et al., 2013). The inherent challenges of principal-led evaluation raise the possibility of considering the underutilized practice of peer coaching and observation (Harvey et al., 2019; Ridge & Lavigne, 2020), of which there is growing empirical support in the literature (Donaldson, 2021; Galey-Horn & Woulfin, 2021; Johnson, 2019; Johnson et al., 2010; Papay & Johnson, 2012).
In their analysis of TALIS 2013 data, Ford et al. (2018) found that teachers who were primarily evaluated by a fellow teacher as opposed to the principal reported higher overall job satisfaction. Motivation theory, while underutilized in understanding the effects of evaluation approaches and tools on teacher motivation (Donaldson, 2021; Ford & Hewitt, 2020; Ford et al., 2018), provides some explanation as to why peer-based evaluation models might make for happier teachers, for generating better feedback, and ultimately inducing teachers to use that feedback to improve their practice. There is ample evidence to suggest that teachers, on the whole, are attracted to and remain in the profession for primarily intrinsic reasons (Kennedy, 2005; Lortie, 1975; Watt & Richardson, 2008). Having entered willingly into a human improvement enterprise (D. K. Cohen, 2011), teachers are naturally enthusiastic about their profession, invested in seeing the children under their care succeed, and are typically willing to use any and all means to make that happen—including working to improve their teaching (Kunter & Holzburger, 2014; Roth, 2014). Teachers, in general, do not object to being evaluated—in fact, many request to be—they object to the way it is often carried out (Ford et al., 2017; Guenther, 2021; Hewitt, 2015; Johnson, 2019). This is partly because many evaluation systems (in the US in particular), continue to misconstrue lack of improvement and/or poor performance as a lack of motivation on the part of teachers and thus continue to heavily rely on extrinsic motivational mechanisms such as rewards (pay-for-performance) and punishment (cut pay, censure, termination, and job/grade changes; Ford et al., 2018). While using reward and punishment might work for individuals who are less-than-enthusiastic about their work, activating existing intrinsic motivation requires a very different approach—one which is focused on ensuring that conditions are ripe for individuals’ innate orientations to their work, learning, and development to emerge (i.e., teachers’ psychological needs; Ryan & Deci, 2017). Extensive research has found that extrinsic motivational mechanisms are likely to: (a) only change behavior in the short-term, (b) damage intrinsic motivation and its long-term development, and (c) induce gaming practices that result from reward or punishment being attached to outcomes instead of behaviors (Ryan & Brown, 2005; Ryan & Deci, 2017; Ryan & Weinstein, 2009).
Addressing teachers’ psychological needs for learning and development is something that teacher/peer-led evaluation seems uniquely suited to do. In studies of teacher evaluation, teachers often desire and request more and better feedback on their practice (Donaldson, 2021; Ford, 2018; Gabriel & Woulfin, 2017; Johnson, 2019). Providing constructive feedback along with support to address the areas of their practice that need improvement is key to building a teachers’ feelings of competence and self-efficacy (Ford et al., 2017; Lavigne & Good, 2015; E. C. Smith et al., 2020). Assuming that teacher-led evaluations foster more content-specific feedback—an important feedback characteristic that principals struggle to provide (Kraft & Gilmour, 2016; Smith et al., 2020)—teaching peers are more likely to have a grasp on subject-specific concerns (Kraft & Gilmour, 2016). Moreover, more experienced teachers, with their wealth of knowledge of practice, have a greater capacity to show teachers where they are missing the mark and model alternative approaches (Johnson, 2019). They are also able to help them set realistic goals for student learning growth and provide a structure, process, and point person (or people) for reviewing progress toward meeting those goals (Longchamp, 2017). Perceived expertise of the evaluator in an of itself enhances the credibility of the feedback and, with it, its likelihood of being used by the evaluee to improve their practice (Donaldson, 2021; Ford, 2018). Second, with models such as peer assistance and review (PAR), observations are more frequent and intentional, yielding the perception that the evaluation process as a whole is meaningful and that the feedback is more representative of their “true” classroom practice (Ford, 2018; Johnson, 2019).
The need for autonomy is also critical for teachers throughout the evaluation process. Having voice, choice, and a semblance of control in how evaluation is conducted as well as the goal setting, assessments, and data used to evaluate is important for meeting this need (Donaldson, 2021; Lavigne & Good, 2015). Having peers lead evaluations can disrupt hierarchical power dynamics between the principal and teacher in traditional administration-led evaluations, increasing the potential for collaboration and engagement so it feels less like evaluation is “done to them” (Danielson, 2010). Furthermore, teacher emotions play an important role in how teachers respond to evaluation and so having some consideration of teachers’ fears and concerns when engaging in evaluation can help them overcome the initial negative responses they often have to receiving feedback. The power dynamic which invariably exists between principals and teachers can make this difficult to do and it can also limit the degree to which teachers perceive the principal as a colleague offering constructive advice instead of a supervisor providing a summative performance evaluation (Donaldson, 2021). Because of the highly personal nature of feedback, teaching peers who share in the experience of teaching and evaluation with all of its inherent challenges and rewards are more likely to develop a relationship that mitigates the typical affective and motivational barriers to change and improvement that arise in evaluation situations (Donaldson, 2021; Lavigne & Good, 2020; Ryan & Deci, 2017). This is not to say that principals could not learn how to fulfill these roles, but there is much evidence to suggest that they are, on the whole, not yet provided the opportunity and structures to do so.
Method
Sample
The primary data source for this study was the 2013 Teaching and Learning International Survey (TALIS 2013), administered by the Organisation for Economic Co-operation and Development (OECD). The data analyzed in this study were collected from all teachers who work in lower secondary schools (level 2 of International Standard Classification of Education [ISCED]) from a selected number of schools and nations out of the 32 included in the TALIS 2013 dataset and international report (OECD, 2014b). While there is a more recent administration of the TALIS (TALIS 2018), TALIS 2013, unlike TALIS 2018, focused heavily on the appraisal of teachers’ work in schools and thus included the detailed information on these processes necessary for the analysis.
For each country sampled, TALIS 2013 set a target size of 200 schools with 20 teachers per school. Schools were selected according to a national sampling plan which used systematic random sampling with probability proportional to size (PPS) within explicit strata which might include school types, regions, or funding (OECD, 2014a). In selecting a sample of TALIS countries appropriate to answer our research questions, we adopted a purposive sampling approach to maximize variability in our focal variable, teacher or leader-led evaluation. Thus, we utilized an empirical approach to country selection, including both countries where a target proportion of the schools in the country reported engaging in either a high proportion of teacher or leader-led evaluation (a high proportion of teacher-led evaluation being a relative criterion as fewer schools and countries engage in this as a primary form of evaluation). This approach was informed by global perspectives on accountability contexts (see Holloway et al., 2017). Using vignettes from South Korea, the United States, Hungary, and Mexico, W. C. Smith (2014) argues that testing for accountability purposes is an increasingly common perspective being adopted globally. W. C. Smith and Kubacka (2017) argue that while rarely used in isolation, the use of student test scores to evaluate teachers is the most common approach and the extent to which test scores are emphasized impacts teachers’ perceptions of feedback effectiveness. Further, having school leaders or other administrators be responsible for leading the evaluation of teachers is part-and-parcel of the “managerial order” characteristic of accountability systems (Holloway et al., 2017).
Thus, our empirical approach was to separately examine all three analytic samples (i.e., teaching observations, assessments of content knowledge, and analysis of test scores) and determine a contrasting set of countries with very little (or no) teacher-led evaluation happening (i.e., primarily or exclusively principal/admin-led) or a relatively high proportion of engagement in teacher-led evaluation. Three countries that consistently reported strong leader-led teacher evaluation across the three analytic samples were the United States, England, and Chile and this appears consistent with what we know of these countries’ emphasis on accountability at the time the data were collected (Falabella & De La Vega, 2016).
For countries to be considered “high teacher-led evaluation” countries, a criterion for inclusion was established separately for each analytic sample relative to the frequency of teacher-led evaluation occurring in that feedback type. For the observation and content knowledge analyses, countries were included if more than 10% of the sample within a country was teacher-led. For the test score analysis, countries were included if they had greater than the average (arithmetic mean) number of teacher-led evaluation schools. Applying these criteria resulted in Brazil, France, Portugal being included as “high teacher-led evaluation” countries for the observation analysis (thus six countries total including the U.S., England, and Chile); Brazil, France, Portugal, Japan, Korea, Mexico, and Romania being included for the content knowledge analysis (thus 10 total countries); and Brazil, Japan, Korea, and Croatia for the test score analysis (thus seven total countries). The base analytic sample from which the sample for each of the three focal analyses was derived was 36,411 teachers within 2,759 schools in 11 countries.
Measures
Items from both the teacher and principal questionnaires were used in this analysis. Descriptive statistics for the pooled sample of 11 countries on all of the variables used in the regression analyses and their corresponding survey numbers in the TALIS surveys are provided in Table 1. Below we discuss specifically the outcome and key independent variable in the analysis.
Descriptive Statistics for Study Variables From the International TALIS 2013 Dataset.
Note. Raw values for all variables in this table, except the outcome, teacher motivation, and change in practice, in which Rasch scale scores are reported. Appropriate weights applied for descriptive statistics and reliability analyses.
TALIS 2013 utilized confirmatory factor analysis (a classical test theory [CTT] approach) for its measure construction and scaling (OECD, 2014a). The Rasch model, in contrast, is an Item Response Theory (IRT) approach and is distinguished from classical test theory in considering the ability of respondents in tandem with the difficulty of the items to which they are responding. Advantages and disadvantages of both notwithstanding (see, e.g., Singh, 2004), both approaches are useful in the development and scaling of latent measures; in fact, other prominent international education datasets have opted to use an IRT approach for measure construction and scaling (e.g., TIMSS and PISA). Our choice to adopt an IRT approach over CFA was due to the exploratory nature of our development of the teacher change in motivation/practice outcome measure, which required a wider range of information to assess person and item performance than is typically provided using a CTT approach.
In addition to a host of other diagnostic information, the WINSTEPS program produces a scaled-score for each teacher in log-odds units which represents where each teacher’s perceptions locate him/her on the continuum of teacher change in practice/motivation (low, negative values reflect perceptions of a less motivation and change in practice, and high, positive numbers reflect perceptions of more motivation and change in practice). We set our threshold at mean-squared values of 0.5 to 1.5—accepted thresholds for Winsteps analysis (Linacre, 2014). Items from both the teacher and principal questionnaires were used in this analysis.
Positive Change in Teacher Motivation and Practice
This was our primary outcome measure, which was constructed from teacher questionnaire items TT2G30F, H, J, L, M, and N. These items elicited teachers’ perception of the degree to which the feedback they received from their evaluation had resulted in a positive change in their: confidence, motivation, satisfaction, and teacher classroom practices including classroom management strategies and use of student assessments for learning. The item stem was: Concerning the feedback you have received at this school, to what extent has it directly led to a positive change in any of the following? (TT2G30). This variable was Rasch analyzed and the items and results of the Rasch analysis are located in Appendix A.
Leader/Teacher-Led Evaluation Practices
This variable derived from principal questionnaire items TC2G28A1-6, C1-6, and D1-6 which ask who in the school performs the following tasks as a part of the formal teacher evaluation process: (a) direct observation of classroom teaching, (b) assessments of teachers’ content knowledge, and (c) analysis of students’ test scores. For the purposes of this analysis, we dichotomized these items which classified schools as sites where evaluation was conducted by: (a) principal or school management personnel, (b) teacher or teacher mentors, and (c) for schools where both of the aforementioned groups conducted the evaluation. Schools that indicated no use of the practice or that it was conducted by external individuals or bodies were excluded entirely from the analysis.
Principal Use of Extrinsic Motivational Tools After Evaluation
One important factor in the effects on teacher motivation and practice following an evaluation is how principals attach rewards or consequences to the results (Ford, 2018). This composite measure includes items that ask principals about the frequency with which they tie extrinsic rewards or punishments to the teacher’s evaluation performance such as: changes to pay (either up, e.g., the use of bonuses, or down, e.g., censure or dismissal, or changes to job responsibilities). These were on a 4-point rating scale from “never” to “always.”
Analytic Strategy
Teachers in the international sample who were missing scores on any one of the items comprising the outcome were removed before analysis (i.e., the values were not imputed). Teacher change in motivation/practice, as well as several other similar perception items toward the end of the TALIS teacher survey (such as self-efficacy and climate perceptions) exhibited a unit nonresponse pattern (Enders, 2010). After reviewing the missing data coding procedure in the TALIS technical manual, the missing codes in the dataset indicated that nearly all of the teachers who were missing scores for the outcome either returned the survey blank or incomplete (OECD, 2014a).
In this case, we endeavored to determine whether or not there were significant differences between teachers who were missing a score for our outcome and other measured variables—in other words, could the data be assumed to be missing completely at random (MCAR), which is needed to justify listwise deletion. To test this assumption, we conducted a series of Bonferroni corrected t-tests and chi-squared tests of independence between teachers who had a complete score for our outcome and those for whom it was missing. We found no significant differences between the groups with respect to these covariates, and thus list-wise deletion was a justifiable missing data handling approach (Enders, 2010). After list-wise deletion was completed, missing data at the teacher level in study covariates was around 3%, while missing data at the school level was around 5%, and both exhibited a general item and unit non-response patterns (Enders, 2010). Then, we employed multiple imputation (MI) techniques to the raw, teacher and school-level data. Multiple imputation is substantially more robust than typical list or pair-wise deletion procedures to missing data bias, and results in multiple versions of the same dataset with different plausible values for the missing data based on available variable data and their underlying covariance structure (Enders, 2010).
Once teachers without scores on the teacher motivation/practice were removed, the remainder of the missing data exhibited a general item non-response pattern (De Leeuw et al., 2003). A general item non-response pattern manifests as gaps in item response that appear to be randomly dispersed throughout the dataset (i.e., MCAR). Because multiple imputation does not require us to invoke an assumption of MCAR for the results to be unbiased, and there was no evidence to suggest they were missing not at random (MNAR), we chose the less stringent assumption of MAR, which allows for missing data on a variable to be related to other measured variables in the analysis (Enders, 2010).
Two-Level HLM Analysis
Because there are typically several different dimensions to a teachers’ evaluation (e.g., lesson observations, test score growth, etc.), we did not assume that all aspects of the evaluation process exhibited differences depending on whether they were conducted by an administrator versus a fellow teacher; rather we investigated each practice separately with respect to this hypothesis. We used variable information from the principal questionnaire on teacher/leader-led evaluation practices (described above) to conduct three separate 2-level HLM analyses for each of our focal areas of teacher evaluation: (1) classroom observations, (2) assessments of teacher content knowledge, and (3) analysis of student test scores. Separate Level 2 files were created for each evaluation type with the same set of imputed level 1 files. A series of “do if” statements were run to flag school principals who indicated that they or members of school management conducted the formal evaluation procedure in question (principal/management), other teachers or assigned mentors (teacher/mentor) conducted the evaluation, or both. Any schools that indicated that the evaluation type in question was not used or done by external individuals or bodies were excluded from the analysis. This procedure led to an effective sample size for the classroom observation analysis of 16,428 teachers in 1,400 schools; teacher content knowledge analysis, 18,615 teachers within 1,441 schools; and for the test score analysis, 21,246 teachers within 1,667 schools.
In addition to the focal dummy predictors of primary evaluator (principal/management, teacher/mentor, or both), we included other important teacher and school characteristics, attitudes, and perceptions about school working conditions and teacher appraisal practices that were presumed to be related to the outcome based on prior research. School-level variables related to the outcome such as: school climate, principal, and school characteristics such as school climate, poverty, urbanicity, sector, and principal satisfaction as well as the primary evaluator dummy variables were used to model between-school variation in teacher change in practice/motivation. The teacher weight, TCHWGT, was incorporated into the final analysis to maintain the intended representativeness of the sample. All other teacher-level effects remained fixed at Level 2, as an analysis revealed that there was little between-school variation in the relationships between each of the teacher-level predictors and positive change in practice/motivation.
Finally, because the sample sizes in this analysis were relatively large and therefore substantial power to detect small effects was present, Keith’s (2015) effect sizes were used to provide additional context for findings. Standardized coefficients of β ≥ .05 reflected a small effect, β ≥ .10 reflected a medium effect, and β ≥ .25 reflected large effects. School-level sample sizes were sufficient to use robust standard errors to adjust for heavier tails present at the highest intervals of many of the climate measures in the TALIS dataset. Robust standard errors can reasonably adjust for any non-normality in the outcome that can lead to heteroscedasticity and thus provide more accurate significance tests and confidence intervals (Hox, 2010).
Findings
The first research question addressed differences in teacher perceptions of positive change in their motivation/practice with respect to who was responsible in the school building for conducting direct observations of their classroom practice. Table 2 displays these results. The null model ICC demonstrates that approximately 16% of the variation in teacher perception of positive change in practice/motivation was between schools. Model 1 presents only the main effects; Model 2 the main effects with teacher covariates; The final model includes all variables including school-level covariates.
HLM Analysis of Teacher Motivation and Change in Practice From Evaluation by Direct Observation of Classroom Teaching.
Robust standard errors reported. Coefficient estimates in this table are the pooled estimates from the five imputed datasets provided by the HLM program, with the teacher weight applied (TCHWGT). All continuous variables standardized. Large cities (1,000,000+) and teachers primarily evaluated by other teachers/mentors were the comparison/holdout groups for those dummy variable sets.
p ≤ .001. **p ≤ .01. *p ≤ .05.
As the final model of evaluation by direct observation of teaching demonstrates, holding other covariates constant, teachers whose principal/administration were the sole observer(s) of classroom practice in their building reported significantly lower overall positive change in practice/motivation, β = −.202, SE = 0.061, p < .001, as compared to those who were observed by a fellow teacher or assigned mentor only, β = .053, SE = 0.095, p = n.s. In schools where both principals/administration and teachers were responsible for classroom observations, the results were very similar to those for principals only—nearly a quarter of a standard deviation lower on average—both of which are medium-to-strong effects, β = −.226, SE = 0.063, p < .001.
Covariates exhibiting medium-to-large positive effects on positive change in teacher practice/motivation were: teacher job satisfaction, β = .220, SE = 0.018, p < .001, shared responsibility and decision making, β = .199, SE = 0.017, p < .001, having been assigned a mentor, β = .196, SE = 0.036, p < .001. Positive medium-sized effects were found in teacher self-efficacy, β = .141, SE = 0.018, p < .001, collaboration intensity, β = .127, SE = 0.020, p < .001, and teaching in a school with a female principal, β = .164, SE = 0.042, p < .001. Teacher age, β = −.087, SE = 0.020, p < .001, and experiencing high barriers to professional development in the past year, β = −.111, SE = 0.019, p < .001, were small-to-medium-sized negative effects related to positive change in teacher practice/motivation. One final small negative relationship key to the investigation was found between positive change in teacher practice/motivation and principal’s use of extrinsic motivational tools after the observation of classroom teaching, β = −.050, SE = 0.019, p < .01.
The second research question addressed differences in teacher perceptions of positive change in their motivation/practice with respect to who was responsible in the school building for conducting assessments of teacher content knowledge. Table 3 displays these results. The null model ICC demonstrates that 20% of the variation in teacher perception of positive change in practice/motivation was between schools.
HLM Analysis of Teacher Motivation and Change in Practice From Evaluation Using Assessments of Teacher Content Knowledge.
Robust standard errors reported. Coefficient estimates in this table are the pooled estimates from the five imputed datasets provided by the HLM program, with the teacher weight applied (TCHWGT). All continuous variables standardized. Large cities (1,000,000+) and teachers primarily evaluated by other teachers/mentors were the comparison/holdout groups for those dummy variable sets.
p ≤ .001. **p ≤ .01. *p ≤ .05.
As the final model of evaluation by using assessments of teacher content knowledge demonstrates, holding other covariates constant, teachers whose principal/administration were the sole evaluator of teacher content knowledge reported significantly lower overall positive change in practice/motivation, β = −.183, SE = 0.057, p < .001. This as compared to those who were observed by a fellow teacher or assigned mentor only, which was positively associated with teacher positive change in practice/motivation, β = .270, SE = 0.089, p < .01. In schools where both principals/administration and teachers were responsible for classroom observations, the results were very similar to those for principals only—one quarter of a standard deviation lower on average, β = −.252, SE = 0.065, p < .001. Thus, these results approach a nearly one-half standard deviation net difference in the positive change in practice/motivation between teachers evaluated by principals only versus teachers only (0.183 + 0.270 = 0.453).
Covariates exhibiting medium-to-large positive effects on positive change in teacher practice/motivation were similar to in the other models: teacher job satisfaction, β = .217, SE = 0.016, p < .001, shared responsibility and decision making, β = .188, SE = 0.018, p < .001, having been assigned a mentor, β = .164, SE = 0.033, p < .001. Positive medium-sized effects were found in teacher self-efficacy, β = .131, SE = 0.016, p < .001, and collaboration intensity, β = .148, SE = 0.019, p < .001. Experiencing high barriers to professional development in the past year, was a medium-sized negative effect related to positive change in teacher practice/motivation, β = −.104, SE = 0.019, p < .001. Emerging in this model again was a key negative association between positive change in teacher practice/motivation and principal’s use of extrinsic motivational tools after evaluation, β = −.090, SE = 0.021, p < .001.
The final research question addressed differences in teacher perceptions of positive change in their motivation/practice with respect to who in the school building is conducting evaluation using the analysis of student test scores. Table 4 displays these results. The null model ICC demonstrates that approximately 17% of the variation in teacher perceptions of positive change in practice/motivation was between schools.
HLM Analysis of Teacher Motivation and Change in Practice From Evaluation Using Analysis of Students’ Test Scores.
Robust standard errors reported. Coefficient estimates in this table are the pooled estimates from the five imputed datasets provided by the HLM program, with the teacher weight applied (TCHWGT). All continuous variables standardized. Large cities (1,000,000+) and teachers primarily evaluated by other teachers/mentors were the comparison/holdout groups for those dummy variable sets.
p ≤ .001. **p ≤ .01. *p ≤ .05.
Holding other covariates constant, teachers whose principal/administration were the sole evaluators involved in analysis of student test score data reported significantly lower overall positive change in practice/motivation, β = −.199, SE = 0.059, p < .05, as compared to those who were observed by a fellow teacher or assigned mentor only, β = .135, SE = 0.107, p = n.s. In schools where both principals/administration and teachers were responsible for classroom observations, the results were very similar to those for principals only—over one-quarter of a standard deviation lower on average, β = −.252, SE = 0.092, p < .01. Both the principal-only and principal-and-teacher evaluation effects were medium-to-strong in size.
Covariates exhibiting medium-to-large positive effects on positive change in teacher practice/motivation were similar to the other two models: teacher job satisfaction, β = .210, SE = 0.016, p < .001, shared responsibility and decision making, β = .202, SE = 0.017, p < .001, having been assigned a mentor, β = .163, SE = 0.028, p < .001. Positive medium-sized effects were found in teacher self-efficacy, β = .111, SE = 0.014, p < .001, collaboration intensity, β = .134, SE = 0.018, p < .001. Principal job satisfaction, β = −.082, SE = 0.024, p < .001, and experiencing high barriers to professional development in the past year, β = −.088, SE = 0.017, p < .001, were medium-sized negative effects related to positive change in teacher practice/motivation. Emerging in this final model, similar to the other two models, a small, negative association between positive change in teacher practice/motivation and principal’s use of extrinsic motivational tools after evaluation, β = −.082, SE = 0.020, p < .001.
In all three models discussed above, some notable small effects were those between teacher age and female principal. All three models clearly show that as teacher age increased, there was a corresponding small decrease of positive change in motivation/practice, β = −.078 to −.087 across the three models all p < .001. With the exception of the assessment of teacher content knowledge analysis, teachers in schools where the principal was female also had a small overall positive effect on teacher perceptions of positive change in motivation/practice, β = .164 and .101 for principal observations and analysis of test scores respectively.
One important consideration given study findings is how teachers might have differed in quality and experiences between schools where teachers were the primary evaluators versus other types of schools. It was certainly possible that schools where teachers were those leading evaluation were unique in some fashion: either having an unusually large number of expert teachers, teachers with other capacities or experience, or else enjoyed collective engagement in teaching and learning/leadership that might have made it more logical for them to become leaders in those activities in their school. A series of Bonferroni corrected independent samples t-tests comparing teacher years of experience, full-time status, perceptions of preparation, collaboration intensity, shared responsibility, satisfaction, learning climate, and shared decision making in teacher-only schools as compared to the others. Table 5 presents the results of this analysis. Overall, we found that teachers in teacher-only evaluation schools were, on the whole, not significantly different or were significantly lower in all of these characteristics than in principal-only evaluation and principal/teacher evaluation schools.
Independent Samples T-Tests of Differences in Characteristics of Teacher-Led Evaluation Schools Versus Other.
Actual p-values reported. Significance markers (asterisks) Bonferroni adjusted. Principal final weight applied for this analysis (SCHWGT).
p ≤ .001. **p ≤ .01. *p ≤ .05.
Discussion
Overall, the findings demonstrate that teacher-led evaluation is associated with more positive feelings of motivation, satisfaction, as well as change in practice. This was true for all facets of the evaluation process we examined—lesson observation, assessments of teacher content knowledge, and analysis of test score data. This evidence suggests, at the very least, we consider more carefully the role that expert peers and mentors can play in the evaluation process to maximize the effects that it can have on teachers’ classroom practice, confidence, and motivation. They also lend credence to our claim that how evaluation is done matters for how teachers ultimately feel about their feedback and how it motivates (or does not motivate) change in practice. Further, our results corroborate and expand upon those found by other researchers who have urged policymakers and leaders to consider models of peer coaching and feedback as a potentially beneficial approach to teacher evaluation (Donaldson, 2021; Galey-Horn & Woulfin, 2021; Johnson, 2019; Johnson et al., 2010; Ridge & Lavigne, 2020).
They also support the findings of other studies which have claimed that assessments of content knowledge and analysis of test scores, because they are context specific, might demand specialized knowledge that administrators might lack in order for the discussions to be beneficial for use in instructional improvement (Kraft & Gilmour, 2016; E. C. Smith et al., 2020). The nature of such discussions seem also to be more time-demanding in order for them to be done well, which, given that principals have incredibly busy schedules and struggle with the time demands of teacher supervision and evaluation (Flores & Derrington, 2017), might be best left to mentors and fellow teachers. Furthermore, the results of the use of test score data in evaluation support past findings critical of now-questionable practices like the use of VAM results which many states have discontinued because of their inherent measurement problems and the visceral reactions of educators to their use (Close et al., 2020; Ford et al., 2017).
Finally, there is the question of the effectiveness of extrinsic motivational tools in accomplishing the main purposes of evaluation. Our study found that principals’ reported use of these tools (either rewards such as increased pay or punishments such as censure, dismissal, or change of job responsibilities) to reward or punish teachers based upon their evaluation results was associated with a modest yet significantly lower motivation and reports of positive change in practice among teachers. While removing ineffective teachers can play an important role in evaluation policy (censure, demotion, reduced pay, and termination), the accumulating evidence—including that presented in this study—suggests that they might be harmful to the intrinsic motivation that a majority of teachers across the world have in abundance (Brookhart & Freeman, 1992; Chong & Low, 2009; Donaldson, 2021; Ford et al., 2017; König & Rothland, 2012; Krecic & Grmek, 2005; Kyriacou & Coulthard, 2000; Pallas, 2021; Sinclair, 2008; Watt & Richardson, 2008).
Implications for Policy and Practice
Our findings suggest, at the very least, that different teacher evaluation and supervision activities might be led by different individuals within the school; thus policymakers and practitioners may want to consider policies and practices that honor that, while administrators bring important school-wide perspectives and knowledge to understanding global aspects of instructional practice and student survey data, teachers and mentors may be able to supplement this expertise in content-based ways that support teacher growth and development, and ultimately student learning.
While not central to our guiding research questions on the relationship of evaluators to important changes in teacher motivation and practice, one notable and consistent finding from our study is the negative relationship between a principal’s use of extrinsic motivational tools after evaluation, such as rewards and consequences, and teacher outcomes (e.g., practice/motivation). Scholars have warned about the negative ramifications of high-stakes evaluation (Ford et al., 2018; Guenther, 2021; Holloway et al., 2017; Lavigne, 2014; W. C. Smith & Holloway, 2020) where teacher evaluation outcomes are tied to extrinsic rewards such as promotion, salary increase, etc., particularly on factors such as teacher affective states, motivation, retention, and school culture and climate. These findings suggest that policies and practices in favor of teacher evaluation systems that prioritize teacher growth and development may be more effective in fostering teachers’ motivation and practice than those models that prioritize holding teachers accountable using high stakes (Ford & Hewitt, 2020; Guenther, 2021; Lavigne & Good, 2020).
Limitations
While we provide theory-based explanations for the results we found, one limitation to the overall study is our inability to provide direct evidence of the mechanism—in other words, how administrator versus teacher-led evaluation results in different outcomes for teachers. For example, we do not know whether or not rich content-based discussions happen when mentors and teachers are leading content-based assessments of knowledge and analysis of test score data. Another limitation is that TALIS2013 data were collected from teachers at the lower secondary level. Perhaps content-based expertise plays a larger role in lower and upper secondary as teachers are more likely to be subject-area specialists. However, this may depend largely on the country or site. For example, as it pertains to the US, elementary teachers often teach in multiple subject areas, but in other countries, elementary-level teachers do have a content area specialty. Thus, the differential role of content may be salient in some places and insignificant in others.
Suggestions for Future Research
To better support states and districts in their use of best practices as it pertains to teacher supervision and evaluation, it would be useful for educational researchers to partner with districts willing to pilot the use of teacher-led evaluation models alongside principal-led evaluation models with the use of randomly assigned schools and teachers in a low-stakes evaluation environment. This would provide an important causal methodological extension to the analysis and research provided here which is primarily correlational.
Including direct measures of feedback from teacher-led and administrator-led components would help illuminate the how and the why. This is particularly important for understanding how teachers and mentors may or may not mitigate the challenges principals face especially as it pertains to content (Kraft & Gilmour, 2016). Outcome measures such as teachers’ perceptions of the usefulness of feedback from teacher- and administrator-led components would be pertinent (Cherasaro et al., 2016) as well as direct measures of teaching and learning especially since current teacher evaluation models have had null effects (Kim & Sun, 2021; Lavigne & Good, 2019). Finally, as supervision continues to travel under the guise of evaluation and few state models distinguish between supervision and evaluation, it would be useful to determine if the inclusion of teacher- and mentor-led evaluations help untangle these two goals and processes (Glanz & Hazi, 2019; Hazi, 1994; Mette et al., 2020).
While we recognize that teacher evaluation often serves dual roles—a summative judgment of a teacher’s effectiveness and formative feedback to improve teaching and learning (Ford & Hewitt, 2020)—the greatest value offered by teacher evaluation to teachers is often the feedback that occurs throughout the year as teachers have immediate opportunities to integrate such feedback into their daily practice. Recognizing this, we believe it is important to better understand and continue to support research that furthers our understanding of the who and how evaluation for teacher practice, motivation, growth, and development.
Footnotes
Appendix A
Item-Level Information on Rasch Measure Positive Change in Teacher Motivation and Practice.
| Concerning the feedback you have received at this school, to what extent has it directly led to a positive change in any of the following? (TT2G30) | |||
|---|---|---|---|
| Item (4-point scale, no positive change to a large change) | D | Infit | Outfit |
| Your classroom management practices (TT2G30H) | 0.50 | 1.03 | 1.04 |
| Your use of student assessments to improve student learning (TT2G30L) | 0.27 | 1.06 | 1.07 |
| Your job satisfaction (TT2G30M) | 0.02 | 0.92 | 0.90 |
| Your teaching practices (TT2G30J) | −.01 | 0.90 | 0.90 |
| Your motivation (TT2G30N) | −.10 | 0.96 | 0.93 |
| Your confidence as a teacher (TT2G30F) | −.68 | 1.12 | 1.11 |
Note. D = item difficulty. Person separation reliability = 0.87; item reliability = 1.00. TALIS teacher questionnaire item numbers in parentheses.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
