Abstract
In recent decades, efforts have been undertaken to improve the quality of principal leadership, including the establishment of the Interstate School Leaders Licensure Consortium standards (Council of Chief State School Officers, 1996, 2008); growth in principal professional development, coaching, and mentoring; and improvement of preparation programs (Kottkamp, 2011). More recently, there has also been a renewed interest in assessment and evaluation as another lever to influence the quality of school leadership.
Assessing principal effectiveness has been an important element of school improvement for more than two decades. However, the knowledge base regarding quality, use, and influence of principal leadership assessment is limited. As early as 1990, in a comprehensive review of the literature on principal evaluation, Ginsberg and Berry (1990) found a wide array of practices reported with little systematic research to support one approach versus another. In a more recent review of leadership assessment in education, Portin and colleagues pointed out that the broad trend of increasing emphasis on learning and school improvement are manifested in five shifts regarding what and how leaders are assessed: assessing behaviors instead of traits, relying on professional standards, focusing on learning results, emphasizing leadership development, and considering organizational context (Portin, Feldman, & Knapp, 2006). The evidence of such shifts, however, is yet to be substantiated by further empirical work on “the evolving nature and uses of leadership assessment approaches” (Portman et al., 2006, p. 26). Although there have been developments of assessment for initial principal certification (e.g., Education Testing Service’s assessments for principals), a recent review of 65 district and state principal evaluation systems found that the instruments consistently lacked alignment with conceptual frameworks for effective leadership and sound methods of establishing validity and reliability (Goldring, Cravens, et al., 2009).
Principal leadership assessment can be an integral part of a standards-based accountability system and school improvement. The need for valid and reliable principal evaluation systems comes at a time when, in preparation for reauthorization of the Elementary and Secondary Education Act, the U.S. Department of Education issued A Blueprint for Reform calling for “states and districts to develop and implement systems of teacher and principal evaluation and support, and to identify effective and highly effective teachers and principals on the bases of student growth and other factors” (U.S. Department of Education, 2010, p. 4). The Blueprint indicates that to measure, develop, and improve the effectiveness of leaders and preparation programs, schools and districts will be required to take two important steps: first, to develop working definitions of effective principal and highly effective principal and, second, to establish district-level evaluation systems that (a) align with the effectiveness definitions, (b) are developed with the knowledge and participation of multiple stakeholders, (c) meaningfully differentiate performance levels, and (d) provide feedback to improve practice and inform professional development (U.S. Department of Education, 2010).
For evaluation tools to meet these criteria and to meaningfully differentiate among performance levels of principals, a rational, sound, and coherent standard-setting process must be implemented during the psychometric development of an assessment instrument. In the past, school principals were often labeled as aspiring, novice, or experienced on the basis of seniority. With a criterion-referenced instrument that measures leadership behaviors, however, a standard-setting process can be used to assess principals on the basis of their proficiency, which sets apart those that are highly effective or effective from the less effective ones.
Psychometric literature supports the notion that valid standard setting adds to the credibility of an instrument that assesses competencies (Berk, 1996; Cisek, 2006). Standard setting typically involves establishing proficiency standards and identifying cut scores on a continuum of measured performance (Downing & Haladyna; 2006; Kane, 1994). The proficiency standards define what behaviors and competencies a person needs to exhibit. Well-specified proficiency standards describe a set of behaviors or competencies that the assessment measures (Haertel, 2002). The cut scores operationalize the performance standards, separating examinees into groups depending on how well they perform on each standard on the basis of a test or rating scale or other type of measurement. In other words, in assessment development, levels of proficiency need to be translated into specific cut points on a scale.
Prior research in psychometrics and personnel psychology indicate that standard setting is arguably one of the most visible but controversial elements of assessment development and reporting (Cizek, 2006; McGinty, 2005). Regardless of the types of methods used, the validity of the standard-setting methodology rests on the degree to which participants can select cut scores that appropriately define proficiency levels (Buckendahl, 2005). Furthermore, the standard-setting deliberation does not occur in a vacuum. That is, the thought processes and experiences of participants during the judgmental tasks of the standard-setting process may reflect the influence of external factors. These factors may include individual personal experiences, the professional roles of the participants, and their awareness of possible consequence of the evaluation results (Ferdous & Plake, 2005).
Although the psychometric dimensions of standard-setting methods have been widely studied, until recently, only a limited number of inquiries have been conducted that examine the participants’ experiences in standard setting as it unfolds (Buckendahl, 2005). Moreover, these studies are limited to standard setting for student achievement assessments, not evaluations of school principals.
This article presents a new study that sheds light on the participants’ experiences with establishing cut scores for principal leadership proficiency. We enlisted a standard-setting process from a large school leadership assessment project as the backdrop for this study. The Vanderbilt Assessment of Leadership in Education (the VAL-ED) is a multirater rating scale that assesses the effectiveness of principals’ behaviors known to influence teachers’ performance and in turn students’ learning (see Goldring, Porter, Murphy, Elliott, & Cravens, 2009). The VAL-ED measures critical leadership behaviors for the purposes of diagnostic analysis, progress monitoring, and summative evaluation. The development of the VAL-ED included a complete set of instrument development steps, among them the process to set cut scores for leadership proficiency levels (Porter et al., 2010). The VAL-ED is one of the few principal leadership assessment instruments that included standard setting in its development.
Specifically, in this article, we present the results of a study in which we observed and documented a standards-setting process for the VAL-ED leadership assessment as it took place and interviewed panel participants during and after the process; this study examines both the observed dynamics and self-reported reflections by panel participants regarding the factors that influenced their decision making of their cut scores for determining proficiency levels on the VAL-ED rating scale.
We first provide a short description of standard-setting processes in general and then identify the main issues that emerge from previous studies. We follow with a detailed portrayal of the VAL-ED instrument and the specific steps involved in setting cut scores, which serve as the primary background for this study. We then explain the analytic strategies for examining the nature of decision making in this very process. Next, we present the findings from our analyses of observation logs and cognitive interviews with panel participants. Last, we discuss the implications of identifying challenges and solutions to setting standards for leadership effectiveness.
Decision Making in Setting Proficiency Standards
Standard setting is the process of establishing cut scores on assessments or rating scales. Cizek (1993) defines standard setting as “the proper following of a prescribed, rational system of rules or procedures resulting in the assignment of a number to differentiate between two or more states or degrees of performance” (p. 100). In addition to Cizek’s procedural explanation of standard setting, Kane (1994) highlights the conceptual nature of the endeavor. He pointed out that the cut score is used to operationalize performance standards, which are indicators of desired levels of competence or proficiency.
Most standard-setting methods start with a test and then involve a panel that selects cut scores that could be earned by someone just at the borderline between competent and incompetent (Green, 2000; Karantonis & Sireci, 2006). Standard-setting panel members are usually selected on the basis of their content area expertise and familiarity with the target group of examinees (Cizek, 2006). However, some suggest that even expert participants may not fully understand their judgments when they engage in the standard setting because of the complexity of the underlying tasks (Shepard, Glaser, Linn, & Bohrnstedt, 1993).
Mitzel, Lewis, Patz, and Green (2001) developed the “bookmark” method in an effort to reduce the cognitive complexity of the judgment tasks. In bookmark standard-setting workshops, panelists follow a set of procedures and make judgments in two or three rounds about cut scores to define performance levels. To identify the points that separate performance levels, panelists work their way through an ordered-item book (OIB), one assessment item per page, where the items are ordered according to difficulty. For example, in bookmark standard-setting workshops for grade-level achievement tests, panelists place their bookmarks on an OIB page that sets apart two competency levels for the particular test. At the end of each round, workshop facilitators calculate the median page number across all panelists. The median is most commonly used to identify the center of the group decisions and to minimize outlying viewpoints (Karantonis & Sireci, 2006). The median (and individual) page numbers are used as feedback in subsequent rounds for forming the convergence of a cut score on the underlying test score scale. Recent empirical studies that compared the results of bookmark methodology and other methods, such as the Angoff method (Angoff, 1971), using independent panels yielded similar cut score recommendations (Buckendahl, Smith, Impara, & Plake, 2002). However, their findings also indicated that the decision-making tasks remained challenging for the panelists across various methods (Hein & Skaggs, 2010).
Although proficiency cut scores are used in various settings for personnel decisions, literature in industrial and organizational psychology and human resource management regarding guidelines for setting cut scores reveals that few guidelines are available on procedures and validity measures (Cascio, Alexander, & Barrett, 1988; Schmidt & Hunter, 1998). In fact, our review of personnel psychology yielded no other studies that examined standard-setting processes empirically for senior personnel (managers or administrators) nor any that connected management or leadership theories with specific cut score methods, such as the bookmark. Although relevant studies largely focused on student assessments and personnel ability testing, various veins of literature point out that concerns about the validity and utility of setting cut scores can be viewed from two angles in relation to the process: the implementation of key procedures (internal) and the consideration of influential factors (external).
Key Procedural Elements of Setting Cut Scores
To select cut scores, panel participants need to conceptualize a given level of mastery on the assessment and identify a particular examinee (such as a principal) who typifies that level. Skorupski and Hambleton (2005) found three procedural steps influential in panelists’ decision making: (a) the specificity of proficiency definition, (b) the training activity that operationalizes the definitions for difficulty levels, and (c) the emphasis on what the examinee would do (manifested behaviors). If these steps are in place, according to the investigators, they are more likely to drive the cut score decisions than are the panelists’ preconceived notions of content mastery (Skorupski & Hambleton, 2005).
Performance level descriptors (PLDs) provide the proficiency definition and anchor the multiple stages of the standard-setting process (Buckendahl, 2005). Examinees (principals in this study) may be labeled according to PLDs such as barely proficient, basic, or highly proficient. Panelists then are asked to consider the likely performance, or effectiveness, of these target examinees and to then estimate, item by item, the performance of such examinees on the assessment. For example, in a study on teachers’ perception of the target examinee in school district standard-setting processes, Giraud and Impara (2005) evaluated the use of the definition of barely proficient and barely master by teacher panelists to set proficiency cut scores for reading and mathematics tests. The study revealed that when the description of the target examinee was less definitive, teachers’ descriptions of the target examinee were more varied than when the description provided by the workshop facilitators was more definitive in terms of expected behaviors. This suggests that “a priori definitions of performance that describe the target examinee in certain ways and more or less exactly can substantially influence judges’ operational notion of target competence” (Giraud et al., 2005, p. 230). In a separate study of the thought process during standard setting, Skorupski and Hambleton (2005) found that panelists with relatively little understanding of the procedures or the PLDs tended to be the most unpredictable with respect to changes in their item ratings from the first to the second round.
External Factors That Influence Cut Score Decisions
Setting proficiency standards through establishing a cut score is a criterion-referenced task gauging levels of competence. It involves a value judgment that is often linked to certain consequences for individuals or organizations. Although value judgments in various forms are embodied in other stages of assessment development as well, standards-based reports send direct messages about adequacy of performance and therefore require attention on the normative nature of the judgments made in the process (Haertel, 2002).
Berk (1996) suggested that consideration of the political, economic, social, or educational outcomes of decisions about examinees is important in setting cut scores. Whether and the extent to which the awareness of consequences influence decision making by the participants in the standard-setting process is a question particularly salient in today’s high-accountability environment. For student achievement tests, for example, the consequences of standard setting may fall disproportionately on certain groups of students, on certain teachers, or on certain schools. Consequently ignoring the roles of teachers, principals, and other stakeholders who participate in standard setting may underestimate the impact of their perspectives on their interpretations and uses of standards or cut scores.
Ann Fitzpatrick (1989) reviewed research from social psychology on the effects of group discussion and of exposure to the opinion positions of other group members. The review indicates that group discussion may polarize preexisting opinions or cause “social comparison,” and subjective judgments are more susceptible to this influence. Because it is most desirable that standard setting be based on diverse and relevant information, Fitzpatrick indicates that procedures should be designed to both minimize the effects of social comparison and maximize the effects of pertinent informational influence on the decisions.
Standard-setting experts emphasize the importance of “due process” in conducting a rational and coherent deliberation where the elicitation and consideration of various viewpoints are provided throughout the process. Giraud (1999) compiled observational data of a standard-setting process and interviewed teachers following completion of the standard-setting tasks. He found that teachers put aside personal views about student performance and adopted the needs and expectations of the school district. However, he also found that the panelists felt pressured to recommend passing scores that would be acceptable to confirm expectations. Haertel (2002) listed various scenarios when stakeholders’ opportunities to express their views might “fall short of authentic participation” (p. 19), for example, when discussions are superficial or rushed.
In the end, the appropriate use of the standards-based reports with constrained normative interpretation hinge on (a) users’ understanding of parameters of the performance domain covered by the assessment and (b) users’ willingness and readiness in associating the standards-based reports with performance (Haertel, 2002; McGinty, 2005). Recent literature on standard setting indicates that panelists’ decisions may also be influenced by their concerns about the lack of necessary conditions to meaningfully implement standards. For example, McGinty (2005) found that in one of the standard-setting studies, panelists (teachers) were skeptical about how their input would actually be used and grew increasingly concerned about whether the standards they recommended would actually be adopted and used by the state. The discussions on the viewpoints of different participants and the influence of group dynamics suggest the need for continued examination of the motivations and beliefs of panelists as they participate in cut score decisions. Yet there is limited understanding of how panelists actually reach their decisions in a standard-setting process.
The aforementioned literature review underscores the challenges of determining criterion-referenced cut scores by expert panels, the importance of effective implementation of key procedural elements of standard setting, and the necessity of taking into consideration external factors that may influence decision making. The review informs the development of our analytical strategies. We aim to “open the black box” of decision making through a qualitative study and to examine both the internal procedural process and the external factors. We address the following research questions:
How did the panelists reach their decisions during the process of establishing cut scores and setting standards of proficiency for a school leadership assessment?
Do external factors, such as the role of the panelist (teacher, principal, or supervisor), consideration of contexts, and the perceived consequence of a decision, influence standard-setting decision making?
In what conditions do panelists believe it would be acceptable to use cut scores for leadership proficiency standards?
We test the hypotheses that setting proficiency cut scores for leadership assessment is a challenging but achievable process when panel participants have thorough understanding and adhere to procedures and when factors influencing decisions are well addressed throughout the process.
The VAL-ED Standard-Setting Procedures
The development of the VAL-ED and its standard-setting procedures, although not the direct purpose of our study, serves as the setting in which we examine how participants determine cut scores for principal leadership proficiency. The VAL-ED is a paper-and-pencil or online assessment of principal leadership behaviors (see Porter et al., 2010, for complete details). 1 The assessment is based on a conceptual framework that is consistent with the Interstate School Leaders Licensure Consortium standards (Porter, Goldring, Murphy, Elliott, & Cravens, 2006), indicating that learning-centered leadership is at the intersection of what principals or leadership teams must accomplish to improve academic and social learning for all students (the core components) and how they create those core components (the key processes). A series of research studies were conducted to establish the VAL-ED as an instrument that (a) works in a variety of settings and circumstances, (b) is unbiased, (c) is construct valid, (d) is reliable, (e) is feasible for widespread use (both online and paper-and-pencil versions), (f) provides accurate and useful reporting of results, (g) yields a diagnostic profile for formative purposes, and (h) can be used to measure progress over time in the development of leadership. 2 The assessment consists of 72 items, 2 for each cell of the 6 × 6 conceptual framework (see Appendix A). The 360-degree assessment surveys the principal, the principal’s supervisors, and all of the teachers in the principal’s school. In each case, the respondent rates the principal on a 5-point scale from ineffective to outstandingly effective for each of 72 behaviors.
As a part of the instrument development and validation for the VAL-ED, an adaptation of the bookmark method was used to establish three cut scores for four levels of leadership proficiency: distinguished, proficient, basic, and below basic (Porter et al., 2008). 3 Twenty-two panel members were recruited to set proficiency cut scores using data from a national field trial in 2008 (Porter et al., 2010). The panelists were asked to make judgments in three rounds about cut scores to define leadership proficiency. Each round was preceded by discussion of feedback among panelists to ensure that they were ready to complete the task. The overall process took 2 days and included the following main steps.
Expert panel selection
The panel members were selected for their expertise in understanding the necessary and exhibited competencies for school leadership. Their task at hand was to make decisions on proficiency cut scores. As summarized by Table1, there were 10 principals, four teachers, four district-level supervisors of principals, two leadership researchers, and two state-level policy makers for professional development from 18 different states and the District of Columbia. Among them were 10 females, one Hispanic, seven African Americans, and 14 Whites. Principals and teachers came from various grade levels and schools of different sizes. One policy maker was at the state level and the other with a large school district. The two researchers are both well-known contributors to the research literature on school leadership. Four out of the 10 principals were recent recipients of Principal of the Year designation for their state. An additional 2 principals and all four principal supervisors received excellence designations by the American Association of School Administrators. All four teachers were recent recipients of teaching excellence awards in their districts or states.
PLDs
The PLDs were initially written by the VAL-ED research team and then went through multiple iterations of modifications based on the feedback of a diverse group of educators (Porter et al., 2008) from five different states. The final version of PLDs describes four levels of leadership proficiency (see Appendix B), whereby the exhibited competencies at each level are described and the differences among the levels are highlighted.
OIB
The bookmark method requires the panelists to work their way through a booklet of the assessment items, one item per page, where the items are ordered according to level of proficiency. For the VAL-ED, the data for the OIB came from a nationally representative field trial completed in the spring and summer of 2008, including 218 schools for which there were principal evaluation results from all three response groups (teachers, principals, and supervisors). Among them, 39% were elementary schools, 32% middle schools, and 28% high schools. Twenty-three percent of the schools were from the West, 30% from the South, 22% from the Midwest, and 25% from the Northeast. There were 39% urban schools, 39% suburban schools, and 22% rural schools (for more detailed description of the national field trial, see Porter et al., 2010).
To order the items by difficulty, an aggregate mean was calculated for each of the 72 items for the VAL-ED across the principal, the supervisor, and the mean for the teachers in the principal’s school. The principal, the supervisor, and the mean of the teachers were thus equally weighted in creating this aggregate mean. Each item’s mean could range from 1 to 5, representing the levels on the effectiveness rating scale, where 1 = ineffective, 2 = minimally effective, 3 = satisfactorily effective, 4 = highly effective, and 5 = outstandingly effective. The lowest mean represents the most difficult item and is placed on the last page and vice versa (Porter et al., 2008). 4
In each round, panelists were instructed to place their bookmark on the page where a just-barely-proficient principal would be rated “at least highly effective.” The OIB did not present item means because panelists were to focus on the nature of the behaviors. Next, panelists placed their bookmarks for where a just-barely-distinguished principal would be rated “at least highly effective.” Finally, bookmarks were placed where a just-barely-basic principal would be rated “at least highly effective” (see Appendix B).
Guiding questions
To place a bookmark on an item that separates two proficiency levels, panelists were guided by two questions that reminded them of the criterion-referenced tasks: (a) What behaviors must a principal exhibit to achieve a rating of at least highly effective (i.e., a 4) on this item? (b) What makes it more difficult to achieve a rating of at least highly effective on this item than on all previous items in this book?
Impact data
Impact data provided an empirical framework for the panelists to see how the cut scores may affect the performance grouping of actual examinees (Cizek, 2006)—how would the cut scores have categorized the 218 principals in the 2008 national trial? Impact data were given at the beginning of Round 2 on the basis of the room’s medians for each of the three cuts in Round 1. Introducing the impact data was intended to help the panel evaluate the reasonableness of the cut scores derived from the standard-setting process.
Group discussions
Group discussion was employed to make standard setting a convergence process (Giraud & Impara, 2005). The panelists were organized into five worktables, where the composition of panelists at each table included random assignment of two principals and a mix of teachers, supervisors, researchers, and policy makers. Each panelist’s bookmark at each round was recorded. Each worktable was informed of its median page number for each of the three cuts. They were also given the median cut for the full group for each cut. The panelists were encouraged to share ideas, insights, and rationales across tables as a whole group before and after their cut score decisions at each round. However, the process was not designed to be a consensus process. Independence of judgments is crucial to the success and defensibility of the process (Ferdous & Plake, 2005). Panelists were instructed to place their bookmarks—that is, make their judgments—independently of all other panelists.
Research Design and Method
Data and Analytic Strategies
The analytical strategies were designed to explore what panel participants did and how they felt about the process, the procedural elements (internal) and the influential factors (external) as related to cut score decision making. To answer the research questions, we relied on qualitative analysis of interviews and observations. We employed research methods that focus on interviewing participants to ask them about the thought processes as they perform cognitive tasks (e.g., Ferdous & Plake, 2005; Green, 2000; Lane, 1991; Skorupski & Hambleton, 2005). Although interviews are often employed to understand participants’ thought processes, such methods have limitations. The two major criticisms are that (a) verbal self-reports may not accurately describe actual cognitive processes (Nisbett & Wilson, 1977) and (b) specific thoughts may be triggered by the interviews, and thus responses may be reactive and not related to the “regular” cognitive processes that would have occurred (Cavanaugh & Perlmutter, 1982). For these reasons, methods that are unobtrusive are also desirable. Therefore, our data collection included observations of the group discussions, structured interview of panelists at the end of each round, and the actual bookmark decisions of where the cut scores should be by each panelist.
At the beginning of the standard-setting workshop, the panelists were invited to participate in the study (in addition to the actual task of setting standards) that would examine the working process itself. The study was voluntary. All 22 panelists agreed to participate in the study (observations, interviews, and the standard-setting process itself) and signed the informed consent agreements. Each participant was interviewed at least once during the process. Five participants were interviewed after each round as the anchor interviewees for each working table throughout the three rounds. To address the limitations of self-reported data, one researcher was assigned to each of the five working tables for concurrent observations and took notes during the group discussions after each round.
Semistructured interview guides (Appendix C) were developed to focus on the areas of inquiry. The observations and interview questions focused on how panelists use the PLDs, the OIB, and the two judgment questions during the decision process, with a particular focus on how clearly the panelists could portray what a target principal should be able to do at each proficiency level. To explore factors that may have influenced the panelists’ decisions, we asked questions and observed issues related to background characteristics and group dynamics. Furthermore, we asked about the influence of impact data and whether the panelists considered the consequences of their decisions.
As suggested by Moustakas (1994) and Ferdous and Plake (2005), interview transcripts and observation logs were read through entirely to get an overview of panelists’ responses, then relevant and salient themes were developed and recorded. Statements that were meaningful in the study context were placed into a thematic framework that was structured on the basis of the research questions (Appendix D). We used a thematic coding framework referencing the “roadmap” of decision-making steps for content analysis suggested by Crano and Brewer (2002, p. 247) and coding procedures recommended by Strauss and Corbin (1998). Two coders worked independently to code data using the thematic framework. When there were differences between the two coders, a third coder was asked to code, independent of the first coders, using the original content of the same thematic framework. The second round of coding results was then compared with the first round to form the final coding decision and reach reconciliation. We also examined the actual bookmark placement of the panelists to (a) seek evidence that could verify or contradict the results gleaned from observations and interviews and (b) see whether any group patterns emerged from the results of bookmark placement that could provide further insight into the findings.
Results
To report our findings, we start first by describing cut score decisions made by the 22 panelists after three rounds of deliberation. Table 2 provides a summary of the bookmark placements (the page number where the cut score is located) at each round, sorted by working table. Overall, the bookmark placement records showed that the 22 panelists converged noticeably around the medians of the three cut scores: barely basic, barely proficient, and barely distinguished. The convergence was demonstrated by the narrowing standard deviations of the three cut scores at each working table from Round 1 to Round 3, ranging from a reduction of 23.9% by Working Table 1 to 100% by Working Table 5 for proficient and a reduction of 100% for distinguished by Working Table 3. The overall reductions of standard deviation of cut score placements across the three rounds are 36.2 % for basic, 43.3% for proficient, and 75.8% for distinguished.
Panel Participants
Note. The panel members came from 18 states and the District of Columbia: Alabama, California, Connecticut, Florida, Kansas, Indiana, Louisiana, Missouri, Massachusetts, New Mexico, Ohio, Oregon, Pennsylvania (3), Rhode Island, South Carolina, South Dakota, Texas, Tennessee (2), and Washington, D.C.
Median Bookmark Placement by Working Table
Note. The median was used to identify the center of the group decisions and to minimize outlying viewpoints (Karantonis & Sireci, 2006).
Using Key Procedural Elements to Determine Cut Scores
Our analyses of the observation logs and panel interviews focused on how the panelists used the PLDs, the two guiding questions for cut score determination, and the OIB to reach their understanding of a targeted principal at each performance level.
Using the PLDs
We found that the PLDs served as the starting point for the conceptual understanding of the performance levels for the panelists. The panelists reported that while they could depict the kind of principals that fit in each category on the basis of their own experience, articulating such a conception and associating it with observable behaviors was not easy. They also reported that identifying the differences between two adjacent categories with observable behaviors was even more challenging.
We found that training and practice on defining what constitutes just barely reaching each level appeared to be instrumental in helping the panelists meet this cognitive challenge. Interview results show that panelists relied on the PLDs to understand the demarcation between the levels. This understanding, however, did not come without extensive explanation and training. The panelists also had to be urged by the facilitator to identify behavioral differences between any two levels by giving specific examples. Panelists stated that “even though we had a definition there, we all had in our minds what we thought.” As one district supervisor put it after Round 1,
I’m having a lot of difficulty with this concept of “barely”—barely basic, barely proficient, barely distinguished. In fact I have to tell you that I don’t think I’ve ever used such terms before for purposes of assessment. Either you have or you haven’t or you’re on your way to improving. It was a little difficult to try to understand what was meant by barely. I think I’m still having difficulty with it, but it certainly did allow us to center our discussions on behaviors. That was where it was helpful.
We found that during the actual standard-setting process, when the panelists were mindful of the PLDs and used them as a reference, they relied less on their previous conceptions and definitions of principal effectiveness that might be more norm referenced.
Observation logs indicate that much discussion about the PLDs occurred only at the beginning of the standard-setting process during training and practices but not as much throughout the bookmarking rounds. It is possible that the panelists referenced the PLDs less as their confidence with the standard-setting process grew and their understanding of the PLDs was internalized and became more intuitive to them. Panelists found the PLDs more influential toward the early part of the process, when they were continually going back to the different kinds of evidence to help them shape their decisions. One principal called it “a long process of familiarization” when interviewed after Round 3:
I think that was part of getting my head around what the task was. We were given the parameters to take our time to get our heads around. We didn’t need to understand everything to begin. I think the performance level descriptors were important earlier in the process and much less so later because I had internalized them and they were more intuitive. I was being more intuitive in their use.
Guiding questions as reminders of criterion-referenced tasks
Several panelists reflected in the interviews that it was easy to revert to an “internal metric” of performance categories, rather than relying on clear understandings of the PLDs. The following description by one supervisor stating how she “internalized” the PLDs appears to fit the definitions of the four levels (after Round 2), whereby a closer look at how she interpreted proficient and distinguished shows more of a norm-referenced judgment rather than behavior-based competencies:
I didn’t use exactly the same language and in mine was, below basic is somebody whose performance would be at best just managing the school but not really moving it forward or devoting sufficient attention to student learning and things like that. Basic was yeah they’re doing fine, not that there is room for improvement, but certainly nothing so terrible that you’d want to get rid of them or anything. Then proficient, they be [sic] a strong performer, above average, you know promote school growth and improvement well and so forth. The distinguished would be almost like superstars and you wouldn’t have very many of those.
The two guiding questions, therefore, are needed as cognitive reminders. The first question is to imbed a cognitive reminder for competency-based deliberation each time the panelist reviews the content of an assessment item. It asks about what behaviors principals have to exhibit to make the cutoff. It reminds the panelist to think about observable behaviors rather than other factors, such as seniority or the performance of their peers, so that decisions stay criterion referenced. Panelists had to constantly remind themselves that it was the levels of difficulty of the items that they were comparing. This is well illustrated by a discussion at one table in Round 2 on the item that measures the extent to which the principal ensures that the school secures the teaching materials necessary for a rigorous curriculum:
I thought of this for the person at the beginning of a goal-setting conference.
Let me give a concrete example. I went to “rigorous” (in the item about curriculum). I figured a really distinguished person would be able do really rigorous curriculum, but I figured a newbie wouldn’t be able to do this because you’ve got so much else to do.
I have two principals who are new and yet are on the high end of this.
I think there are newbies who could be at the top of this; it’d be rare but it could happen. I’m looking at my job here as evaluating the difficulty of the task in each item. I’m not sure if this is right but that’s how I’m thinking about it.
This (the process) is not meant for a 1st-year versus more experienced principal comparison.
You should take this and create the context in which you use this.
I just don’t like the concept of the immaculate principal who just comes out of the womb and is an expert.
Well, I’ve got two of them.
We see that Panelists 1 and 2 were taking experience measured by seniority into consideration, a more norm-referenced value judgment. Panelists 3 and 4, however, tried to stay focused on the behavior. Keeping in mind the question of what behaviors must be exhibited proved to be helpful. When this working table moved to higher page numbers (more difficult items), they were better able to focus on behavioral competencies. For example, on the item regarding monitoring the accuracy and appropriateness of data used for student accountability, the same panelists used very specific examples, such as “not just meet minimal prescribed formal evaluation each year but go beyond,” “walk-throughs,” “formal observations with follow-ups,” “informal feedback,” and “leave note on desk saying ‘Great job,’ ‘See me,’ etc.”
The second guiding question asks the panelist to identify what makes each item more difficult to achieve than all previous items in the OIB. Here the panelist is reminded that the bookmark divides two item sets, not just two items. Answering this question was almost unanimously claimed by the panelists as the harder task in the interviews. The observation logs also indicate that it appeared to be a conceptual leap that some panelists found difficult to make. Most of the panelists described the cut score decision as an iterative and arduous process of comparison. For example, one principal said after Round 3,
In particular in Round 1 when I started looking at specific items and trying to figure out where the cuts were, I actually had made notations. When we went through the discussion process and answered the two questions for each item, when I started feeling like boy this would be harder for a principal to do, this is more difficult, and when I started seeing a threshold to the next category, I actually made a little notation on the page numbers. So in Round 1 when I went to say well where are the cuts, I looked at those page numbers and then started weighing them against items that came before, items that came in after, is that the right place?
Using the OIB conceptually
Observation logs and interviews show that more than half of the panelists initially felt that the OIB appeared to be out of order at various places, a possible reason for the difficulty in seeing a distinct separation point for the cut score. Panelists reflected in the interviews that it became problematic when they perceived that there were conceptually hard-to-achieve indicators toward the beginning of the book.
Despite the disagreement with the page sequence of the OIB, however, most of the panelists were able to separate two sets of items on the basis of the behavioral demarcation. One supervisor’s description (at Round 3) of finding the intangible “hierarchy” of the items after discussions and using the two guiding questions is representative of the group reflections:
Hardly anybody was in agreement as to the hierarchy of the way the things were put together. We all wanted to pull out the staples and move things around because we didn’t think it necessarily built a hierarchy. It made a little more sense after we had a chance to answer the two questions on each page, that there was some good hierarchy going on there in terms of this skill is a lot less difficult than this skill and this skill. I think that after that discussion we had a little more clear idea of what each of those things meant. Now the tough part was to weigh them against the ideas of behaviors.
Panelists also came to a better understanding that the OIB was based on empirical data from the national field trial of 218 school principals. With small differences among time averages, the panelists were instructed that it was okay to look past an item that seemed out of order. The real task was to be able to identify the separation point that could differentiate two sets of behaviors.
External Factors Influencing Cut Score Determinations for Proficiency Levels
We examined whether and the extent to which factors outside of the standard-setting process influenced the cut score decisions: (a) the role and the perspective of the panelist, (b) the panelist’s consideration of school contexts, and (c) the introduction of the impact data. The results include findings from the observation logs, interview transcripts, and the bookmark results from the three rounds.
Diversity of panel composition
The data suggested three main findings regarding the influence of the roles and perspectives of the panelists: First, the panelists consistently used their own experiences to guide the decisions of where to put the bookmarks. Second, the five groups of panelists (principals, teachers, supervisors, policy makers, and researchers) started their bookmark ratings differently, but the differences narrowed from Round 1 to Round 2 and stayed largely the same in Round 3. Third, in conjunction with the use of well-guided discussions, the diversity of perspectives was an enhancing factor to reach convergence. The standard-setting training started off with this question from one of the principals:
If we’re working together at the table, wouldn’t it matter if the principal was elementary, middle, or high school? It would be a different set of skills for an evaluation of an elementary school of 400 students as compared to a high school who would be working with department heads. I will be thinking about behaviors differently than other people at my table, right?
In the interviews, panelists were asked how they interpreted the PLDs and assessed the difficulty levels of the items. Panelists consistently referenced their own experiences. The interview results after each round are consistent with the observation logs, whereby panelists, regardless of their roles, cited personal encounters to substantiate their decisions regarding cut scores.
We found not only that the panelists drew from their own experiences but that their awareness of background differences led them to make certain assumptions about role-based decisions by others during the process. In fact, during the first-round interviews, teachers predicted that they would have the highest expectations of the principals and therefore would set the highest bookmarks for the cut scores. This assumption was echoed by two supervisors who also predicted that the teachers would be “tough” on the principals. One teacher shared this perspective and said,
Well the theory that I had was that my bookmarks would be tighter than his because he’s a principal and he would give himself more leeway. I as a teacher want my principal to be effective and I think I’m going to be tighter in my bookmarks than him. It would be interesting to see if the other teachers do the same.
Interestingly, such assumptions were not supported by the actual bookmark placements. Table 3 presents the cut scores sorted by role. The data show that in the first round, the five role groups (teachers, principals, supervisors, policy makers, and researchers) started at different group median points. Teachers actually set a much lower median benchmark at page 12. For the proficient cut score, supervisors and the principals came out very close with expectations higher than the other groups at pages 41 and 40, respectively. The teacher median page was 29.5. For the distinguished cut score, supervisors set the highest bar at page 61, whereas the principals were in the middle (page 56), and the teachers had the lowest median benchmark at page 44. Adding the two researchers and the two policy makers into the mix, we see that the researchers stood out as setting the lowest threshold for basic (page 5.5) but a high bar for distinguished (page 60.5), and the policy makers had very similar scores to those of the teachers at all three levels. The differences among groups on cut scores may be attributable to small numbers but reflect results from the national field trial data of 218 schools, whereby supervisors and teachers consistently rated the principals higher than principal self-ratings (Porter et al., 2010).
Median Bookmark Placement by Role
Note. The median was used to identify the center of the group decisions and to minimize outlying viewpoints (Karantonis & Sireci, 2006).
From Table 3, we also see marked convergence occurring at Round 2 with distinguished and proficient cut scores and to a lesser extent with basic. Furthermore, role-based differences were largely reduced by Round 2 and stayed mostly unchanged in Round 3. In other words, convergence occurred rather quickly despite different perspectives. Observation logs and interviews revealed that structured and monitored group discussion played the essential role in facilitating convergence. Panelists pointed out that accepting the diversity in perspectives and understanding how they contributed to the interpretation of the performance standards were assets to accomplishing convergence in the final cut score decisions.
Observations and interviews also dealt with whether the panelists were unduly influenced by others’ opinions. We found that this was not the case. Panelists consistently recalled making the placement adjustment as an independent decision. In fact, several of them expressed a certain level of amazement with how close their decisions became:
My bookmark for the distinguished was far above [higher than] the two principals and the teacher in my group, which I was a little surprised at. I thought they might actually be higher. None of them have actually evaluated principals and I was just surprised that theirs were lower. I don’t know what I was expecting. It was interesting. Now as we talked we were pretty much right on target with each other. I came down a little bit and they came up a little bit.
The discussion, like the other principal said, around our table kind of forced us to look at that hard and seriously when we were putting those bookmarks in there. It was surprising when we went to Round 2 on proficient. We were separated from 42 to 30, I think, and we had a lot of discussion. When Round 2 came back we all put down 39 and we didn’t know that we had done that. I felt like we accomplished something!
Concerns about the accountability context
The school context in which principal leadership assessment takes place was another important element of the panelists’ cut score deliberation when weighing the levels of difficulty of certain assessment items. We found that principals and teachers were more attentive to the accountability environment than other role groups. The researchers, policy makers, and supervisors took the vantage point of their roles. One noted that
the indicators were demands for any principal regardless of where he or she is placed, whether it is a high poverty area or a low poverty area. They were generic enough and important enough to me to say this is something all principals need to pay attention to.
The principals and teachers struggled with this notion, but most arrived at the same conclusion, as reflected during the interviews. Overall, although school context and related factors were speculated, they did not influence the task of setting cut scores.
The principals frequently pointed out in their interviews that they were cognizant of the differences between schools that are rural or urban, well run or low performing, with a strong teacher community or a hostile staff environment, and so forth. Several principals very frankly shared their experiences of working in drastically different school settings that called for distinctive sets of leadership skills. One emphasized that school size would make a difference:
I’ve got 203 kids and you know other schools have got a thousand kids in their elementary school. I’m thinking holy moly what would I do with a thousand kids? You know I feel pretty inadequate when I start talking to people who have those kinds of numbers that they’re trying to deal with. When I started thinking about who’s distinguished and who’s not, it would be easy for me to be distinguished because I don’t have to deal with the tons of discipline and paperwork that someone else is doing.
More comments and comparisons were made on working with teachers and the organizational culture. Specifically, panelists felt that items on monitoring and accountability might actually have varying levels of difficulty depending on the context. For example, one panelist said,
You know there are different levels of monitoring and different levels of teacher quality. All of those things come into play. Those things sort of define the role of the principal because if you have a staff that is filled with quality instructional people then your monitoring process is going to be a little less stringent. It’s going to have to be as diligent, but not as stringent or not as planned out as it would have to be if you had a school where you had many teachers that were ineffective.
Observations and interviews indicate, however, that such awareness of varying school and district contexts served more as background references without direct influence on where the panelists located the bookmarks. In fact, in a manner very similar to what had occurred with the role-based discussions, panel discussions during the three rounds about whether certain items were context dependent prompted more in-depth probing into the item content. Panelists concurred that whether it was “the ruins of Rome needing to be put together again” or “this little well-run machine inherited from somebody,” certain sets of skills essential to school leadership would be required.
Concerns about consequence
For the VAL-ED standard-setting workshop, the impact data presented at Round 2 revealed to the panelists that their Round 1 median cut scores would have classified approximately 30% of the principals as below basic and another 30% as distinguished on the basis of the 2008 national field trial data. Table 2 shows that the introduction of the impact data at the beginning of Round 2 did not change the medians for the whole group for the three cut scores. The working-table medians changed more than the overall median, especially from Round 1 to Round 2. Using the cut score for basic as an example, although the overall median bookmark for the cut score moved only slightly from page number 14 to 15 in Round 2, among the five working tables, there were much bigger changes, ranging from lowering the cut-score bookmark (easier to pass) by 16 pages at Working Table 4 to increasing the cut-score bookmark (harder to reach) by 5 pages at Working Table 5.
Panelists reflected on their reactions after seeing the potential impact of preliminary standards set by the group. Most of the concerns were with the high percentage of principals classified as below basic because they considered this was a “sink-or-swim” threshold. The panelists were torn by their understanding of the competency-based criteria for basic and the norm-referenced information provided by the empirical data. A principal explained why she debated changing her bookmark placement for basic:
I went from 20-something down to 13 I think. I mean we talked it through and I see their point. They were worried they were gonna be let go—oh if you do that you’re gonna throw away an awful lot of good principals, you know that just aren’t right there where they need to be yet. They were worried about throwing people out.
Panelists were not as concerned with the high percentage of distinguished even though that seemed to be equally unrealistic, because the consequence of having 30% of below-basic principals was considered more severe than having 30% of “principals with blue ribbons around their necks,” as noted by a principal. A supervisor commented,
I thought that the number of people rated distinguished was too high. But I said my God it ought to be higher than it is. I don’t have any problem with lots of people being distinguished. If we could populate this country with distinguished principals we would have a much better public school system.
In the end, only a few of the panelists were swayed by the impact data. The fact that the room medians remained steady indicates that it was probable that the panelists understood the task at hand and acted on the instructions. The following comment by a supervisor, however, echoed the concerns of the panelists about the impact data:
I just thought it was an interesting conversation but it didn’t really influence where we placed our bookmarks. I think people are concerned that basic and below basic seemed to have significant percentages aligned to them and we’re not questioning what that means. My sense was if this is an accurate indicator, we’re in trouble.
This concern was reflected in the postsession evaluation filled out by 21 of the 22 panelists. Although virtually all respondents rated highly the helpfulness of the workshop leaders, the accessibility of training, and the clarity of instructions (Porter et al., 2008), 24% expressed some concern about thresholds for below basic and distinguished.
Necessary Conditions for Performance Standards Implementation
The panelists appreciated that a competency-based assessment had the potential to better define principal effectiveness in behavioral terms; to facilitate constructive dialogues among principals, teachers, and supervisors in the school system; and to guide professional development. A supervisor commented that the PLDs would help her operationalize state and district performance goals with her principals. Principals, in turn, welcomed the use of a tool that brings objectivity and also facilitates conversations among key stakeholders. However, the discussions and interview sessions also featured an undertone of concern about how the leadership performance standards would be used in the school districts. Although these concerns did not appear to have directly influenced the cut score decisions of the panelists, they are related to the ongoing efforts to enhance the external validity of assessment instruments, such as the VAL-ED, as viable tools of measuring and development school leadership.
The feedback of the panelists on enhancing the external validity of performance standards on school leadership can be categorized into several themes. First, there might be a need to consider weighing results along the career trajectory. The professional development policy maker repeatedly raised the question that the standard-setting process “sort of indicates that all items are equal and that people graded them and have higher grades for these, but are they the appropriate higher grades for the appropriate time in this principal’s leadership?” Second, panelists pointed out that there were items that might be outside of the control of the principals, namely, teacher recruitment, hiring, and firing, and that these items could skew the results. They recognized that these factors might be connected with the school contexts and largely depended on the district policies. Third, panelists emphasized that it could be “problematic” if an assessment was used for summative evaluation without incorporating other aspects of school performance beyond student academic learning and achievement.
Last, panelists listed various necessary administrative conditions for implementing leadership performance standards. Several panelists expressed concerns about the varying capacity of users in using criterion-reference standards to interpret assessment results:
There’s no statistical instrument that takes the place of good administration and good administrators understand that and can use an instrument like this for good administration. Poor administrators don’t necessarily understand that and could use this instrument poorly. I worry about that.
As I said earlier today just before we broke for lunch, this is a great instrument for checking the temperature of things, but it is really going to come down to using this along with your boss, along with what’s going on at your school, what the climate is really like.
Consequently, the panelists strongly recommended providing training and guidance for the future users of leadership assessments such as the VAL-ED. The state professional development policy maker suggested that being evaluated and knowing one’s proficiency level should be embedded in a carefully planned system, whereby there is “an entry portal in which we have the conversation about performance and what expectations are and then a time frame in which this is done and then a way of really talking through what just happened.”
Overall, our findings demonstrate the complexity of setting performance standards for school principals. In reflecting on how they reached decisions on the VAL-ED cut scores, panelists provided valuable insight into a process that requires cognitive attention to not only specific competencies but the context in which the behaviors take place.
Discussion
Understanding the processes in setting cut scores for proficiency levels—a standard-setting process—is imperative for developing the types of principal evaluations called for in A Blueprint for Reform (U.S. Department of Education, 2010). Set against the backdrop of the VAL-ED development and validation, this qualitative study addresses panel participants’ experiences in setting proficiency standards for school leadership. We examined the understanding and adherence of cut score procedures and the influence of external factors, including the heterogeneity of the panel members and concerns about potential consequences in panelists’ decision making. The study provides a roadmap to the inner workings of standard setting and illuminates the necessary elements to this complex process beyond the technical procedures involved. Its results indicate that setting cut scores for proficiency can be a feasible process that contributes to the credibility and utility of a comprehensive evaluation system regarding school leadership.
Our study suggests that standard setting is a cognitively demanding process that requires panelists to think and make multiple important judgments. Assessing principal leadership is a challenging task first at the conceptual level. To set performance standards for the VAL-ED, for example, what constituted an appropriate level of mastery was conceptually defined first by the assessment developers on the basis of research linking effective leadership behaviors with student learning (Goldring et al., 2009) and articulated by the PLDs given to panelists during the workshop (Porter et al., 2008). In defining a highly proficient principal or a below-basic principal, the challenge was to translate what the assessment measures into value judgments that reflected the description of manifested proficiency and can be defined by cut scores. Panelists from various backgrounds must be flexible enough in their thinking to revise their judgments on the basis of feedback from the facilitator and other panelists. Moreover, they must balance their judgments with their interpretations of the impact data.
Our analyses also substantiate findings from previous studies on standard setting that the PLDs, guiding questions, and the logic of instrument item ordering are essential to providing a disciplined structure to the process. These elements, if effectively employed, add clarity to the tasks, help the panelists focus on performance competencies, and lower the decision variance among panelists. Our observations of group dynamics and discussions reveal the nature and topics of deliberations and shed light on the formation of convergence on cut scores. Our interviews provide more in-depth understanding of how the panelists reached their decisions. We found that although some panelists struggled with the decision-making tasks, overall, the panelists were able to stay with the criterion-referenced judgments.
Setting performance standards must also include consideration of the influence of multiple external factors. We found that the participants recognized the diversity in roles, experiences, and perspectives of the panel but embraced the differences as an asset that added to the coherence and credibility of the process. Well-structured and facilitated group discussions were also credited by the panelists as key to providing breadth and depth of understanding to their cognitive quests. Moreover, the panel deliberation of the cut scores took school contexts and the impact of standards into consideration, but such awareness did not affect where cuts were set.
Linn (1978) suggested that cut scores must not be set so that unacceptable numbers of students were labeled as incompetent. Giraud and Impara (2005) stated that “the cut score was not simply a point at which individual students could be identified as masters or not, but was a symbol of high standards, an economic signpost, a measure of district accomplishment” (p. 310). For principal leadership standards, the stakes are equally high. We observed the panelists making strides to reach a level of balance between considering the meaning of a cut score on the basis of competencies but also taking account external factors, such as principal shortage and district reputation. Such consideration accentuates the political flavor of these types of performance standards and serves as yet another reminder that the cut scores are not simply markers for the mastery of skills and that standard setting is not purely psychometric (McGinty, 2005).
This study has its limitations. Selection of the participants to participate in the study as well as in the standard-setting process for the VAL-ED was based on their content expertise and their leadership positions in practice. Although the formation of the panel was diverse in terms of ethnicity, professional role, and school context, we were not able to investigate the extent to which many other important external factors, such as state or district accountability schemes, might have influenced their standard-setting decisions. Our findings could be biased because of the fact that information was drawn from one expert panel and therefore lacked generalizability to be applied to other standard-setting groups. We also cannot know the extent to which personal biases influence individual decisions. On the technical front, the VAL-ED is one of the few competency-based assessment instruments for school leadership newly introduced to the field. The standard-setting process requires summarizing many item-level judgments on what and how a principal ensures that the school performs effectively as a single point on a one-dimensional scale. Whether it is truly possible to capture valid opinions about differences between performance levels requires further exploration of the underlying constructs. With more field data, future research may need to test the conceptual distinctions between the empirical ordering of a set of items versus the ordering defined by panelists’ judgments of importance or relevance to a PLD.
Furthermore, the feedback from the participants and recent developments in understanding educator accountability systems underscores that assessment results must be interpreted not only in relation to the specific context but in connection with other performance measures that, collectively, may capture the complex nature of school leadership (Kowal & Hassel, 2010). Consistent with the national standards for personnel evaluation, an evaluation system for principals should include multiple measures of performance. No high-stakes decisions should be based on one source of data or one data point. Value added to student achievement can be weighted differently for principals depending on the context of their work. Districts should consider multiple dimensions to principal evaluation, whereby in addition to a multirater competency-based leadership assessment, student learning outcomes, organizational attainment objectives, and professional development goals may also be incorporated into a comprehensive evaluation system (see Porter, Murphy, Goldring, & Elliott, 2011). For example, for a new principal in a turnaround context, student achievement might be weighted less in a principal’s evaluation than for a veteran principal who has been at the same school for a long period of time, whereas leadership assessment results and reaching school improvement goals could be weighted more for the novice than for the veteran.
We recognize that there is limited empirical work to inform policy makers as to how various components of a principal evaluation system, including value-added student achievement measures and multirater systems, should be weighted; this is a ripe area for further research. A rigorous standard-setting process is one step in the ongoing development of assessments for educator effectiveness.
Footnotes
Appendix A
Appendix B
Appendix C
Appendix
Thematic Framework for Observation and Cognitive Interview
| Principals | Teachers | Supervisor | Researchers | Policy Maker | |
|---|---|---|---|---|---|
| Procedural elements of setting cut score | |||||
| How does the panelist use the PLDs throughout the process? | |||||
| How does the panelist use the two judgment questions? | |||||
| What does the panelist focus on when using the OIB? | |||||
| How clear can the panelist portrait what a proficient leader is able to do? | |||||
| How comfortable is the panelist in rating someone on the basis of her or his behavior? | |||||
| Factors influencing cut score decisions | |||||
| Does the panelist feel that his or her views of where to place the bookmark were similar or different from others in his or her role? Why (or why not)? | |||||
| Do the opinions and views of others influence the panelist’s decision about where to put the bookmark? | |||||
| Does the panelist use his or her role/background/experience to explain where the bookmark is put? | |||||
| What does the panelist feel about whether or how school contextual factors should be taken into consideration? | |||||
| Is the consequence of the performance report a concern to the panelist when setting standards? | |||||
| Does the impact data change the panelist’s decision on the cut score? Why or why not? | |||||
| Using proficiency cut scores to evaluate principals | |||||
| What does the panelist say about the actual use of the proficiency cut scores? Any recommendation on necessary conditions? Any concerns? | |||||
Declaration of Conflicting Interests
The authors declared the following potential conflict of interest with respect to the research, authorship, and/or publication of this article: The Vanderbilt Assessment of Leadership in Education (VAL-ED) instrument is authored by Drs. Porter, Murphy, Goldring, and Elliott and copyrighted by Vanderbilt University, all of whom receive a royalty from its sales by Discovery Education Assessment. The VAL-ED authors and their research partners have made every effort to be objective and data based in statements about the instrument and value the independent peer review process of their research. With any publication, readers in the end must judge the facts and related materials for themselves.
Funding
The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: The Wallace Foundation and the Institute of Education Sciences of the U.S. Department of Education to Vanderbilt University and University of Pennsylvania.
