Abstract
State policy makers are constantly looking for ways to improve teacher quality. An oft tried method is to increase the rigor of licensure exams. This study utilizes state administrative data from Arkansas to determine whether raising the cut-scores on licensure exams would improve the quality of the teacher workforce. In addition, the study explores the trade-offs of such a policy decision. It is concluded that raising the required passing score on the Praxis II would increase the quality of the teacher workforce, as measured by value-added student achievement. This change, however, would be accompanied with an important trade-off as it would reduce the number of minority teachers and potentially lead to negative outcomes in disadvantaged schools.
In every profession, the difference between an effective employee and an ineffective one can be substantial. This is especially true in education, where the quality of a teacher can have a significant impact on a child’s life. In fact, Hanushek and Rivkin (2006) found the difference between an effective teacher and an ineffective teacher can be as much as a year’s worth of learning. They note that students in the average teacher’s classroom learn a year’s worth of material, while students in an ineffective teacher’s classroom learn only half a year’s worth of material and students in a great teacher’s classroom learn a year and a half worth of material. As a result, being in an ineffective teacher’s classroom for 2 years could put a student a full year behind their average classmate and even further behind students who have had highly effective teachers.
President Obama noted the importance of teachers in his 2012 state of the union address when he cited a study that linked the academic records of 2.5 million students to adult outcomes (Chetty, Friedman, & Rockoff, 2011). The study found that students who had highly effective teachers, as measured by value-added student achievement, were more likely to go to college, earn higher salaries, and save more for retirement. In effect, teachers who produced the most learning gains in students also contributed to better later life outcomes for their students.
Throughout the United States, policy makers recognize the importance of teachers and have enacted many policies to improve the quality of education that students receive. Arkansas is one such state and is used as an example in this research. In many areas, Arkansas is behind. According to Education Week’s “Quality Counts” (2013) Report, Arkansas’s graduation rate is just 69.7% ranking the state 35th in the country. The state was also well below the national average in the percent of students passing their Advanced Placement exam. Nationwide, 21.9% of test takers passed, while just 15.5% of Arkansas students earned a passing score. Arkansas also falls below the national average on the National Assessment of Educational Progress (NAEP) exam. The NAEP, known as the nation’s report card, “is the largest nationally representative and continuing assessment of what America’s students know and can do in various subjects” (National Center for Education Statistics, 2013b). It is the best measure for comparing one state with another and for comparing a state with itself over time. Random samples of fourth and eighth graders take the test in math and reading every 2 years. In 2013, Arkansas eighth graders scored significantly below the national average in both subjects (National Center for Education Statistics, 2013a).
To improve student achievement, Arkansas policy makers should continue to seek policies that will help improve the quality of the teaching profession. The question is: How can we improve the quality of the teacher workforce? There are really three broad methods or strategies that might be attempted. States can try to screen out ineffective teachers on the front-end, help current teachers improve via professional development, or remove ineffective teachers from the classroom. There may be benefits and drawbacks to each of these strategies and each deserve careful study.
The goal of this article is to examine one front-end strategy that raises the barrier to entry-increasing licensure exam cut-scores. When cut-scores are increased, it becomes more difficult to enter the profession and it is believed that this will reduce the number of low-ability individuals from becoming teachers. This strategy has been used in the past by Arkansas and other states, such as Michigan and Missouri. By design, increasing the rigor of licensure exams leads to a decrease in the passage rate; thus, keeping more prospective teachers out of the profession. In Michigan, for example, passage rates dropped from 82% to less than a third when state officials changed the state’s exam (French, 2015). Missouri saw similar declines in passage rates when the state shifted from the Praxis exams by Educational Testing Services to new, more rigorous exams created by Pearson (Williams, 2015). The assumption is that the individuals who are weeded out by these more difficult exams are of lower quality. As a result, the quality of those entering the profession is supposed to increase.
Arkansas teachers are required to pass a series of licensure exams before they can become a certified teacher. They first must pass a Praxis I examination in mathematics, reading, and writing. These tests are essentially basic skills tests. Most colleges of education require prospective teachers to pass these exams before they are admitted to the education program. Teachers are also required to take a series of Praxis II exams. There are numerous Praxis II exams. All teachers are required to pass a Praxis II exam on pedagogical knowledge. They also must pass a Praxis II exam in their content area.
The research questions related to these exams are below.
The analytic strategy and data used are detailed in the “Method” section of this article. First, however, a review of research that has explored this topic in a similar fashion is provided. The findings are presented in the “Results” section and are followed with a discussion of the policy implications and some conclusions.
Literature Review
Occupational licensing is used in many trades. In fact, roughly 29% of the entire United States’s workforce is comprised of individuals in trades that require an occupational license (Kleiner, 2000). The intent of occupational licensing is to keep low-performing individuals out of the profession. In education, teacher certification is a form of occupational licensing, and licensure exams are one of the most common screens used in the certification of new teachers. As such, the purpose of certification is to make sure that teachers have the requisite skills to be an effective teacher. For certification to be effective in this task, the certification process must be correlated to effectiveness in the classroom. That is, the screen that keeps individuals out of the classroom should be an indicator of teacher quality. Therefore, how well an individual performs on a licensure exam should correlate to how effective they are in the classroom.
Some might contest this assertion. For instance, Gitomer, Brown, and Bonett (2011) suggested that the Praxis I exams were not designed to be a predictor of teacher quality. They wrote, “Although the tests are not intended to warrant expert or even competent teaching practice, they do set a minimum floor of knowledge that all licensed teachers must demonstrate” (p. 433). It may be the case that licensure exams were not created to predict teacher quality. Nevertheless, the logic model is clear. These tests which assess the “minimum floor of knowledge” are put in place because we assume teachers need these skills to be successful (Gitomer et al., 2011, p. 433). In other words, they are related to teacher quality at some level.
Many have examined the relationship between teacher licensure exams and teacher quality (Clotfelter, Ladd, & Vigdor, 2006, 2010; Goldhaber, 2007; Shuls & Trivitt, 2015a, 2015b, etc.). Each of these studies use large administrative datasets from state departments of education in North Carolina (Clotfelter et al., 2006, 2010; Goldhaber, 2007) and Arkansas (Shuls & Trivitt, 2015a, 2015b). These analyses measured teacher quality in terms of value-added student achievement as measured by standardized exams. Most of the studies have treated the licensure exam score as continuous, rather than a binary pass/fail. Generally speaking, these analyses have demonstrated that licensure exam scores are significantly correlated with teacher effectiveness, although the correlations are relatively small. Similar studies have examined the relationship between the SAT (Boyd, Lankford, Loeb, Rockoff, & Wychoff, 2008) and ACT (Ferguson & Ladd, 1996) on teacher effectiveness as measured by value-added student achievement. They too have found significant relationships between a teacher’s performance on exams and their performance in the classroom.
Fewer studies have examined the efficacy of the cut-scores used to determine whether a prospective teacher passes or fails the exam. Using administrative data from Texas, Hanushek, Kain, O’Brien, and Rivkin (2005) found no difference in terms of effectiveness between teachers who passed the state’s licensure exam and those who had not. Goldhaber (2007) took this line of analysis a bit further. Using North Carolina data, he examined the difference in terms of effectiveness between those who had passed the state’s licensure exam and those who had not. North Carolina had raised the cut-score needed to pass the exam. This allowed him to compare the performance of those individuals who passed the exam under the old regime but would have failed under the new with individuals who would have passed in both circumstances. In reading, he found no relationship between passing the exam and teacher effectiveness. However, in math, passing the exam was positively correlated to performance in the classroom.
Goldhaber (2007) noted that many states require teachers to take the same licensure exam, but often have different passing cut-scores. Connecticut and North Carolina required the same tests, but the passing cut-score was much higher in Connecticut. Therefore, Goldhaber also examined the impact of raising North Carolina’s passing score to the level of Connecticut. He found there was no difference in terms of effectiveness between North Carolina teachers who would have failed based on Connecticut’s cut-scores and those who passed in reading. In math, however, teachers who failed the exam were less effective. He concluded, “If states are seeking criteria to ensure a basic level of quality, then licensure tests appear to have some student achievement validity” (p. 788). In summary, teacher licensure exams may provide useful information when considered as a continuous number; but the scores may have less relevance when treated as a binary pass-fail.
Trade-Offs for Test-Based Certification Policies
Although licensure exams have some student achievement validity, Goldhaber noted they are not a perfect measure of quality. That is, the screen may keep out some potentially effective teachers. By raising their cut-score to the level of Connecticut, North Carolina would not dramatically improve the teacher workforce. It would, however, eliminate many potentially effective teachers from the labor force. This is just one potential drawback with making teacher licensure exams more difficult to pass or increasing barriers to entry.
The United States has long had a shortage of minority teachers. According to Villegas, Strom, and Lucas (2012), “the number of minority teachers nearly doubled” from 1987 to 2007 (p. 296). Although this appears to be an impressive growth in the number of teachers of color, minority teachers are actually more underrepresented in 2007 than they were in 1987. During this time frame, there has been rapid growth in the percentage of minority students. As a result, there is a large mismatch between the make-up of the nation’s student body and the teacher workforce.
For those interested in improving student achievement, especially for students of color, this is a problem. There is ample reason to believe students are more apt to connect with a teacher who shares their culture and can identify with his or her experiences (Gay, 2010; Villegas & Lucas, 2002). This view is supported by quantitative research which has examined the impact of having a teacher who is of the same race as the student.
Using data from the National Education Longitudinal Study of 1988, which contained survey data from more than 24,000 eighth-grade students, Dee (2005) found that teachers who did not belong to the same race or ethnic group as their students were more likely to view the student as disruptive and inattentive. Dee noted, “The results presented here indicate that the racial, ethnic, and gender dynamics between students and teachers have consistently large effects on teacher perceptions of student performance” (p. 163). Gershenson, Holt, and Papageorge (2016) found similar results with data on 10th-grade students from the Education Longitudinal Study of 2002.
Given these facts, we might expect students who have a teacher who is of the same race to fare better on standardized exams than teachers who do not have a teacher who matches their race. This is what Egalite, Kisida, and Winters (2015) found using a longitudinal dataset from Florida which contained observations on “over 2.9 million students linked to more than 92,000 teachers” (p. 46). Using a fixed-effects analysis, the researchers found student’s value-added scores on standardized exams improved more when they had a teacher whose race or ethnicity matched their own.
Although minority students may benefit from having minority teachers, licensure exams make this possibility more difficult. Minority teacher candidates perform worse, on average, on teacher licensure exams (Latham, Gitomer, & Ziomek, 1999; Gitomer et al., 2011; Goldhaber & Hansen, 2010). There has long been an achievement gap between White and Black students (Fryer & Levitt, 2004), and this gap may explain some of the differences in performance on teacher licensure exams. It is also important to recognize the possibility that stereotype threat (Nguyen & Ryan, 2008; Steele & Aronson, 1995) or test bias (Tellez, 2003) may impact the performance of prospective minority teaching candidates.
Petchauer (2012) suggested that students of color face additional obstacles when taking licensure exams, such as a lack of peers who have successfully completed the exams. Bennett, McWhorter, and Kuykendall (2006) similarly noted these challenges in their longitudinal qualitative study of 44 African American and Latino students who were preparing to be teachers. They suggested rigid licensure exam requirements continue to limit quality minority candidates from entering the teaching field.
For these reasons, there is some concern about using licensure exams as a screen to the teaching profession. Memory, Coleman, and Watkins (2003) suggested licensure exams may prevent capable minority teachers from entering the profession. They observed and evaluated the teaching performance of 161 prospective teachers and compared teaching performance with performance on the Pre-Professional Skills Test (PPST), also known as the Praxis I. The authors noted a one point increase in the passing cut-score on the PPST would yield a more effective teaching field by 0.01 standard deviation. It would also reduce the number of minority teachers entering the field.
Although Goldhaber (2007) and Memory et al. (2003) suggested that increasing the rigor of licensure exams would improve the overall quality of the teacher workforce, there is evidence to suggest these findings mask the impact on schools serving minority students. Goldhaber and Hansen (2010) noted that “these potential average gains would likely come at the cost of adverse effects for minority students” (p. 244). Their research utilized North Carolina’s longitudinal administrative data which spanned an 11-year period. Similar to Egalite et al. (2015), they found Black students significantly benefited from having teachers of the same race. As a result, reducing Black teachers who perform poor on exams, but are otherwise effective teachers, would put Black students at a disadvantage.
The logic model or rationale for teacher licensure exams is apparent, so is the rationale for increasing licensure exam scores. We want highly qualified teachers in the classroom and it makes sense that teachers should have some level of requisite knowledge. However, it seems there are clear trade-offs to this type of policy. Namely, test-based certification requirements lead to a reduction in the number of minority teachers entering the profession. This study adds to this discussion by examining the impact of raising licensure exam scores in Arkansas.
Method
In this analysis, teacher performance is measured in terms of a teacher’s impact on student achievement. This process is called “value-added modeling” or VAM. In an ideal VAM, students would be randomly assigned to teachers and the data would clearly link students to teachers. Unfortunately, Arkansas’s data are not suited for this type of analysis. Students are not randomly assigned to teachers, and the data do not provide student–teacher links. Therefore, this analysis utilizes a two-step strategy for estimating a teacher’s impact on value-added student achievement, whereby value-added is calculated in the first model and then included as the dependent variable in the second model. The impact of passing a licensure exam on student achievement is estimated in the second model. This strategy has been used by Shuls and Trivitt (2015a, 2015b).
VAM—First Stage
In the first stage, individual student test scores are regressed on 1- and 2-year lagged student test scores. The difference between the actual test score and the test score predicted from the model—the residual—is interpreted as the value added by the student’s teacher. This is a way to control for a student’s background, prior performance, and personal characteristics. The value-added measure is standardized with a mean of zero and can take on both positive and negative values, indicating whether a teacher has produced learning gains above or below the average teacher.
This analysis follows McGee and Costrell (2010) and Shuls and Trivitt (2015a, 2015b) by using a parsimonious model that uses multiple prior test scores to control for a student’s prior achievement and unobservable characteristics. Using multiple years of prior student achievement captures prior cumulative inputs of the individual and their school (Ballou, Sanders, & Wright, 2004).
Models of this form have been used by Aaronson, Barrow, and Sander (2007) and Hanushek et al. (2005), among others. Two years of prior test scores in both math and language arts are used as controls on the right side of the equation to capture student and school time invariant characteristics. The model used in this estimation is as follows:
Value-added estimates are generated at the school-grade level using a random effects estimator. This provides quality estimates at the school-grade level that are normally distributed with a mean of zero. This means that all teachers in the same grade at a school will have the same value-added scores imputed for them. This adds imprecision to the measure, but should not bias the outcomes.
Regression on Teacher Characteristics—Second Stage
In the second stage, the value added generated at the school-grade level in the previous equations becomes the dependent variable. Teacher-level data are regressed on the value-added scores. This strategy provides coefficient estimates that should be unbiased even though the standard errors are high. Upon generating the value-added estimates for the school-grade level, the following equation is estimated:
The dependent variable, u
j,k,t
, is the residual captured as value-added in Equations 1 and 2.
The Data and Descriptive Statistics
To examine the impact that raising the bar on teacher licensure exams will have on the teacher workforce, a number of datasets are used. The data were provided by the Arkansas Department of Education. The data indicate where a teacher teaches and the subject he or she teaches in a given year. The data also include descriptive characteristics, including race, gender, and whether the teacher has an advanced degree.
The data also indicate how well a teacher performed on various licensure examinations. Scores are provided for the three sections of the Praxis I, mathematics, reading, and writing. Scores are also provided on a variety of Praxis II examinations. The Praxis II tests are in content areas and professional knowledge. It is these scores that are used to explore the research questions.
Both the teacher data and student data are panel data. The student data use a unique 10-digit student identifier. These data indicate which grade and school a student is enrolled in during a given year. The data also indicate whether the student is enrolled to receive free or reduced-price lunch (FRL) under the National School Lunch Act. In addition, the data indicate whether the student has an Individualized Education Program (IEP) or is an English language learner.
Students in Arkansas take exams in Grades 3 through 8 in language arts and mathematics. These tests are known as the Benchmark Achievement Exams. The data include student records on these exams from 2005 to 2008. The Benchmark is a vertically aligned test; nevertheless, student test scores are standardized within each grade and year with a mean of zero and a standard deviation of one.
Student demographics in each year are relatively consistent. In each year, there are more than 200,000 student records. Table 1 displays relevant demographic statistics for the 2007-2008 school year. As can be seen in the table, the majority of students are eligible to receive FRL. The majority of Arkansas students happen to be White, while Black, and Hispanic students also make up a sizable portion of the student population.
Demographics for Arkansas Students in Grades 3 to 8, 2007-2008.
Note. FRL = free or reduced-price lunch; ELL = English language learner; IEP = Individualized Education Program.
As noted above, 2 years of math and language arts test scores are used as controls in the first model. Tables 2 and 3 display the value-added coefficients for these controls. In both subjects, all four prior test scores are significant predictors of future student performance. As expected, a student’s performance in ELA in the previous year is the strongest predictor of how well he or she will perform this year on the ELA exam. Similarly a student’s math score in the previous year is the best predictor of how well he or she will perform this year in math.
School-by-Grade Value-Added Coefficients for the Benchmark ELA Exam.
Note. ELA = English language arts.
School-by-Grade Value-Added Coefficients for the Benchmark Math Exam.
Note. ELA = English language arts.
Licensure Exam Data
As of September 1, 2010, Arkansas required prospective teachers to take the Praxis I pre-professional skills assessments in reading, writing, and math. The Praxis series is the most widely used licensure exam, but passing scores are set at the state level and these decisions are made somewhat subjectively. To pass the exams in Arkansas, teachers must score a 172, 173, and 171 on the respective exams. In the neighboring state of Louisiana, teachers must score a 176, 175, and 175. These tests are designed to provide more detailed information around the cut-point; they are not designed to give detailed information about test takers at all points in the distribution. That is, they do not do a good job distinguishing between individuals at the ends of the spectrum. In fact, individuals often max out or hit the ceiling on the test by scoring the maximum possible score. This results in a negatively skewed distribution, rather than a normal or bell-shaped distribution of scores (Figure 1)

Histogram of Arkansas teacher scores on the Praxis I mathematics exam.
On the Praxis I mathematics exam, there is a large spike in the number of teachers at 171, Arkansas’s passing score. This occurs primarily because individuals who do not score 171 typically will not become teachers. They are screened out. If Arkansas were to raise the score required to pass to the level of Louisiana, it may have the same effect of weeding out the individuals who fall below the dashed line in Figure 1. The analysis presented here makes use of Arkansas and Louisiana’s cut-scores to determine the effectiveness of the Praxis I as a licensure screen and the impact of raising the passing score on Arkansas’s exam.
It is not clear who the teachers are who fall below the passing line. These individuals could be teachers who were needed to fill a vacant position and were awarded a temporary license, teachers who passed the test during a time period when the cut-score was lower, and teachers who moved to Arkansas from a state that required lower scores, but had reciprocity with Arkansas. It is not clear how the difference between these teachers and the rest of the teaching population may differ on unobservable characteristics. Therefore, the results that compare these teachers should be interpreted with some caution.
The Praxis II exams are similarly distributed, but unlike the Praxis I, no state exactly matches the required exams of Arkansas. Therefore, there is not another state that serves as a comparison for the cut-score. More to the point, teachers of various subjects take a sundry of different Praxis II tests. Thus, the same strategy cannot be used for the Praxis I and Praxis II analysis. For the Praxis II, all exam scores are standardized and the impact of raising the cut-score by 0.25 standard deviations is estimated. This is slightly less than the equivalent of the increase to Louisiana’s passing mark on the Praxis I tests.
Results
In this section, the results of the analyses as they relate to the two research questions are presented. The first research question asks whether teachers who pass the licensure exams outperform teachers who did not pass licensure exams. The second research question asks what impact raising the required score to pass the licensure exams would have on the teacher workforce. Prior research has demonstrated that a relationship exists between a teacher’s performance on licensure exams and their performance in the classroom (Shuls & Trivitt, 2015a). Before examining the cut-score as an indicator of teacher quality, the relationship between performance on licensure exams and performance in the classroom is explored by sorting teachers into quintiles based on their performance on a licensure exam.
Descriptive statistics of the quintiles for teachers who took the Praxis I and Praxis II exams are presented in Table 4. In this analysis, the Praxis I score is a composite of the three Praxis I subtests in reading, writing, and mathematics. There are numerous Praxis II exams, and teachers are required to take a different number of tests depending on their certification subject. The Praxis II quintiles in Table 4 are based on performance on the Praxis II professional knowledge exam. This is the one type of Praxis II exam that all teachers in Arkansas are required to take.
Praxis I and Praxis II Quintile Descriptive Statistics.
Note. Standard deviations in parentheses.
As you can see, teachers in the lower quintiles were more likely to be minorities. In the first Praxis I quintile, just 75.2% of the teachers were White. This compares to 95.9% of the teachers in the fifth Praxis I quintile. Teachers in the lower quintiles also tended to work in schools that served higher percentages of minority students and economically disadvantaged students, as measured by the percentage of students receiving FRL. Thus, it seems that lower performing teachers are more likely to be employed by school districts that are serving low-income, minority students.
The relationship between these quintiles and value-added student achievement are examined using the regression methods outlined above, see Table 5. Quintile 1 serves as the base in this analysis. In both math and ELA, there is no significant difference between teachers in the second, third, or fourth Praxis I quintiles and those in quintile one. However, teachers in the fifth Praxis I quintile are significantly more effective, on average. The results are similar for the Praxis II; teachers in the fourth and fifth quintile are significantly more effective, on average, than teachers in the first quintile. In both subjects, the Praxis II exam tends to be a stronger predictor of performance in the classroom than the Praxis I.
Relationship Between Licensure Exam Performance and Teacher Effectiveness.
Note. Robust standard errors in brackets.
p < .1. **p < .05. ***p < .01.
There is a positive correlation between performance on licensure exams and performance in the classroom. The relevant questions are whether teachers who pass the licensure exam are significantly more effective than those who fail and what effect raising the passing score would have. These questions are examined next, but first the descriptive statistics of teachers who passed and failed the various exams at the specified cut points are displayed in Table 6. As noted above, there are teachers currently working in Arkansas who have failed a licensure exam. It could be that these individuals subsequently passed the exam or it could be that they have obtained some type of emergency certification. Whatever the cause, this allows for an examination of the difference between teachers who passed their Praxis I or Praxis II licensure exam and those who failed.
Praxis I and Praxis II Pass/Fail Descriptive Statistics.
Note. Standard deviations in parentheses.
p < .1. **p < .05. ***p < .01.
In addition, a similar methodology to Goldhaber (2007) is used to assess the impact of raising the licensure exam score. During the years included in my data, Louisiana and Arkansas required teachers to take the same Praxis I exams. However, Louisiana set the passing score several points higher than the Arkansas requirement. Louisiana’s scores (Praxis I) are used to estimate the impact of raising Arkansas’s cut-scores. The Praxis II is not as clean and tidy as the Praxis I where there is one test for all prospective teachers. On the Praxis II, states may require the same tests for some subjects but not others. This makes it difficult to conduct the exact same type of analysis as with the Praxis I. Therefore, the impact of raising the various licensure exam scores by 0.25 standard deviations is estimated. This produces essentially the same effect.
Table 6 presents descriptive statistics of teachers who passed and failed each exam at the two different cut points. When the cut-scores are increased, the number of teachers who would have failed climbs, from 291 on Arkansas’s Praxis I exam to 1,150. For each of the descriptive statistics, a t test is used to examine whether the teachers who passed the exam were significantly different than those who failed the exam. On both the Praxis I and Praxis II, teachers who failed the licensure exam were more likely to be minorities, more likely to be a non-traditional teacher, and had fewer years of experience. Teachers who failed licensure exams were also more likely to work in schools serving higher percentages disadvantaged and minority students.
The results of the analyses indicate Arkansas’s cut-scores on the Praxis I exams are not a significant indicator of teacher quality. The results of theses analyses are displayed in Tables 7 and 8. In both math and ELA, the difference between those who pass the exam and those who failed the exam is negative, but the difference is not statistically significant. This means that Praxis I exam is screening out some individuals who may be effective teachers. Even when the cut-score is raised to the level used in Louisiana, the difference between those who passed the Praxis I and those who did not is not statistically significant, but more teachers fail the exam.
Relationship Between Failing a Licensure Exam and Teacher Effectiveness in Mathematics.
Note. Robust standard errors in brackets.
p < .1. **p < .05. ***p < .01.
Relationship Between Failing a Licensure Exam and Teacher Effectiveness in English Language Arts.
Note. Robust standard errors in brackets.
p < .1. **p < .05. ***p < .01.
The Praxis II exam tends to be a more effective licensure screen than the Praxis I. Teachers who failed an Arkansas Praxis II exam tended to be significantly less effective in math. However, the difference between those who passed and those who failed was not significant in ELA. If Arkansas were to raise the score required to pass the various Praxis II exams by 0.25 standard deviations, the higher barrier to entry would be a more effective screen. Teachers who failed the Praxis II at the increased level were, on average, significantly less effective than those who passed. This was the case in both math and language arts.
Although the difference between teachers who passed the Praxis II and teachers who failed were statistically significant, there is a question as to whether these differences are practically significant. In other words, are the differences large enough to warrant action? Figure 2 presents a scatterplot that displays the relationship between a teacher’s effectiveness and their performance on the Praxis II professional knowledge exam. Both of these measures are standardized with a mean of zero and a standard deviation of one. At all points of the distribution, teachers vary in effectiveness. Imagine drawing a vertical line between the −3 and −2 on the horizontal axis. The individuals to the left of the line would fail the licensure exam, and the individuals on the right would pass the exam. Now imagine sliding that vertical line to the right. As the required passing score increases, you would remove more low-performing teachers, but you would also remove many high-performing teachers.

Scatterplot of teacher effectiveness in mathematics and performance on the Praxis II professional knowledge exam.
A more concrete illustration of this is presented in Figure 3. This figure plots the relationship between performance in math and a teacher’s score on the Praxis I mathematics exam. A vertical line is included at 171, Arkansas’s cut-score in 2010. The dashed line represents Louisiana’s cut-score in 2010. The individuals between the two lines represent the individuals who passed Arkansas’s current licensure exam, but would fail under Louisiana’s requirements. As you can see, these individuals vary in terms of effectiveness from a half of a standard deviation above the mean to a half below.

Scatterplot of teacher effectiveness in mathematics and performance on the Praxis I math exam.
Summary
There is a positive relationship between performance on licensure exams and performance in the classroom. However, the cut-scores on the Praxis I do not effectively weed out the lower performing teachers. Arkansas’s current Praxis II cut-scores are a more effective screen. Individuals who pass are significantly more effective in math, but not in language arts. Increasing the Praxis II cut-score by 0.25 standard deviations would improve the screen, but would also remove some effective teachers from the classroom.
Neither the Praxis I nor the Praxis II is a perfect licensure screen. That is, some low-performing teachers pass the exams and become a teacher and some highly effective teachers fail the exam and may be kept from the classroom. Moreover, the individuals who failed the exam tended to be significantly different than those who passed the exam. Namely, those who failed the exams are more likely to be minorities. Teachers who have performed worse on the exams are also more likely to be employed by more disadvantaged schools.
Policy Implications and Conclusion
The implicit purpose of licensure exams is to ensure a minimum quality among teachers by preventing inadequate teachers from entering the profession. For this to work, the licensure exam must be highly correlated with the desired outcome—performance in the classroom. If the exam and teacher effectiveness are not perfectly correlated, then the test will produce false positives and false negatives. That is, it will let some ineffective teachers into the classroom and it will keep out some highly effective teachers.
The Praxis licensure exams used by Arkansas to screen potential teachers are correlated to performance in the classroom, but that correlation is relatively low. Indeed, individuals who passed the Praxis I exams are not significantly more effective than those who failed the exam. This is not changed when the score required to pass the Praxis I exam is raised to the level of Louisiana. The Praxis II exam is a more effective screen, but is far from perfect. If Arkansas were to raise the score required to pass the Praxis II exams by 0.25 standard deviations, it would weed out significantly more ineffective teachers, but it would also remove some effective ones.
Moreover, low-scoring teachers are not randomly distributed among Arkansas schools. Just as Goldhaber and Hansen (2010) noted, minority teachers perform worse on licensure exams and they are more likely to work in schools with minority students. Presumably, these schools did not desire to hire the low-scoring teachers. More likely, they hired the teachers who were available or the teachers who fit their demographics. If Arkansas could raise the cut-score on the licensure exams, and replace the teachers who failed the exam with an average teacher, the state and the disadvantaged schools would be better off. To put it another way, if the number of teachers who passed the exam increased enough to make up for the increased number of teachers who will fail a more difficult exam, Arkansas would be in a better place. This, however, is not likely. Moreover, the minority teachers would be most at risk of being screened out by the increased rigor of the licensure exams.
Arkansas already suffers from teacher shortages, especially in rural and poverty stricken areas in the Mississippi Delta (Maranto & Shuls, 2012). Increasing licensure exam cut-scores will likely lead to greater teacher shortages in hard to staff subjects and regions of the state. These shortages will disproportionately impact the most disadvantaged schools.
If policy makers in Arkansas or other states want to improve teacher quality through a front-end policy, simply raising the score needed to pass the licensure exams does not appear to be an effective strategy. Although the current licensure exams are correlated with performance, there is tremendous variation in teacher quality among those who pass and those who fail the exams. Rather than consume time, effort, and political will to pursue policy changes unlikely to increase teacher effectiveness, Arkansas’s leaders and citizens would be better served by looking in new directions, searching for something more elusive: an innovative approach to improving teacher quality.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
