Abstract
We present the results of a study of local education agency (LEA)-based interventions to improve the reading outcomes of struggling readers in sixth grade. The sample included 1,076 intervention students and 3,644 comparison students. Regression discontinuity was used to evaluate intervention impact. The study contributes to the field by demonstrating (a) the importance of conducting rigorous evaluations of existing school interventions to understand the true impact of evidence-based practices when implemented under local circumstances and constraints, and (b) the use of regression discontinuity as an evaluation design is feasible within typical school settings. Implications for the study are discussed, especially in terms of providing a model for evaluation of LEA-driven interventions in the middle grades.
The risks for students with poor academic performance or low literacy skills and the concomitant costs to the individual and to society are enormous (Alliance for Excellent Education, 2011; Hernandez, 2011). Yet a sizable proportion of U.S. students are unable to meet reading proficiency standards in middle school and high school. According to the National Center for Education Statistics (NCES) report (Snyder & Dillow, 2013), only 34% of America’s eighth-grade students read at a level judged to be proficient for reading and completing school work at grade level. Nearly 25% of eighth graders did not reach a basic level of reading performance.
Reading Interventions With Secondary Level Students
For most middle school students, the problem is not simply decoding. They also struggle with fluency and comprehending what has been read (Cirino et al., 2013). They lack the background knowledge necessary to connect what they are reading to an established knowledge base, and they struggle with the vocabulary used (Biancarosa & Snow, 2006; Vaughn & Fletcher, 2012). Struggling adolescent readers are less likely to possess the engagement and motivation to read that is necessary to become frequent and fluent readers and comprehenders of grade-level text (Torgesen et al., 2007). By middle school struggling readers have endured a long history of negative academic experiences. Disengagement from school and the risk of dropping out become a serious problem, manifested by behavioral problems, truancy, and low motivation (Dynarski et al., 2008). Thus, multicomponent interventions are needed that address challenges with both academic skills and school engagement (Dynarski et al., 2008).
Several key reports addressing adolescent literacy delineate teaching strategies and practices that can improve outcomes for struggling readers. In Reading Next (Biancarosa & Snow, 2006), many of these strategies focus specifically on instructional practices in the classroom. The Institute of Education Sciences (IES) practice guide on adolescent literacy (Kamil et al., 2008) and a meta-analysis of 33 middle and high school studies (Slavin, Cheung, & Groff, 2008) concur that what happens in the classroom on a daily basis to challenge students’ meta cognition, facilitate comprehension and vocabulary, and provide frequent opportunities to interact with text, matters most in improving the literacy skills of struggling adolescent readers.
At the same time, however, literacy interventions implemented by educational staff have, on average, a much lower effect size than interventions implemented by researchers (Scammacca et al., 2007). The effects of evidence-based practices can be diluted or even nullified once they are implemented at scale by practitioners, or subjected to the local constraints, interpretations, and management of a particular school or district (Mckenna & Walpole, 2010). In a meta-analysis of reading interventions in late elementary, middle school, and high school, Scammacca and colleagues examined the effectiveness of interventions on reading comprehension as measured by a standardized, norm-referenced instrument. They found the mean effect size was 1.08 (95% confidence interval [CI] = [.57, 1.59]), when the intervention was delivered by a member of the research team, but it was only .21 (95% CI = [–.09, .50]) when it was delivered by the practicing teacher.
This narrative poses a significant conundrum. Reading deficits in middle school are especially difficult to remediate, and what happens in the classroom on a daily basis is the most impactful contributor to reading improvement. At the same time, interventions implemented by teachers are less effective than interventions implemented under the controlled research conditions that form the basis of the field’s understanding of best practices.
We have worked with many administrators and teachers who strongly believe, without causal evidence, that their intervention practices are essential for their struggling readers. It feels like the interventions are making a difference. As a result, these well-meaning educators continue to invest significant time, money, and resources on existing programs without any evaluation data to test the impact of those investments. Educators cannot afford to waste limited resources on ineffective interventions. For students’ sakes, it is necessary to evaluate if an intervention is having its intended effect. Only if this question is answered conclusively will educators know whether or not to continue an intervention, or if adjustments and changes are required, posthaste, to improve upon their existing practices. In this article, we demonstrate a feasible approach to provide local education agencies (LEAs) with the rigorous evaluation data they need to address this conundrum.
Notably, the IES underscored the need to embed ongoing program evaluation into the continuous improvement of intervention practices, stating “substantial improvements in student outcomes can be achieved if state and local education agencies rigorously evaluate their education programs and policies” (IES, 2010, p. 6). Despite this call for rigorous evaluations of state education agency (SEA) and LEA practices, a comprehensive literature search found only one published study that met the conditions of causal evaluations of intervention practices implemented under local conditions for the purpose of ongoing program evaluation. In the Chicago Public Schools, Nomi and Allensworth (2009) evaluated the effects of an education policy requiring double-dose algebra on academic math outcomes in ninth and 10th grades.
Middle School Intervention Project (MSIP) Overview
The MSIP was established, in part, to test and demonstrate the feasibility and practicality of conducting an impact evaluation of reading interventions under local conditions using in-place reading assessment measures. Our school partners led the direction of the evaluation by choosing foundational parameters of the evaluation, that is, their greatest need, practices to be assessed, grade levels of the students, and assessment measures to be included. During an extensive, collaborative planning period with the research team, the district superintendents identified the improvement of reading achievement of struggling middle school students as their most pressing initiative, for which they desired credible evidence of impact on student outcomes.
Each district had already made a serious commitment to implementing academic and behavioral interventions in the middle grades, with a focus on struggling readers and also had two common measures of reading assessment in place, districtwide—the statewide reading assessment and a measure of oral reading fluency. Each district also faced significant uncertainty regarding the actual impact of these interventions. This illustrates precisely the practice-to-research gap that has been called out by IES (2010).
MSIP represented an opportunity to conduct a valid evaluation of existing education practices under local conditions and resources which could subsequently be used as a model for other LEAs. Through our collaboration, participating MSIP districts could (a) discover if their investment of time, resources, and personnel was improving student outcomes; (b) receive data about actual implementation of intervention practices in each school; and (c) collaborate as a consortium to learn about practices implemented by other school districts within the state to address common challenges.
At the outset of MSIP, district leaders understood that helping struggling middle school readers requires data-driven, intense interventions (Kamil et al., 2008), which also address issues of school engagement (Dynarski et al., 2008). Each district already had multiple supports in place for their struggling readers. We documented the elements of their multicomponent interventions that were similar in scope across all three districts. The intervention as documented in each school district was comprised of three components: (a) targeted Tier 2 reading interventions, coupled with (b) Tier 2 school engagement interventions, and (c) data-based decision-making (DBDM) teams to review and respond to student data. All intervention students received all three components as part of the districts’ approach to provide intensive support to these students. However, some comparison students also received access to a variety of school engagement interventions and/or were discussed in DBDM team meetings. As the reading intervention was the only element on which the two groups completely differed, this evaluation provides an assessment of the impact of the reading intervention component.
Evaluation Design
To meet our goal of providing a model for evaluation that could be replicated by other LEAs, it was essential to choose an evaluation design that was feasible and acceptable to school leaders, while also providing causal evidence regarding impact. To do this, we incorporated a working definition of LEA that was not limited to a single school or district, but in this case included a “combination of school districts . . . recognized in a State as an administrative agency for its public elementary or secondary schools” (regulatory language regarding LEAs; Local Education Agency, 34 CRF 303.23, 2016). This evaluation feature provided the sample size necessary to detect small to moderate intervention effects.
The effect of reading interventions was assessed using regression discontinuity (RD). Like randomized control trials (RCT), RD provides an unbiased estimate of the causal effect of an intervention (Bloom, 2012). In RD (as in RCT,) some students receive the treatment while others do not. This decision is based on the student’s performance on a continuous variable relative to a cut point chosen for that variable. Students on one side of the cut point are assigned to intervention, while students on the other side are assigned to the comparison group (Bloom, 2012; P. Schochet et al., 2010). This approach to intervention assignment aligned well with MSIP districts’ existing intervention practices and is likely to be attainable for other LEAs that include a cut point approach to placing students into intervention. For example, this approach would work well when applied by LEAs that use benchmarks on a norm-referenced test to determine if students are at, above, or below expectations in a particular content area. RD, if applied appropriately, could allow educators to use their own assessments to evaluate their own practices and move forward in an ongoing and iterative improvement process.
In addition to Nomi and Allensworth (2009), a few studies have used RD to evaluate interventions with middle or high school populations. Luyten, Peschar, and Coe (2008) used RD to evaluate the effects of 1 year of schooling on reading performance and related activities in British teenagers, and Dougherty (2015) evaluated the impact of a double-dose literacy intervention on student reading test scores for middle school students.
The MSIP evaluation was conducted using in-place assessment measures following in-place assessment procedures on two standardized measures of reading. These LEAs were already experienced with providing differentiated services and instruction based on students’ performance on key reading measures. Relying on only a single cut point (rather than additional sources such as teacher input) to determine intervention status was new, yet it aligned with the districts’ approach to providing reading intervention based on demonstrated need. Thus, this assessment and the RD approach to condition assignment could co-exist with existing school practices and was palatable to administrators and other decision makers.
We used a pilot year of the study to understand and address specific issues related to the use of RD with our participating LEAs. This pilot year helped us identify questions that should be answered by LEAs interested in embarking on an RD evaluation of their own practices and programs. These guiding questions are delineated in Table 4. One issue was whether the cut point for condition assignment would be a single fixed point or flexibly set at multiple points on the distribution. Some RD studies have incorporated the use of multiple cut points for assignment (e.g., Deke, Gill, Dragoset, & Bogen, 2014; Wei, 2012), while others use a single cut point to assign relevant units to condition (e.g., Baker, Smolkowski, Chaparro, Smith, & Fien, 2015; Nomi & Allensworth, 2009). A strength of the single cut point approach is that it describes intervention impact at a specific point on the reading performance distribution; however, impact at other points on the distribution is unknown. A potential benefit of multiple cut points is generalizability of findings across a broader spectrum of students if intervention effects are significant at different cut points. Furthermore, use of multiple cut points allows for the inclusion of schools or districts that implement comparable educational interventions but use different cut points to determine placement.
In the pilot year, we determined we would need to allow for multiple cut points on the distribution because districts and schools differed in their criteria for placing students in reading intervention, even though they used the same statewide assessment measure. For example, some schools only provided intervention to those students who performed below a specific benchmark on the statewide reading assessment. Other schools threw a wider intervention net. Schools also differed in the percentage of students in need of intervention, with low performing schools needing to serve a higher percentage of students than higher performing schools. Thus, the specific cut point used at each school to assign students to intervention or comparison condition was determined by school leaders. Once again, this approach aligned with existing school practices, which made the evaluation more authentic for our school partners.
Research Questions
First, we ask whether schools could satisfactorily comply with the RD design, to consider if this approach is feasible for an LEA-based evaluation of intervention practices and to determine the degree to which causal associations between the reading intervention and outcomes are justified. Second, we ask to what extent were reading intervention practices effective in boosting scores on a measure of passage reading fluency and on the statewide reading test for students who had demonstrated low reading proficiency at the end of fifth grade.
Method
Setting
The intervention was implemented during the 2010–2011 school year in 40 schools located in five school districts in the Pacific Northwest. Across these districts, there were six schools (out of 46) that opted not to participate in the project. Sixteen schools were traditional middle schools with a grade 6–8 configuration. Two schools were K–8 and 22 schools were elementary schools in one district, serving K–6 students. District size ranged from 5,725 to 38,668 students and school size ranged from 179 to 1,096 students (Oregon Department of Education [ODE], 2010–2012). Across the districts, 35% to 61% of students were eligible for free or reduced-price lunch (ODE, 2010–2012), and 3% to 13% of students received English as a Second Language (ESL) services (ODE, 2010–2012). The percentage of minority students ranged from 27% to 47%. In each district, the largest minority student population was Hispanic (16%–33%; ODE, 2010–2012).
Description of Intervention
In response to the needs of their struggling readers, all schools implemented an intervention that consisted of three components: (a) a reading intervention, usually implemented as a reading class for at-risk students; (b) a school engagement component; and (c) a DBDM system for tracking the progress of at-risk students. Building-level implementation of the multicomponent intervention was largely defined by the existing practices of each participating school. Because it was possible that comparison students could also participate in the engagement and/or data team component, the key differentiator between intervention and comparison students was their participation in a reading intervention. The engagement and DBDM components remain important to document and describe to preserve a complete picture of the supports provided to these struggling students.
Reading intervention component
Variability among districts and schools in how the reading interventions were designed and implemented included type of program (e.g., published curricula, teacher-developed strategies, etc.), intervention dosage, staff training, and student–teacher ratio. Table 1 provides information on how those elements differed between comparison (English Language Arts [ELA]) and reading intervention classrooms, project-wide. Across the 40 schools and 232 reading intervention classrooms, reading interventions averaged 27 weeks in length (SD = 10.61) and met 4.8 times per week (SD = 57), for an average duration of 211.94 min per week (SD = 65.23). The mean student to teacher ratio was 9.47:1 (SD = 5.37).
Descriptive Summary of Reading Instruction for Comparison Reading (i.e., English Language Arts; n = 229) and Reading Intervention (n = 232) Classes, and Observations of Data-Based Decision-Making Team Meetings (n = 111).
Note. Data on program length and proportion of classes using published programs were not collected for English Language Arts (ELA) classes. Most ELA classes lasted for the full academic year (i.e., approximately 40 weeks).
Published programs were used in 75% of interventions. The most common programs were Language! (26 schools), Rewards (14 schools), Corrective Reading (nine schools), and Read Naturally (nine schools). In total, 23 different published programs were used across the 40 schools.
DBDM component
DBDM teams at each school collected and summarized achievement and engagement data for intervention students. Teams used these data to evaluate the effectiveness of the interventions and modify interventions as necessary. On average, these meetings lasted 55 min (SD = 14), and nine to 10 intervention students were discussed per meeting (SD = 6.8). At many meetings, comparison students or students in other grades were also discussed. Regardless of intervention status, within each meeting, quantitative data were discussed for about two thirds of the students and qualitative data were discussed for about half the students (see Table 1).
Engagement component
The purpose of the engagement component was to strengthen intervention students’ behavioral and psychological connections to school. Some examples include check-in/check-out programs, tutoring, homework club, and social skills groups. Engagement programs that included intervention students lasted an average of 22.0 weeks (SD = 10.8) and included meetings and activities that occurred an average of 2 times per week (SD = 1.6) for about 68 min each week. On average, the student–adult ratio in engagement interventions was 9.25:1 (SD = 7.89; see Table 2 for more details).
Descriptive Summary of Engagement Programs by Treatment Condition.
Note. Comparison engagement programs are activities identified by participating schools as an engagement intervention option for students, but in which no intervention students participated. Program categories identify the primary focus of the engagement programs. Academic programs included study hall, tutoring, and activities focused on math, reading, and writing. Behavioral programs included formal check in programs, counseling, and student mentoring. Leadership programs included social skills training, volunteering, and service activities. Student interest programs included arts, music, technology, and sports. Categories and program structure represent the proportion of programs by category, and thus sum to 1.
A key takeaway from schools’ multicomponent approach to intervening with struggling readers is that schools expended significant time and resources supporting these students and could reasonably expect these efforts to have an impact on students’ reading achievement. Through this evaluation partnership, schools and LEAs were able to obtain causal evidence of impact of their reading interventions.
Description of the Comparison Group
Table 1 illustrates how reading intervention students and comparison students differed. Comparison students received their reading services in ELA classes. Table 1 describes these services in terms of duration, frequency, and student–teacher ratio. Some comparison students received some attention in data team decision making (Component 3), but this support was less extensive than the support received by intervention students.
Student Assignment and Outcome Measures
We accessed reading proficiency data on students in the study when they were in fifth grade and again in sixth grade. Fifth-grade data were used to assign students to the intervention or comparison condition, and sixth-grade data were used to evaluate the impact of the intervention. In four districts, the fifth-grade measure of reading fluency was the easyCBM Passage Reading Fluency (PRF) measure (Alonzo, Tindal, Ulmer, & Glasgow, 2006). The fifth district administered the Dynamic Indicators of Basic Early Literacy Skills Oral Reading Fluency (DIBELS ORF; 6th edition; Good & Kaminski, 2002) measuer. In sixth grade, all five districts used the sixth grade easyCBM PRF measure. We also accessed fifth- and sixth-grade reading performance data on the statewide reading test.
Oregon Assessment of Knowledge and Skills Reading/Literature (OAKS)
The OAKS is a criterion-referenced test aligned to grade-level content standards. The Reading and Literature assessment is a multiple-choice test focusing primarily on reading comprehension (79%) and to a lesser extent on vocabulary knowledge (21%) with a statewide mean and standard deviation of 225 and 9, respectively, for Grade 5 and 229 and 9 for Grade 6 (ODE, 2012). Students could take the test up to 3 times within a year, in an effort to “meet” or “exceed” expectations (ODE, 2012). Statewide, school districts retain the best OAKS score in their databases; thus, our analyses are also based on the students’ best OAKS score for the year.
easyCBM PRF
easyCBM PRF (Alonzo et al., 2006) is a standardized, individually administered measure of the speed and accuracy with which students read connected text aloud from a 250- to 350-word, grade-level, narrative passage. The final score is a rate of correct words per minute (CWPM). The average correlation between a reference passage and 19 other fifth-grade passages is .86 (SD = .04; Alonzo & Tindal, 2008).
DIBELS ORF
DIBELS ORF (6th edition; Good & Kaminski, 2002) is a different standardized measure of oral reading fluency, administered and scored in the same manner as the easyCBM PRF. Single probe reliability for fifth-grade ORF was .93. Multiprobe reliability for fifth-grade ORF was .98.
Procedure for Assigning Students to Condition
Transformation of OAKS and PRF/ORF into z scores
The OAKS score was standardized by subtracting the fifth-grade population mean of 225 from the student’s individual score and then dividing this difference by the fifth-grade population standard deviation of 9. PRF and ORF were standardized similarly. Students’ standardized z scores for each measure (OAKS and PRF or ORF) were then averaged to create a combined performance z score for each student. Thus, this overall z score included information on three components of reading proficiency: comprehension, vocabulary, and oral reading fluency. For students who took only one assessment, their overall z score was set to equal the z score for that single measure.
Schools use of z scores to place students into condition
The overall fifth-grade z score was used for assignment to intervention or comparison condition. We instructed each school to pick a specific z score to establish the preliminary cut point. The school’s final cut point was then placed half way between a student at that z score and the next higher student. All students with z scores below this cut point were expected to be placed into the intervention condition, and students with z scores above the cut point were placed in the comparison group. The full range of cut points chosen by the schools was 0.03 to −1.15. The inner 80% of the distribution runs from –.35 to –.95. These values correspond to the 35th and 15th percentiles, respectively, of the MSIP student cut score distribution. In other words, the vast majority of schools were intervening with readers in the bottom third of the distribution.
Strategies to Increase Study Compliance Among Schools
Schools wanted assurance that assignment to condition using a cut score was a valid way to assign students to intervention and aligned well with their existing practices of placing students into intervention based on demonstrated need on key reading assessments. The research team used several strategies (described below) to increase understanding and compliance among district and school personnel. These strategies are pragmatic and could be implemented by LEAs seeking to proactively plan for a rigorous evaluation of their own practices.
A reliable and valid indicator of reading proficiency
To demonstrate to school leadership that the overall composite score was a reliable and valid indicator, we fit a multilevel confirmatory factor model to the three standardized measures in our original sample of 6,690 fifth-grade students nested in 97 classrooms. The fit was acceptable given the large sample size (chi-square = 33.28, df = 8, p < .001, root mean square error approximation [RMSEA] = .022, comparative fit index [CFI] = .995, Tucker–Lewis index [TLI] = .996). The estimated reliabilities for fluency (.67) and OAKS (.63) as indicators of overall reading proficiency imply that the reliability of the composite cut score is about .79, higher than either measure alone.
Use of multiple cut points
Each school chose their cut point on the z score distribution for assigning students to intervention and comparison conditions. This concession was critical to maintain participation by all schools in the sample. For analyses that pooled students across schools, we centered the z scores within each school by subtracting the individual school’s cut point so that all schools had the same cut point (i.e., 0), for purposes of analysis and interpretation. Comparable to the use of school mean centering in a multilevel model (Raudebush & Bryk, 2002 ), we included the school cut point value in the multilevel RD model as an additional school-level predictor of RD impact.
At the initiation of MSIP, schools agreed to follow design guidelines so that a minimum of 20% up to a maximum of 80% of students received the intervention to help maintain statistical power for the RD. Some schools placed close to, but not quite, 20% of students in intervention. Across the five districts, 22.6% of the sample was placed in intervention.
Use of wild cards in the assignment process
Schools were allowed to exempt up to 5% of their sixth graders from placement into the condition indicated by the student’s z score. We referred to these exemptions as wild card slots. Some school staff had strong opinions that certain children should or should not be placed into intervention and felt in some cases that their professional judgment needed to override the use of a cut score. The wild card solution accomplished this while maintaining school leaders’ willingness to use the cut point approach to intervention placement. Schools were required to identify which students were given wild card status prior to receiving information about their z scores and choosing their cut point. This was done to avoid dependence between students’ scores and cut point selection (i.e., to eliminate the possibility that staff would choose a specific cut point to purposely include one or more students in the desired group).
Of those students designated as wild cards, the strongest predictor of wild card status was the dichotomous variable indicating whether a student met expectations on the fifth-grade OAKS-R (a score of 218 or higher). This variable positively predicted being wild carded out of the intervention for students who were below the school cut point and negatively predicted being wild carded into the intervention for students who were above the school cut point.
Final Student Sample
Students had to meet three criteria for study inclusion. Students had to (a) have a valid fifth-grade reading score on at least one of the two reading measures; (b) attend an MSIP school for two of three or three of for academic terms; and (c) not have taken the alternate form of the OAKS (a different version of the OAKS with a different scoring and scaling procedure, administered to students with special needs and Individualized Education Programs [IEPs]) in fifth or sixth grade. These inclusion criteria resulted in a total of 4,720 sixth-grade students in the study of which 1,076 (22.8%) received intervention.
Across both conditions (and including these planned wild card students) at the end of the school year, 4.7% of the sample were noncompliers. This is an adequate compliance rate and another indicator that this evaluation design can be implemented with LEAs. Of the 1155 students below the cut point, 13.0% were noncompliers (wild cards 7.4% and unplanned 5.6%), that is, they should have participated in the intervention but did not. Of the 3,565 students above the cut point, 2% were noncompliers (wild cards 1.2% and unplanned 0.8%). They participated in intervention although they should have been placed in the comparison group.
Fuzzy RD Design and Analysis
The fuzzy RD (Bloom, 2012) accommodates imperfect compliance with group assignment (i.e., noncompliers/wild card students are included in the analyses) under the same background assumptions as the standard or sharp RD but with two additions. First, the effect size estimate only pertains to compliers, and second, there has to be a significant gap in the probability of receiving the intervention (PRI) at the cut point despite the noncompliance (Bloom, 2012). Notice that in the fuzzy RD, it is important to distinguish between intervention assignment (below the cut point) and receipt of intervention, because noncompliance makes these two different. The fuzzy RD requires estimating two regression equations, for the outcome and for the PRI. The analysis for the outcome is the same as in the sharp RD design and consists of estimating the difference in mean outcomes (outcome gap) between the intervention and comparison groups at the cut point. The analysis for the PRI is very similar but consists of estimating the difference in mean probability (probability gap) of receiving the intervention above, but at the cut point versus below, but at the cut point. The fuzzy RD effect is estimated by the raw outcome gap divided by the probability gap (Bloom, 2012).
For a standard RD analysis, it is important to accurately model the relation between the cut score and the outcome on both sides of the cut point. For a fuzzy RD analysis, the same concern applies for the relation between the PRI and the cut score. We used the generalized additive model (GAM) procedure as implemented in R by Wood (2006) where the degree of smoothing was determined by the data, using generalized cross-validation procedures and we used separate smoothing procedures on each side of the cut point.
Because our design is multilevel with students nested within schools, we used the multilevel extension of GAM (GAMM4) for three reasons. The first is to obtain standard errors that take in to account the nonindependence of students within schools. The second is to allow for variation across schools in the mean outcome and mean outcome gap (i.e., RD effect) and the mean PRI and PRI gap at the cut point. Finally, a multilevel approach lends itself to looking at RD effects in individual schools (Gelman, Hill, & Yajima, 2012), which we intend to pursue in future research. Our district-level analyses used separate multilevel GAMs.
For the outcome models, we use a multilevel Gaussian GAM, and for the PRI model we use a multilevel logistic GAM. We make the standard assumption that random effects are multinormally distributed. Because of the cut point centering procedure, we included the school cut point as a school-level predictor of mean outcomes and PRI on both sides of the cut point.
Results
Initial Checks of Assignment Integrity
We performed fuzzy multilevel RD analyses on observed pseudocovariates to check that assignment was based solely on the cut score (Imben & Lemieux, 2008). We checked gender, special education status (SPED), limited English proficiency (LEP) status, and for those districts (three out of the five) that provided the information, free or reduced-price lunch status. We found no significant evidence of fuzzy RD effects on any of these covariates. That is, it appears schools did not rely on any of these variables in addition to the cut score to assign students to condition.
We also performed the McCrary test (Dimmery, 2013; McCrary, 2008) to check for discontinuities in the cut score distribution at the cut point to assess for deliberate manipulation of intervention assignment status. There was no visual or statistical evidence of any discontinuity or irregularity in the cut score distribution. We checked that the cut score was relatively continuously distributed. The cut score took on 3,517 (75%) unique values among the 4,720 students in the study. Within –.5 and .5 points of the cut point, 1,175 out of 1,563 (75%) scores were unique and showed very little clumping of points at identical values. The results of the McCrary test and of the multilevel RD analyses on observed pseudocovariates revealed no evidence to undermine the internal validity of the RD. With these tests, we address and affirmed one aspect of schools’ compliance with the RD evaluation.
Descriptive Statistics
We present descriptive statistics in Table 3. Students were assigned to the intervention based on having a low cut score. Thus, the large mean differences between intervention and comparison groups on the cut score (and the OAKS and PRF outcome scores) reflect this aspect of the RD design. At the end of Grade 6, the comparison group is still substantially higher than the intervention group (11 points for OAKS and 68 points for PRF). This observed difference would not necessarily preclude a significant discontinuity at the cut point.
Descriptive Statistics by Group Status.
Note. Q25 = 25th Quartile; Q50 = 50th Quartile; Q75 = 75th Quartile. OAKS = Oregon Assessment of Knowledge and Skills; PRF = Passage Reading Fluency.
Fuzzy RD Results
Project wide analyses
Preliminary analyses included the cross-level interaction between school cut point and the intervention assignment variable (less than or equal to the cut point) to check for variation in RD intervention effects that might be linearly related to the level of the school cut point. We also checked for the same kind of effects on the PRI gap. Note that this is a specific form of heterogeneity in RD effects across schools in addition to the school-level random effect in the GAM for any form of heterogeneity in RD effects. No significant effects were detected so we dropped cross-level interactions from all the models.
Regarding the fitted smooth functions and the probability gap from the multilevel logistic GAM for the PRI, the mean PRI above the cut point but right at the cut point was .104 and below the cut point but right at the cut point the corresponding probability was .839 resulting in a gap of .735, which was highly significant (z = 18.44, p < .001), as required by the assumptions for fuzzy RD (Marmer, Feir, & Lemieux, 2014). The variance component for school-level variation in the PRI gap was also highly significant (p < .001).
In Figure 1, there is virtually no gap visible in mean OAKS at the cut point. The raw outcome gap from the multilevel Gaussian GAM was .520, which was not significant (p = .301). After dividing by the probability gap, it became .702, which was also not significant (p = .302) using the approximate standard error from P. Z. Schochet (2008). The population standard deviation for the OAKS is 9 so the standardized fuzzy RD effect size is .078, very small.

Fitted GAM for the Grade 6 OAKS reading score as a function of the cut score.
In contrast, in Figure 2, there is a small negative gap visible in mean PRF at the cut point. The raw outcome gap from the GAM model was −5.302 CWPM, which was significant (p = .007). After dividing by the probability gap, it became −7.156 CWPM, which was also significant (p = .007), using the approximate standard error from P. Z. Schochet (2008). Dividing the fuzzy RD effect by the standard deviation of PRF in our sample, 49.34, leads to a standardized effect size estimate of −0.145. This is a small (negative) effect of the intervention on passage reading fluency.

Fitted GAM for the Grade 6 PRF score as a function of the cut score.
Districtwide analyses
We also checked for variation in RD impact across districts by estimating the GAMs described above for each outcome separately for each of the five districts. Districts varied widely in both student sample size (372 to 1,559) and number of schools (3 to 22). Across both outcomes and five districts, only one significant effect (negative) for PRF was detected. Cochran’s Q test (Cochran, 1954) for differences among the districts for both outcomes was nonsignificant (p > .50). Overall, results for both OAKS and PRF on the estimates for the five districts were essentially the same as for the entire sample. Although districts used their own building-based practices for the reading interventions, all districts were similarly ineffective in improving student reading outcomes at the break between intervention and comparison groups.
Schoolwide analyses
For both the PRF and OAKS models, the estimated variance component for the school-level raw RD intervention effect was strongly significant (p values < .001), indicating that the raw RD intervention effects varied significantly across schools. For the OAKS, although there is no overall raw RD intervention effect, the significant variation suggests that positive raw RD effects in some schools were canceled out by negative raw effects in other schools. The GAM results suggest that the middle 90% of the school population would have raw OAKS effects that range from −2.61 to 3.65. For PRF, the overall intervention raw effect was small and negative, but the significant variation suggests that larger raw negative effects in some schools were balanced out by smaller raw negative or even positive effects in other schools. The GAM results suggest that the middle 90% of the school population would have raw PRF effects that range from −11.50 to 0.90. Although it is beyond the scope of the current article, future research efforts will be devoted to estimating individual school-level effects and identifying the elements of the reading interventions utilized at schools with high, positive effects.
Discussion
Summary of Findings in Response to Research Questions
Compliance with RD design
First, we asked how closely schools would adhere to the parameters of condition assignment. Many earlier RD studies have relied on a cut point based on a clear criterion driven by local educational policy and still have found variability in cut score adherence. For example, using a cut point based on birthdate, Wong, Cook, Barnett, and Jung (2008) found assignment to the correct group was generally followed. More than 90% of participants complied with the cut score. Other researchers have found less compliance, particularly when the cut point was based on a measure of performance. For example, in a literacy intervention evaluated by Dougherty (2015), in sixth grade, more than 20% of students eligible for the intervention did not receive the intervention, while 55% of students who scored above the cutoff did receive one or two semesters of the supplemental reading intervention.
MSIP, in contrast, applied RD in a prospective evaluation that was designed collaboratively with the school district partners. The extent to which schools could maintain fidelity to the RD design parameters was an important question for demonstrating the practicality of RD as an option for local, rigorous impact evaluations. We worked closely with school staff to explain the RD design, to increase the likelihood schools would comply with group assignment. All schools insisted on needing leeway to use their professional judgment to override some assignments based on a cut score. Consequently, we worked out a compromise whereby schools could use their professional judgment to place up to 5% of students in the intervention or comparison condition regardless of the cut score. We believe the allowance for wild card slots was instrumental in maintaining participation and compliance with the evaluation. In total, 95.3% of students received the correct treatment based on their assigned condition.
This extensive, proactive work with the districts and schools led to compliance rates that are sufficient to use a fuzzy RD to estimate a causal association between the intervention and student reading performance. Furthermore, results of the McCrary test and of the multilevel RD analyses on observed pseudocovariates revealed no evidence that school personnel manipulated condition assignments or used student characteristics other than the cut score (e.g., ethnicity, gender) to assign students to the intervention or comparison conditions.
Intervention impact
Our second research question asked the extent to which LEAs’ sixth-grade reading interventions impacted scores on two measures of reading performance. Sample-wide, the reading interventions did not have a positive impact on reading achievement. There was no significant difference on the OAKS between groups at the cut score. On the PRF, there was a small, negative effect of the intervention. There was no significant variation between districts in the size of intervention effects, except for one district which had a larger negative effect on reading fluency than the other districts. Secondarily, we found significant variation in raw impact across schools for both reading outcomes, but we did not conduct tests of hypotheses about specific sources of variation in intervention effect size across districts or schools.
Implications for Practice
Guiding questions
By providing a demonstration in this article of how LEAs can use their own data to rigorously evaluate their own practices, we expect that educators may (a) apply the model in their own districts and (b) use the results of their own evaluations to make data-driven decisions to modify and improve their interventions, as needed. In Table 4, we provide a set of guiding questions that districts can use to lay the foundation for their own evaluations. These questions guided the MSIP study, as demonstrated throughout this article.
Questions for Determining Readiness for LEA Evaluation of Existing Practices.
Note. LEA = local education agency.
The value of evaluating the impact of existing school practices
The evaluation demonstrated null or negative impact of the intervention on student outcomes—crucial information for these LEAs to have when making decisions about instructional programming and budgeting of resources in upcoming academic cycles. It is not clear why the intervention suppressed reading fluency growth, as suggested by the small difference favoring the comparison group on the passage fluency measure. Classroom observations suggested that attention to fluency development was about the same in reading intervention and ELA classrooms, 18% and 12% of time, respectively. Furthermore, observations indicated that both reading intervention and ELA classes spent quite a bit of class time on reading comprehension, about 30%, which was the primary focus of the state assessment. Intervention and ELA classrooms also spent substantial time on other important aspects of literacy instruction (about 28%), which included vocabulary content, another important focus on the state reading assessment (for a detailed account of the observation tools and procedures, see Nelson-Walker & Crone, 2011).
In the present study, one objective was to partner with districts and show by way of example how to move forward with evaluations given that schools and districts were collecting all the necessary data for this type of evaluation. Funding was used for documenting intervention implementation (which, while advisable, is not essential for a rigorous evaluation) and for methodological expertise to conduct the RD evaluation. We found that with some help and guidance from us, districts and schools were quite capable of engaging in intervention implementation that would allow for the use of an RD design to measure impact. However, most districts will not have personnel to conduct the complex statistical analyses involved in RD. Consequently, we, as the evaluation team, conducted all the statistical analyses. While we did not conduct a cost/benefit analysis in this study, the greatest expense for future evaluations would be in contracting for statistical expertise, assuming districts do not have the expertise for this. Another assumption regarding cost is that districts and schools can use existing practices and existing data (as suggested in Table 4) for the analysis.
The large initiatives the participating districts had in place during the study reflect the way many LEAs operate in trying to provide the range of programs, services, and supports expected of them. In the current study, the district superintendents rated middle school reading interventions as their number one concern. Simultaneously, district and school personnel had expressed concerns about the potential harm of withholding an intervention from students who needed it. These interventions had been in place for some time, but the LEAs had no credible evidence they were improving reading achievement. We believe the costs/benefits of conducting an evaluation of the impact of an intervention must be considered in relation to the costs of implementing interventions that are not improving student outcomes, as was found in this study.
Bridging the practice to research gap
Establishing the cut point in an RD design is often the result of an educational or economic policy instituted from the top, down . For example, in the double-dosed algebra intervention evaluated by Nomi and Allensworth (2009), the entry criteria set by the school district was all students who scored below the national median on the math portion of the eighth-grade Iowa Test of Basic Skills (ITBS). For the supplementary reading intervention evaluated by Dougherty (2015), the cut point set by the district was the 60th percentile on the reading portion of the fifth-grade ITBS.
In contrast, in MSIP each school chose their own cut point. Each school had procedures in place to provide reading intervention to students, and a flexible cut score enabled them to serve the number of students they typically would serve. A single, uniformly applied cut point was not palatable to the participating schools, and if used, the evaluation would not have mirrored the intervention placement decisions schools made to determine how many students could be served and what level of reading performance corresponded to a need for intervention.
Rigorous evaluations of LEA-based practices that fit comfortably within the existing contextual parameters of daily school operations are sorely needed. This study makes a valuable contribution to the literature and to practice because it demonstrates a feasible model for successfully conducting a rigorous evaluation of teacher-delivered, existing practices. Furthermore, the study demonstrates that with a reasonable amount of planning, a high-quality RD can be implemented with sufficient compliance to reach valid conclusions about intervention impact.
RD designs work well to evaluate school-based practices and interventions organized around the concept of providing more intense services and interventions to those students who demonstrate low performance or insufficient progress (e.g., tiered interventions in response to instruction [RTI] service delivery models). A cut score could be used, based on a score from a test(s) administered at a single point in time, to determine which students would receive a more intense intervention and which students would not (see also Baker et al., 2015).
Limitations
We address three limitations important to interpreting the findings. First, while we demonstrate lack of intervention effects, RD findings only apply to those students close to the cut score. Our use of multiple cut points, however, means the lack of a positive intervention effect does not apply to a single point on the cut score distribution, but across multiple points associated with the different cut scores used by the schools, and only for those students close to the cut score used in their school.
A second limitation is related to the complexity of the interventions that schools implemented. Districts and schools defined the interventions used in this study, and a range of programs and practices constituted the treatment. There was no single program or set of programs that could be identified and considered in relation to the findings. Given the range of programs used, it is very difficult to identify which specific programs or practices may have been responsible for the findings. We recognized this limitation prior to implementing the evaluation. We understood we would not be able to make statements about specific programs or practices; rather, we intended to examine the overall effect of existing practices centered on a common goal, and having common critical features.
Furthermore, we cannot describe the extent to which published programs were implemented with fidelity, nor comment on whether teachers received adequate professional development (PD) to implement the intervention practices. This was not the point of the study, and it would not have been possible to collect fidelity data across the wide range of programs implemented in the 40 schools. Rather, the purpose of the study was to provide schools with data on the causal impact of their own chosen practices. Upon learning about the null to negative impact, educators would be responsible to make decisions about discontinuing or altering current practices. Potential changes might include attending to fidelity, dosage, PD, or trying different programs. Future LEA-based evaluations could limit the scope of their study and examine intervention impact in the context of a specific program and in relation to fidelity to that program.
Footnotes
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: The research reported here was supported by the Institute of Education Sciences, U.S. Department of Education, through Grant R305E100041 to the University of Oregon. The opinions expressed are those of the authors and do not represent views of the Institute or the U.S. Department of Education.
