Abstract
We present results from the first randomized experiment of a remedial inquiry-based science education program for low-performing elementary students in a developing country. Among third-grade students in 48 low-income public elementary schools in Metropolitan Lima who score in the bottom 50% of their school baseline science distribution, half are randomly assigned to receive remedial inquiry-based science education in after-school sessions, and the remaining half to business as usual control conditions. Assignment to treatment increased endline science achievement by 3 percentiles (0.12 SD) with greater gains for students who attended at least one remedial session, and a concentration of gains among boys. We cannot reject the null hypothesis of no indirect science achievement gains among nonparticipants.
Keywords
I
Recent rigorous evidence from various developing countries suggests that pedagogical shifts that facilitate customized teaching to the needs of every child can dramatically improve learning (Banerjee et al., 2017). In particular, remedial education, by which students receive customized, self-paced teaching, shows promise at improving short- and medium-term academic performance of low-achieving students in a variety of contexts (Abdul Latif Jameel Poverty Action Lab [JPAL], 2018). The evidence to date on remedial education, however, is mostly limited to improving basic mathematics and literacy skills. 1 The evidence on remedial mathematics and literacy education suggests that direct instruction may be an effective pedagogical model for low-achieving students (Houtveen & van de Grift, 2007, 2012; Kaiser, Palumbo, Bialozor, & McLaughlin, 1989; Linan-Thompson & Vaughn, 2007). However, research on whole-class science instruction suggests that inquiry-based instruction—in which students engage in hands-on practical work with different degrees of teacher guidance—improves learning more than traditional classroom practices (Brickman, Gormally, Armstrong, & Hallar, 2009; Ergül et al., 2011; Harris, Penuel, DeBarger, D’Angelo, & Gallagher, 2014; Hmelo-Silver, 2004). It remains unclear whether inquiry-based instruction is an effective pedagogical approach to improve science skills among low-achieving, early-grade students (Hmelo-Silver, 2004).
We present experimental evidence on the learning impacts of a remedial inquiry-based science education program for low-achieving third-grade students in 48 low-income public elementary schools in Metropolitan Lima, Peru. Peru is a relevant study context for several reasons. Teaching in Peru is governed by centralized national curricular standards. As a result, many students typically fall behind the curriculum. In the 2015 application of the Program for International Student Assessment (PISA) test, Peru ranked 62nd in mathematics and 64th in natural science out of 70 participating nations, and 58% of Peruvian students were low achievers in science as compared with 21% for Organisation for Economic Co-Operation and Development (OECD; 2016) students. In the 2013 TERCE regional standardized test, close to 40% of sixth-grade Peruvian students scored at the lowest level of achievement in science (Laboratorio Latinoamericano de Evaluación de la Calidad de la Educación [LLECE], 2015).
Recent policy efforts in Peru to address these dismal results have primarily focused on training and coaching teachers on how to best deliver national curricular standards for science in early grades. Prior randomized evaluations have concluded, however, that these policy initiatives only improved science achievement among early-grade students with above-average baseline performance (Beuermann, Näslund-Hadley, Ruprah, & Thompson, 2013).
The remedial science education program we evaluate aims to improve scientific skills among low-performing third-grade students through supplemental, customized instruction delivered through a student-centered inquiry-based approach. Although the program shares some features of typical remedial education interventions such as supplemental, small-group instruction, to our knowledge this is the first rigorous study to date to document impacts of a remedial inquiry-based model aimed at improving scientific proficiency in early grades, in which struggling students are taught in smaller groups according to their learning ability.
Among third-grade students who score in the bottom half of their school’s distribution on a science test administered at baseline, half of them are randomly assigned to receive 16 supplemental remedial science sessions (90 min/session), while the remaining half are assigned to business as usual control conditions. The remedial sessions follow an inquiry-based format and take place in schools—typically in the afternoon—in groups of nine students, on average. Tutors are public-sector elementary school teachers selected among volunteer candidates. Prior to the start of the remedial sessions, selected tutors received content knowledge, pedagogical training, and detailed, highly structured materials that included flipcharts with student-centered activities for each session as well as formative evaluation rubrics.
Estimates of the program’s direct effect suggest that assignment to treatment increased endline science achievement of low-performing students by about 3 percentiles or 0.12 SD, with slightly greater gains of 0.13 to 0.16 SD for students who attended at least one session, and likely even greater gains for those with complete remedial session attendance. The direct effects of remedial education participation, however, are concentrated primarily among boys. The program’s direct effect probably represents the total effect because we cannot reject the null hypothesis of no indirect gains among nonparticipants. Although the point estimate on indirect test-score gains is small in magnitude, one caveat is that statistical power for the test is likely low due to insufficient classroom-level variation in the fraction of students assigned to remedial sessions within schools.
In addition to our main substantive findings, we make two methodological contributions. The first is to introduce an empirical framework and application to estimate nonexperimentally the indirect effects of remedial education on nonparticipants, which can be viewed as an ancillary test of the stable unit of treatment value assumption (SUTVA; Rubin, 1980). This framework draws on prior work by Hudgens and Halloran (2008); Sinclair, McConnell, and Green (2012); and VanderWeele, Hong, Jones, and Brown (2013). The second methodological contribution is to use our experimental design and eligibility threshold rules to illustrate an empirical validation of the regression discontinuity design (RDD), similar to Cook and Wong (2008) and Buddelmeyer and Skoufias (2004).
The remainder of this article is organized as follows. The second section discusses previous efforts in Peru to identify an effective primary education science model and the inquiry-based remedial science education approach we design and evaluate. The third section describes the sample and experimental design. The fourth section describes the data and analytical approach. The fifth section reports our main findings, and the sixth section concludes with a discussion and directions for future research.
Background and Program Description
In this section, we describe science classroom practices and recent efforts to boost science skills in Peru that motivate the present study, and describe the program we evaluate.
Pedagogical Practice in Science Classrooms and Efforts to Boost Science Skills Among Peruvian Students
Peruvian students have poor overall performance on international assessments. In the last application of the PISA test, for example, Peru ranked 62nd in mathematics and 64th in natural sciences among the 70 participating nations (OECD, 2016). One fifth of Peruvian students place in the lowest proficiency level of PISA in science, which means that they do not master even the most basic skills. The PISA assessment indicates that Peruvian students lack critical reasoning skills and the ability to analyze and synthesize information, and to apply new knowledge in real-life settings.
Lack of adequate teaching skills among Peruvian teachers may help explain students’ poor performance on comparative science and mathematics assessments. Teachers typically teach the curriculum without setting aside time for struggling students. Evidence suggests that Peruvian teachers overemphasize the least cognitively demanding topics. Teachers frequently assign learning tasks that are not cognitively demanding, students rarely get teacher feedback and, when they do, it is often erroneous (Cueto, Ramirez, & Leon, 2006). Moreover, half of mathematics teachers nationwide cannot perform basic arithmetical calculations (Alfonso, Bos, Duarte, & Rondon, 2012).
In Peru, scientific learning typically follows a teacher-centered instruction model. Teacher lectures take up most of class time and limited time is devoted to practical work. To the extent that they do, teachers conduct practical work themselves, limiting student opportunities for hands-on learning (Loera, Näslund-Hadley, & Alonzo, 2013; Näslund-Hadley, Loera, & Hepworth, 2014).
To address some of the country’s mathematics and science educational challenges, Peru’s government piloted a program in 2010 that aimed to promote critical thinking and scientific reasoning skills among third-grade students. This program, based on the 2008 national curriculum that includes areas such as the physical world, the human body, and living beings and environment, aimed to teach children about scientific models and their applications. A key component of this pilot science program was teacher training, with a specific focus on mastering the structure and content of inquiry-based learning approaches (Tutwiler & Grotzer, 2013). 2
A school-level randomized evaluation of the pilot teacher-training model in 62 districts of the Region of Lima concluded that the program only improved science test scores among boys in urban areas and for students with above-average baseline performance. Government program administrators made some adjustments to place an increased focus on girls’ confidence in their science skills, and working groups were separated by gender for some activities to ensure that girls got hands-on experience. To close the geographical divide, efforts were made in a subsequent 2012 pilot to increase compliance in rural areas to ensure that all teachers benefited from mentoring. These adjustments to the program made the average gender and geographical gaps insignificant. However, among students in the bottom half of the baseline score distribution, the pilot program still had no impact (Beuermann et al., 2013). Thus, the pilot program widened the science achievement gap between high and low performers.
The fact that earlier efforts to improve science learning only improved performance of above-average students motivated the government to identify pedagogical approaches to specifically benefit low-performing students. This context motivates the present study, which investigates whether remedial inquiry-based science educartion for the lowest performing students helps improve their science achievement.
Program Description: The Remedial Inquiry-Based Science Program
The remedial science program we study aims to help early-grade low-performing students master theoretical and practical science skills through student-centered, hands-on inquiry-based methods. The goal is that, when confronted with an unfamiliar situation, students can pose questions, formulate hypotheses, think critically, and work collaborative to produce relevant answers. The remedial sessions are offered as supplemental instruction after regular school hours and do not substitute any existing science instruction. As a by-product, the program seeks to promote healthy study habits, academic motivation, and love of learning.
Universidad Cayetano Heredia—a private, science-centered, leading research and teaching university in Peru—developed the structure and contents of the remedial science program. The program has four components: (a) development of pedagogical materials, (b) selection and training of tutors, (c) selection of students, and (d) implementation of remedial sessions in schools.
Pedagogical Materials
To develop the pedagogical materials, Universidad Cayetano Heredia employed two local pedagogy specialists, one specialist in primary education and one in science education. These two specialists developed the tutoring materials, the contents of which are based on the 2008 National Curricular standards for teaching science to third-graders.
Based on the National Curriculum, to bridge the issue of gaps in tutors’ content knowledge, the specialists developed detailed and highly structured materials that included flipcharts with activities for each session and formative evaluation rubrics. That is, the materials combine elements of explicit instruction with inquiry-based activities (Hmelo-Silver, Duncan, & Chinn, 2007). In this inquiry-based approach, tutoring sessions begin with a challenge/question. For example, as part of a weather module, students explored why Lima is covered in fog. The tutor guided them in the formulation of hypotheses, design of experiments, and discussion of their findings as the students made their own fog in jars. Students were then encouraged to formulate preliminary answers based on prior knowledge, acquire new information through experimentation and reading, restructure prior knowledge, establish conclusions, and apply the new knowledge to unfamiliar situations.
Selection and Training of Tutors
Universidad Cayetano Heredia selected 16 tutors—15 of whom were women, like the majority of public school teachers in Peru. Universidad Cayetano Heredia assigned the male tutor to schools located in high-crime areas. Tutor selection took place between March and May 2014. Selection criteria included (a) minimum of 2 years of primary school teaching experience, (b) positive attitude toward the teaching and learning of science, (c) assertive communication and class-management skills, and (d) ability to create respectful, empathetic, and tolerant relationships with children.
Tutors are local primary or secondary public school teachers, although not necessarily teachers in the schools in which they provide tutoring. Tutors are paid an hourly wage of US$10 for their services and transportation, which is slightly below what primary education teachers earn on average (US$14 per hour). The pay was not linked to the performance of the tutor, nor was there any prospect of continued employment after the program ended. Tutors were assigned to target remedial education schools based on geographic proximity to their residential location.
Once selected, tutors participated in a training workshop organized by Universidad Cayetano Heredia. The two education specialists led the workshop. The workshop took place before the start of the 2014 school year and lasted 20 hours, split over 6 days. The general goal of the workshop was to train tutors in the pedagogical and didactical foundations of inquiry-based learning. As such, tutors were encouraged to incorporate seven principles into each session: (a) learning builds on prior knowledge, (b) learning is a restructuring of prior knowledge, (c) learning takes place in the interaction with the object of study, (d) learning requires language and communication, (e) emotions affect learning, (f) learning is a social as well as a psychological process, and (g) learning requires self-regulation (meta-cognition).
Tutors were instructed on possible approaches to apply these foundational principles to each of the tutoring activities to engage students. Some of these approaches include encouraging and discussing different points of view, sequencing contents to follow the children’s logic and applying new perspectives to unfamiliar situations. In the workshop, the specialists and tutors also reviewed the content and activities for each session.
During the workshop, tutors received an instructional guide summarizing principles, pedagogical approaches, and activities for each tutoring session. The two specialists also provided ongoing support to tutors during the implementation of the remedial sessions.
Selection of Participating Students
Third-grade students from 48 public elementary schools in Metropolitan Lima, who scored in the bottom half of their school’s baseline assessment score distribution, were eligible for remedial science education (sample selection details below). Baseline performance was assessed through a written test administered during class in May 2014. Among eligible students, a subset was randomly assigned to receive the intervention, which was an offer to attend the remedial sessions (details below).
As Figure 1 shows, there is considerable across-school variation in the halfway cutoff point that determines eligibility. For the majority of schools, the halfway cutoff point is around the 50th percentile of the overall science baseline score distribution. There are, however, schools in which the halfway cutoff is closer to the 25th percentile of the overall distribution and others in which the cutoff is closer to the 75th percentile of the overall distribution.

Distribution of baseline science test cutoff score by school.
The variation in cutoff points across schools does not represent a threat to interval validity. This variation, if anything, helps toward allaying external validity concerns because it suggests that results represent an average across a heterogeneous population. In this context, findings should therefore be interpreted as expected achievement gains for an average low-performing student in an average sample school—one in which the halfway cutoff is close to the 50th percentile of the overall student distribution. However, heterogeneity across schools in the eligible student population may lead to heterogeneous treatment impacts because, as we explain below, experimental variation in treatment status is all within-school. Note that this within-school assignment rule is akin to that of other experimental evaluations of educational interventions such as tracking by school-specific ability in Kenya (Duflo et al., 2011) and remedial tutoring within schools in India (Banerjee et al., 2007). In the empirical section, we explain how we test for heterogeneity in treatment effects across various levels of baseline school achievement.
Implementation of Remedial Sessions
Remedial sessions took place in each of the 48 participating public elementary schools in Metropolitan Lima. There was a total of 70 remedial groups. All remedial groups within a school were assigned to the same tutor, so there is no variation in tutor quality within schools. Each tutor was assigned on average to five tutoring groups (some as few as three and some as many as seven). There were more tutoring groups than schools because some of the schools had very large third-grade classes or more than one third-grade classroom (Table 1). Anywhere between 3 and 17 students were assigned to each tutoring group (always in the school they attended), with a mean group size of 9 students. The variability in group-size was a function of (a) the size of the third-grade class in participating schools (some schools had small student populations), and (b) scheduling conflicts because some schools were unable to accommodate sessions for multiple groups.
Sample School Statistics
Note. Table shows descriptive statistics for a total of 2,399 eligible and noneligible students from the 48 schools in the final sample.
Evidence on the optimal duration of remedial education programs is mixed. One meta-analysis found that the most important learning effects are achieved from remedial education programs of at least 16 hours of duration (Jun, Ramirez, & Cumming, 2010). Ultimately, the duration of the Peru remedial science education program was only partly based on what the prior literature considers to be the optimal length. Equally important were logistical and operational constraints, which were discussed in advance with school leaders based on group size, scheduling restrictions, pace, and what was feasible among third-grade students in Lima. Hence, although the planned format consisted of 16 weekly 90-minute sessions, some remedial groups met more times for shorter sessions. The planned format of 16 sessions provided treatment students up to 24 hours of additional remedial instruction, a 14% increase in total instructional time relative to the regular science schedule over the school year. Remedial sessions began in July 2014, halfway through the Peruvian school year, which begins in March. Appendix Table A1, available in the online version of the journal, provides a timeline of program and evaluation activities.
Remedial sessions took place at each school’s premises. Most sessions were scheduled in the afternoon (at the end of the school day). In a few cases, sessions were scheduled in the morning (for students attending school in the afternoon) or on Saturdays. In the first session, students received a workbook called “Making and Learning Science,” which describes various scientific inquiry activities that students could pursue independently.
Each tutor was responsible for coordinating and scheduling sessions with his or her groups. Tutors initially approached school principals and third-grade teachers to explain program details, seeking support to promote attendance of eligible students. Tutors also invited parents of eligible students to tutor information sessions to explain the goals of the tutoring program, the approach, and the expected benefits.
To ensure that all parents were informed about the availability of the remedial science education program, students were also asked to bring home an information sheet that parents were supposed to sign and send back. Some tutors also visited the students’ homes to deliver the information sheet. In total, about 50% of the parents of students assigned to remedial science sessions signed and returned these forms. This suggests that at least 50% of parents knew about the availability of the program for their children. The take up rate at the student level is discussed below.
Evaluation Sample, Experimental Design, and Student Characteristics
Evaluation Sample
We collected baseline test score data from third-graders in 51 public elementary schools in Metropolitan Lima in May 2014 to determine student eligibility for the program. Of these 51 schools, 39 had participated in the 2012 Science Education Teacher-Training Program. We chose these 39 schools to facilitate access to the tutors, as these schools had prior contact with the training staff from Universidad Cayetano Heredia. The remaining 12 schools were randomly chosen among comparable schools in the poorest localities in Metropolitan Lima. After baseline data collection, we discarded two schools because they had less than eight third-grade students, thereby minimizing the risk of stigmatizing one or two students with eligibility for participation. We further discarded one school because we were unable to contact tutoring-eligible children. The final evaluation sample is, therefore, drawn from the remaining 48 public elementary schools in Metropolitan Lima.
The average school in the sample has a remedial education eligibility cutoff at the 50th percentile of the overall baseline science score distribution, and the median school has a cutoff at the 49th percentile of the overall distribution (Figure 1). The average school in the evaluation sample has three third-grade classrooms (minimum = 1, maximum = 6); the average classroom has 24 students (minimum = 6, maximum = 38; Table 1). The principal of the average school in the sample has 6.3 years of experience as school principal and teachers have 5.6 years of experience, 4.5 years of which are in the current school. About 8% of sample teachers had participated in the 2012 Science Education Teacher-Training Program.
Experimental Design
In the baseline test, we assessed a total of 2,399 third-grade students in the 48 schools of the evaluation sample. The science, mathematics, and Spanish language sections of the test were simplified versions of the tests administered as part of the evaluations of the 2010 and 2012 science pilot programs implemented in Lima. The baseline test was only used for the purposes of the present evaluation. The test was developed to measure third-grade skills based on Peru’s 2008 new basic education curriculum and national study plan for third-grade mathematics, Spanish, and science. In science, the third-grade curriculum includes the Physical World and Preservation of the Environment, the Human Body and Health, and Animals and their Environment.
Test questions addressed a mixture of content and critical thinking skills. Content questions included, for example, questions about how different food groups can help us stay healthy, and the identification of Peruvian animals. As an example of a critical thinking question, students were asked why a snow cone turned into red water when a little girl left it on a bench while playing. A supervisor monitored and timed the students as they individually completed the learning test in writing.
A total of 1,219 students were eligible for the program. Among these 1,219 eligible students, a subset was randomly assigned to receive the intervention, which was an offer to attend the sessions. We stratified randomization by school and gender. In practice, we only had 95 lotteries (48 × 2 − 1) because in one school only boys were eligible. In the final evaluation sample, we have 609 students assigned to treatment (331 boys and 278 girls) and 610 students assigned to business as usual control conditions of regular classroom instruction (337 boys and 273 girls).
Baseline Characteristics of Eligible Students
About 46% of eligible students are girls. Average age among eligible students is 8 years. Just over 90% of eligible students attend school during the morning shift. In the average eligible student’s household, there are close to two adults present, and 83% of them have a father present in the sample. At baseline, eligible boys and girls score at comparable levels in science, mathematics, and reading. These characteristics are balanced across eligible students assigned to treatment and those assigned to control. Standardized differences are small (0.01–0.04 SD for sociodemographics; 0.00–0.13 SD for baseline scores). We cannot reject equality across randomization groups of student characteristics separately or jointly (Table 2).
Sample Baseline Student Characteristics
Note. Table shows results of raw mean comparisons (i.e., not adjusting for the stratified research design) across students assigned to remedial science education and to control conditions. Sample is 1,219 third-grade students who score in the bottom 50% of the baseline science test administered in May 2014 to third-grade students in 48 public elementary schools in Metropolitan Lima. Joint significance test corresponds to F-stat of joint hypothesis that coefficients in Panel A are all zero.
Data and Empirical Strategy
Data
We use three data sources to document program impacts. The first data source is the baseline test and sociodemographic questionnaire collected from third-grade students in the 48 schools in the sample, designed and validated for use among elementary students. For example, it did not include questions about income, but rather about the dwelling of the household.
The second data source is student attendance records collected by the tutors at remedial group meetings (i.e., compliance with treatment assignment). As Figure 2 shows, however, the remedial session attendance data are incomplete. For students in some remedial groups, we only have student attendance records for less than five sessions. For the majority of students, we have attendance records for at least six sessions, including for some of those remedial groups that met more than 16 times in sessions of less than 90 minutes of duration.

Distribution of number of sessions per assigned student with recorded student attendance.
To estimate Treatment-on-the-Treated (TOT) impacts of remedial participation, we deal with missing attendance data by computing student attendance bounds. Specifically, we compute a lower bound on student attendance by assuming that students were absent from sessions without attendance records (worst-case attendance bound). For example, if a student attended 5 sessions and the tutor only recorded attendance data for those 5 sessions, in the worst-case attendance bound we assume that the student was absent in the remaining 11 planned remedial sessions. Similarly, we compute an upper bound on student attendance by assuming that students were present from sessions without attendance records (best-case attendance bound). In the example, the best-case bound would assume that the student was present in the remaining 11 remedial sessions.
The final data source is endline test and student survey data, which were collected exclusively for this evaluation in November 2014, about 5 months after the start of the tutoring sessions (see online Appendix Table A1 for timeline of activities). Achievement outcome measures with subject-specific sections were collected in the same manner and conditions for the treated and control students. Among eligible students, about 92% took the endline test, resulting in an overall attrition rate of 8%. Differential attrition between students assigned to treatment and to control was one statistically insignificant percentage point. We only exclude from the experimental analysis sample students with missing endline test scores. Characteristics of students who take the endline test are balanced across treatment and control, and standardized differences are small and statistically insignificant (see online Appendix Table A2). In addition, sample attrition is uncorrelated with treatment status (online Appendix Table A3). On the whole, the study meets conditions for high evidence standards, including treatment assignment determined through a random process, outcome data collected in a comparable manner for treatment and control groups, a combination of overall and differential attrition rates that is well within tolerable limits for threats of bias, and sample exclusions that are uncorrelated with treatment status (Institute of Education Services [IES], 2017).
Empirical Strategy
Definitions
In our analysis, it is useful to define three student populations of interest. The first group consists of eligible treatment students (“eligible treatments” henceforth), who are the students who score below the school-specific halfway eligibility cutoffs and are randomized into treatment (remedial education). The second group consists of eligible control students (“eligible controls” henceforth), who are students who score below the school-specific eligibility cutoffs and are randomized into control conditions (business as usual). The third group consists of ineligible students, who are student who score above the school-specific eligibility cutoffs.
To conceptualize our empirical strategy, it is also helpful to define direct, indirect, and total effects of the remedial education intervention (Hudgens & Halloran, 2008; Sinclair et al., 2012; VanderWeele et al., 2013). The total effect is the sum of the direct and indirect effect. The direct effect is the difference, all else constant, between potential outcome for a student under treatment, compared with potential outcome under control. The indirect effect (i.e., interference) is the difference between potential outcome for an untreated student when most students in her classroom receive remedial education treatment and potential outcome for that same student when few students in her classroom receive treatment. Without interference, the direct effect is the total effect of the treatment.
Experimental identification and estimation of the direct effect of remedial education on achievement
Identification of the direct effect of remedial education on student achievement requires that potential student outcomes are orthogonal to remedial education assignment. Under student-level randomization of treatment, the direct effect of remedial education on student achievement in our setting is the mean difference in outcomes between eligible treatments and eligible controls. Under no interference, the experimental estimate of the direct effect is also the total effect of treatment.
To estimate the direct effect of remedial education on achievement, we begin by showing unadjusted mean differences in outcomes between eligible treatments and eligible controls. Our preferred models, however, are test-score value-added specifications of the following form:
where
We present TOT estimates for the direct effect of remedial education for students who attended at least one remedial session. 3 We estimate lower and upper bound versions of TOT, depending on whether we use as the endogenous regressor, respectively, the upper bound or lower bound estimate for whether a student attended at least one remedial session. We estimate TOT in the experimental data via two-stage least squares.
SUTVA: Nonexperimental identification and estimation of the indirect effect of remedial education on achievement
As noted, when remedial education assignment of a student only affects her achievement through her participation in the remedial sessions, the total effect of remedial education equals the direct effect because the indirect effect is zero (i.e., there is no treatment interference). The no-treatment interference assumption or SUTVA (Rubin, 1980) may be challenged if, for example, nonparticipants benefit indirectly through improved regular classroom learning as a result of a lower fraction of underperforming students delaying the pace of learning. Although SUTVA is ultimately untestable, one of our methodological contributions is to introduce an empirical approach to indirectly test SUTVA that exploits naturally occurring variation in the fraction of treated students within classrooms. Specifically, for schools with more than one third-grade classroom, our research design creates (nonexperimental) variation within classrooms in the fraction of students receiving treatment because randomization stratifies treatment assignment by school and gender but not by classroom. Recall that the average sample school has three third-grade classrooms, with a standard deviation of 1.2 classrooms and a maximum of six third-grade classrooms (Table 1). Given that many schools in the sample have multiple third-grade classrooms, there is in our sample nonexperimental variation in the classroom-level proportion of treated students, as Figure 3 shows. In some classrooms, no students are assigned to receive remedial education. There are several classrooms in which anywhere between 20% and 60% of students are assigned to remedial sessions. In one classroom all students are assigned.

Variation in the fraction of students in a section assigned to participate in the targeted, inquiry-based remedial science program.
Our approach to estimate the indirect effect—an ancillary test of SUTVA—consists of a conditional contrast of achievement outcomes for untreated students in classrooms with high and low proportions of treated students. Identification of the indirect effect requires an assumption that the fraction of treated students in a classroom is orthogonal to classroom potential outcomes. This would be the case, for instance, in a multilevel randomized intervention in which teachers and students are randomly assigned to classrooms, classrooms are assigned to a fraction of students treated, and within classrooms, students are randomly assigned to treatment or control (Miguel & Kremer, 2004; Sinclair et al., 2012).
Without multilevel randomization, we obtain identification of the indirect effect of remedial education by assuming that potential outcomes of nonparticipants are orthogonal to the fraction of students treated in the classroom, conditional on covariates—a group-level version of the selection on observables assumption (Rosenbaum & Rubin, 1983). Specifically, under a linear-in-means peer-effects model, if (positive) spillovers exist, student achievement should be higher in classrooms with a higher fraction of students assigned to remedial sessions. We estimate this indirect effect by modifying Equation (1) as follows:
where
Empirical validation of a nonexperimental estimator of the effect of remedial education on achievement
The remedial education program’s eligibility rule generates an alternative nonexperimental research design because eligibility varies discontinuously as a function of a student’s baseline science percentile achievement. Students whose percentile score is below the school’s 50th percentile are a mix of approximately 50% eligible treatments and 50% eligible controls, given that about half of eligible students were originally randomized to receive remedial education. Students to the right of the cutoff are ineligible students, none of whom participates in treatment. We leverage this nonexperimental variation in an RDD framework to estimate the effect of remedial education eligibility on achievement for students on the margin of their school’s halfway cutoff point. Because cutoffs are school-specific, we follow common practice and normalize scores across schools to be zero at the 50th percentile of the school-specific baseline science achievement distribution. Given the cutoff normalization, the RDD estimator represents the weighted average across cutoffs of the local average treatment effect for all units near each particular cutoff value (Cattaneo, Keele, Titiunik, & Vazquez-Bare, 2016). In this RDD framework, we limit our sample to include only all eligible treatments and all ineligible students to the right of the cutoff (i.e., we exclude eligible controls). In this sample, we estimate ITT effects of remedial education using the following sharp RDD model:
where
Results
In this section, we present results on student attendance, direct effects on endline achievement, heterogeneity, indirect effects, and, finally, nonexperimental RDD estimates.
Student Attendance to Remedial Science Education Sessions
Figure 4 shows bounds on student-level attendance for each of the program’s scheduled remedial session. The bounds diverge over time, indicative of tutors being more diligent at collecting student attendance data at the beginning of the program than toward the end of it. Our bounds suggest that somewhere between 40% and 60% of students attended the first few initial remedial sessions. Halfway through the program (Remedial Sessions 6–10), on average between 25% and 75% of students were attending sessions. Toward the end of the program, when most student attendance to remedial sessions is missing, attendance bounds are fairly wide and, consequently, not very informative, because they suggest that anywhere between 5% and 95% of students assigned attended remedial sessions nearing program completion.

Bounds on student-level attendance to remedial sessions.
Table 3 shows mean lower and upper bound estimates for various metrics of student attendance averaging over the remedial program’s duration. Lower bound estimates of attendance suggest, for example, that 75% of eligible treatments attended at least one remedial session (columns 1 and 2, Panel A, Table 3). In the lower-bound scenario, eligible treatments attended, on average, between four and five remedial sessions (columns 1 and 2, Panel B, Table 3), corresponding to about 28% of the 16 total sessions initially planned (columns 1 and 2, Panel C, Table 3), or an additional 425 minutes of supplemental instruction time (columns 1 and 2, Panel D, Table 3).
Lower and Upper Bounds of Student Attendance to Remedial Sessions, OLS Various Metrics
Note. Standard errors clustered at the school level in parentheses. All models are OLS regressions. Columns 1 and 3 have no controls; columns 2 and 4 include full controls for baseline scores, stratification fixed effects and other student sociodemographic characteristics not shown in the table including school shift, Spanish speaking, adults in household, and father present in household. Lower bound obtained by assuming all students were absent in sessions with missing attendance data. Upper bound obtained by assuming all students were present in sessions with missing attendance data. OLS = ordinary least squares.
p < .1. *p < .05. **p < .01.
Upper bound estimates of student attendance suggest that 96% of eligible treatments attended at least one remedial session (columns 3 and 4, Panel A, Table 3). In the upper-bound scenario, eligible treatments attended, on average, about 11 remedial sessions (columns 3 and 4, Panel B, Table 3), corresponding to about 65% of sessions initially planned (columns 3 and 4, Panel C, Table 3), or an additional 1,015 minutes of supplemental instruction time (columns 3 and 4, Panel D, Table 3).
As Table 3 also shows, compliance with treatment assignment among eligible controls was very high. On average, eligible controls attended 0.04 remedial sessions (Panel B, Table 3) or alternatively, received 4 to 5 additional minutes of total remedial instruction time (Panel D, Table 3).
The low and declining continued attendance to the remedial program among eligible treatments is likely the result of a combination of factors. Absences at the beginning of the program could be due to a failure to effectively promote the program and its benefits among students and parents. Students may also have time conflicts with other responsibilities, as 16% of Peruvian 8- to 9-year-olds are economically active, generally combining school with work. Although child labor is 30 percentage points more prevalent in rural areas (Instituto Nacional de Estadística e Informática [INEI], 2015), children in urban areas are also economically active, mainly as street vendors (International Labor Organization [ILO], 2009). Moreover, children may need to help at home in the afternoon or during weekends by taking care of younger siblings while their parents are working. Diminished attendance toward the end of the program could be due to—at least in part—a group of schools scheduling the remedial science sessions later in the afternoon to prioritize the use of the tutors for test preparation sessions with higher grades.
Endline Achievement of Remedial Education Participants
Assignment to the inquiry-based remedial science program increases endline science scores. When measured in percentiles of the test score distribution, the ITT estimate for the direct effect of treatment is 3 percentiles in the preferred specification with the full set of control variables (column 2, Panel A, Table 4). When measured in SD units, the estimate of treatment assignment is 0.12 SD in the preferred specification (column 4, Panel A, Table 4). These estimates are similar with and without control variables. Conclusions about the statistical significance of these estimates are robust to clustering standard errors by tutor instead of school. Clustering by tutor is an even more conservative estimator of effects’ variance as each school only had one tutor assigned to all its groups and some tutors had groups in more than one school (online Appendix Table A4). 4
Experimental Estimates of the Direct Effect of Remedial Science Education Participation on Endline Science Achievement
Note. Standard errors clustered at the school level in parentheses. Table shows science endline impact results. In columns 1 and 2, outcome variable and lagged test-score regressor are expressed in percentiles. In columns 3 and 4, outcome variable and lagged test-score regressor are expressed in SD units. Columns 1 and 3 do include additional controls; columns 2 and 4 include controls for baseline scores, stratification fixed effects and other student sociodemographic characteristics not shown in the table including school shift, Spanish speaking, adults in household and father present in household. Panel A, Intent to Treat (OLS); Panel B, Treatment on Treated (Upper Bound, 2SLS); Panel C, Treatment on Treated (Lower Bound, 2SLS). Lower bound obtained by assuming all missing sessions were attended. Upper bound obtained by assuming all missing sessions were missed. Treated is defined as at least attending one tutoring session. OLS = ordinary least squares; 2SLS = two-stage least square.
p < .1. **p < .05. ***p < .01.
Panel B and Panel C of Table 4 show 2SLS TOT estimates for the direct effect of treatment with binary treatment compliance. In this binary case, compliance is measured by whether eligible students attended at least one remedial education session. Because between 75% and 95% of eligible treatments attended at least one session—depending on assumptions about attendance of sessions with missing attendance records—and there are very few crossovers among eligible controls, it is not surprising that these TOT estimates are very similar to the ITT estimates. For students who attended at least one session, the upper bound TOT estimate of participation suggests that participation increased endline science achievement by 4 percentiles (column 2, Panel B, Table 4) or 0.16 SD (column 4, Panel B, Table 4). The lower bound TOT estimate suggests that for students who attended at least one session, participation increased endline science achievement by 3 percentiles (column 2, Panel C, Table 4) or 0.13 SD (column 4, Panel C, Table 4). 5
Heterogeneity in Endline Science Achievement Impact Estimates
Gender
The effects of the inquiry-based remedial science program on endline science achievement are entirely driven by gains among boys. For girls, both ITT and TOT program effect estimates are small in magnitude and not statistically significant (columns 1–4, Table 5). For boys, assignment to remedial sessions increases science scores by about 5 statistically significant percentiles (Panel A, columns 5 and 6, Table 5) or 0.19 to 0.22 SD (Panel A, columns 7 and 8, Table 5). The implied TOT estimate for boys who attend at least one remedial session is bounded between 5.5 percentiles (Panel C, column 6, Table 5) and 7.1 percentiles (Panel B, column 6, Table 5), or between 0.23 SD (Panel C, column 8, Table 5) and 0.29 SD (Panel B, column 8, Table 5).
Gender Heterogeneity in Remedial Education Impacts on Endline Science Scores
Note. Standard errors clustered at the school level in parentheses. Table shows heterogeneity in science endline impact results by gender. Columns 1 to 4 show results for female students, columns 5 to 8 for male students. In columns 1 and 2 and 5 and 6, outcome variable and lagged test-score regressor are expressed in percentiles. In columns 3 and 4 and 7 and 8, outcome variable and lagged test-score regressor are expressed in SD units. (1), (3), (5) and (7): no controls except for stratification fixed effects; (2), (4), (6) and (8): controls for baseline scores, stratification fixed effects, and other student sociodemographic characteristics not shown in the table including school shift, Spanish speaking, adults in household, and father present in household. OLS = ordinary least squares; 2SLS = two-stage least square.
p < .1. **p < .05. ***p < .01.
One possible explanation to the impact heterogeneity by gender is differences in treatment intensity (compliance) between boys and girls. We do not find empirical support for this conjecture. Eligible boys and girls are equally likely to attend remedial sessions according to various metrics of attendance (online Appendix Table A7).
We performed additional tests to explain this finding. First, we analyzed the gender composition of the overall sample (that is, the sample not restricted to eligible students) and found it does not differ statistically from that of the experimental sample. Thus, the gender differences between low performers and high performers cannot explain the differential effect for boys and girls. Second, we analyzed the attendance patterns by session between boys and girls and could not find any significant variation in attendance patterns by gender, suggesting that differential impacts are not driven by differences in which sessions boys and girls attended. Third, we analyzed the distribution of test scores at baseline by gender. We found that the distribution of scores in science and reading is skewed to the left for boys, suggesting that boys have lower performance at baseline. Finally, we estimated heterogeneity of tutoring intensity controlling for quantiles of the baseline score interacted with the female dummy and found similar results to those presented in Table 5, suggesting that differential impacts are not driven by gender differences in achievement at baseline.
Remedial Session Group Size
To explore heterogeneity by group size, we modify regression equation (1) by replacing the treatment indicator variable with two mutually exclusive indicators of whether eligible treatments were assigned to a small remedial science education group (10 or less students) or a large group (11 or more students). The omitted group is eligible controls. Remedial group size, however, was not randomly assigned, so this evidence is an observational comparison. One candidate confound, for example, is tutor quality, to the extent it is correlated with remedial education group size. With this caveat in mind, the evidence is consistent with the greatest achievement gains accrued by students assigned to small remedial groups who score, on average, 5 statistically significant percentiles higher (0.23 SD) on the endline science test than their peers assigned to the control group. Students assigned to large remedial groups saw considerably smaller gains (1.8 percentiles/0.05 SD), none of which are statistically significant (Table 6). This pattern of results is potentially consistent with the hypothesis that hands-on, inquiry-based remedial education approaches are more effective through smaller groups to the extent that these facilitate greater student participation. However, we caution against overinterpreting the evidence, as we cannot formally reject in any specification statistical equality of treatment effects among students assigned to small and large remedial groups at conventional significance levels (second to last row, Table 6).
Intent-to-Treat Remedial Education Impacts on Endline Science Test Scores, by Initial Remedial Group Size (OLS)
Note. Table shows OLS results for a test of treatment effect heterogeneity by remedial group size. Results are based on a modified version of Equation (1) in which we replace the indicator for being assigned to remedial sessions with two mutually exclusive indicators for being assigned to a small or a large remedial group. The omitted category continues to be students randomly assigned to control. Standard errors clustered at the school level in parentheses. In columns 1 and 2, outcome variable and lagged test-score regressor are expressed in percentiles. In columns 3 and 4, outcome variable and lagged test-score regressor are expressed in SD units; (1) and (3): no controls except for stratification fixed effects; (2) and (4): controls for baseline scores, stratification fixed effects, and other student sociodemographic characteristics not shown in the table including school shift, Spanish speaking, adults in household, and father present in household. The last row shows the p value resulting from testing the null hypothesis that the coefficients for the two group size categories are equal. OLS = ordinary least squares.
p < .1. **p < .05. ***p < .01.
Baseline School Achievement
To test for heterogeneity in impact estimates by schools’ baseline achievement, we estimate Equation (1) separately for each school in the sample to obtain school-specific impact estimates on endline science achievement. We then estimate a random-effects intercept and slope model at the school level in which we meta-regress each school’s science endline impact estimate on a constant and each school’s average baseline science achievement. Each school’s impact estimate is weighted in proportion to its estimation precision (i.e., by the inverse of the estimate’s variance). Figure 5 shows results of this approach. Each circle in the figure is a school-specific impact estimate, with size proportional to estimation precision. The trend line depicts the estimated slope for association between effect sizes and baseline school achievement. As Figure 5 shows, there is no systematic relationship between remedial education effect sizes at the school level and schools’ baseline achievement (Figure 5a shows results in percentiles; Figure 5b in SD). The random-effects slope coefficient on baseline school achievement is close to zero and is not statistically significant. Moreover, the p value shown in Figure 5 indicates that we cannot reject the null hypothesis of no residual heterogeneity in the distribution of school-level remedial education effect estimates. These results suggest the possibility that remedial science education was equally effective across the distribution of science school achievement at baseline. 6

School-level treatment effects (ITT) of remedial education on endline science achievement by school’s baseline science achievement.
Estimates of the Indirect Effect of Remedial Education on Nonparticipants
We can take advantage of nonexperimental variation to estimate indirect effects on nonparticipants’ science achievement arising from classroom interaction with participants of the remedial program (Figure 3). The variation arises from naturally occurring variation in the classroom fraction of remedial education participants for schools with more than one third-grade classroom, because randomization stratifies treatment assignment by school and gender but not by classroom.
In general, estimates of the indirect effect are substantially smaller in magnitude relative to estimates for the direct effect, and are not statistically significant (Table 7). For instance, in the model with full controls, going from a classroom with no participants to one in which all students participate in remedial sessions is associated with a 2.3 percentile (column 2, Table 7) increase in endline science achievement of nonparticipants (0.06 SD, column 4, Table 7). However, in the average classroom the fraction of participants is 0.5 with a standard deviation of 0.12. Therefore, a more meaningful way to assess the magnitude of the indirect effect is to express it relative to a one standard deviation increase in the fraction of participants in a given classroom. In this case, the interpretation would be that, all else constant, a one standard deviation increase in the classroom fraction of participants is associated with an increase of 0.28 percentiles (2.3*.12) in the endline achievement of nonparticipants (0.007 SD = 0.06*.12). The estimates of the indirect effect on nonparticipants are orders of magnitude smaller than those of the direct effect and are not statistically significant.
Estimates of the Indirect Effects of Remedial Education on Nonparticipants’ Science Endline Achievement
Note. Standard errors clustered at the school level in parentheses. Table shows estimates of Equation (3) for the effect on nonparticipants’ endline science achievement as a function of the classroom fraction of students assigned to remedial sessions. In columns 1 and 2, outcome variable and lagged test-score regressor are expressed in percentiles. In columns 3 and 4, outcome variable and lagged test-score regressor are expressed in SD units; (1) and (3): no additional controls except for stratification fixed effects; (2) and (4): controls for stratification fixed effects and other student sociodemographic characteristics not shown in the table including school shift, Spanish speaking, adults in household, and father present in household.
Section averages obtained for each student excluding their own value, both for section fraction assigned and section average baseline science score. Section fraction assigned to remedial sessions has a mean of 0.50 and a standard deviation of 0.12.
p < .1. **p < .05. ***p < .01.
These results potentially support the SUTVA assumption, implying that the experimental estimates of the direct effect of remedial education participation possibly represent the program’s total effect on endline achievement. However, we caution against a strong interpretation of this result. Although the point estimate on indirect test-score gains is small in magnitude and not statistically significant, one caveat is that statistical power for the test could be low due to insufficient classroom-level variation in the fraction of students assigned to remedial sessions within schools.
RDD Estimates of Remedial Education Participation
Figure 6a shows in graphical form how the remedial education program’s eligibility rule generates a discontinuous relationship in the probability of attending at least one remedial session as a function of students’ normalized baseline science percentile achievement. Panel A of Table 8 shows the regression analog to Figure 6a. Using the optimal bandwidth and the lower-bound attendance measure, estimates indicate that eligible treatments have a 42 statistically significant percentage point greater chance of attending at least one remedial session than ineligible students above the cutoff. In the case of the upper-bound attendance measure, the difference at the cutoff in the probability of attending at least one remedial session is 53 statistically significant percentage points (column 1, Panel A, Table 8). Panel B of Table 8 indicates that eligible treatments just below the cutoff score 4.2 percentiles higher in the endline achievement test than ineligible students just above the cutoff. The RDD-based ITT impact estimate on endline achievement using the optimal bandwidth (N = 518) is not statistically significant. However, the magnitude of the point estimate is comparable to the experimental ITT estimate of 3.2 to 3.7 percentiles shown earlier in Table 4. Note that this need not be the case because the RDD estimator with multiple cutoffs represents the weighted average across cutoffs of the local average ITT effect for all units near each cutoff value, while the experimental ITT estimate represents the average effect of treatment assignment.

Regression discontinuity graphical evidence of remedial education impacts on endline science achievement as a function of eligibility cutoff.
Regression Discontinuity Estimates of Remedial Education Eligibility on Endline Science Achievement
Note. The group of students above the normalized school-specific percentile halfway eligibility value is the omitted group in the regression models used to estimate the coefficients reported in the table. Standard errors clustered at the school level in parentheses. In all models, the dependent variable is endline science test-score percentiles. Regressions include as additional control variables a nonparametric function of the normalized baseline science achievement science percentile score. Column 1 uses the default bandwidth of Calonico, Cattaneo, and Farrell (2017), designed to minimize mean square error in a sharp regression discontinuity design. Column 2 uses half the optimal bandwidth. Column 3 uses double the optimal bandwidth. ITT = intent-to-treat; TOT = treatment on the treated.
p < .1. **p < .05. ***p < .01.
At the optimal bandwidth, the implied lower-bound TOT estimate for students at the margin who attend at least one remedial session is 8 endline science test percentiles, whereas the upper-bound TOT estimate is 10 percentiles (column 1, Panel B, Table 8). These RDD-based TOT estimates are not statistically significant. The magnitude of these point estimates is roughly twice the magnitude of the experimental estimates of attending at least one remedial session (lower bound is 3–4 percentiles, upper bound is 4–5 percentiles; Table 4). However, the RDD estimates are imprecisely estimated and the confidence intervals comfortably include the experimental point estimates. Doubling the optimal bandwidth (N = 1,106) results in RDD point estimates that are perhaps even more comparable to the experimental estimates. The ITT estimate of eligibility at the cutoff is about 3 percentiles, whereas the lower and upper-bound TOT estimates of attending at least one session for students near the cutoff are, respectively, 4 and 5 endline science test percentiles (column 3, Panel B, Table 8). Doubling the bandwidth still results in RDD point estimates that are not statistically significant with wide 95% confidence intervals. Reducing the estimation bandwidth to include only 261 observations within one-half the optimal bandwidth results in negative point estimates and extremely wide 95% confidence intervals. Ultimately, the RDD estimates in our sample are statistically underpowered. This is perhaps not surprising because for an RDD to produce the same minimum detectable effect as an otherwise comparable randomized trial requires an RDD sample that is three to four times the size of experimental sample (Bloom, 2012). By contrast, in our setting the RDD sample size (optimal bandwidth) is one-half the experimental sample size. Overall, our interpretation of these RDD validation results is that the RDD estimator can potentially approximate the experimental estimates of remedial science education participation, but to statistically detect impact estimates of the magnitude of those in the experimental sample would require a substantially larger RDD estimation sample.
Discussion
Society and economies benefit from a scientifically literate population. In light of Latin America’s meager results on international standardized science assessments, it is important to identify learning models that help ensure that all children boost their scientific skills. The 2010 and 2012 primary science education pilots implemented in Peru helped improve science learning. Yet, the pilots revealed that one size does not fit all. Although the learning models that were piloted provided differentiated instruction according to the abilities of diverse groups of students, they were not effective for students who scored at the bottom of the distribution in the baseline assessment. These students needed extra help in the areas where they were struggling.
There is a growing literature on the effectiveness of remedial education, showing that remedial education based on direct instruction tends to be more effective than unstructured pedagogical approaches (Houtveen & van de Grift, 2007, 2012; Kaiser et al., 1989; Linan-Thompson & Vaughn, 2007). However, there is scant literature on remedial science education in elementary grades. In the absence of rigorous evaluations of remedial science education in elementary grades, we turned to research on what works in regular classroom science instruction in elementary grades, which points to the effectiveness of inquiry-based classroom practices. This raises the question of whether inquiry-based approaches can effectively be used also in remedial education to improve learning among low-performing students in elementary grades.
The results we present here are the first that measure the educational impacts of an inquiry-based, remedial science education program focused on low-performing students in early grades. It is also the first randomized experiment of a science tutoring program for small groups of low-performing students in Latin America.
Our results indicate that, although most eligible treatments participated in at least one remedial session, overall compliance among eligible treatments was not perfect, ranging, on average, from a worst-case scenario of attendance of 5 remedial sessions to a best-case scenario of 11 sessions. In this sample of low-performing students and relative to eligible controls, eligible treatments score 3 percentiles or 0.12 SD higher in an endline science test, with slightly greater gains for students who attended at least one session, and likely even greater for those with complete remedial session attendance. The direct effect of remedial education participation is, however, concentrated primarily among boys, with additional potential heterogeneity by tutoring group size. Our estimates of the direct effects of remedial education possibly capture the total effect of the program due to substantively small and not statistically significant indirect effects on nonparticipants, although this null finding is also consistent with limited statistical power to detect an indirect effect. What is clear, however, is that the remedial inquiry-based science education program improved science achievement among low-performing students who did not benefit academically from prior universal interventions focused on curricular standards and teacher training.
Our findings, therefore, suggest that low-performing students can learn through inquiry-based pedagogical approaches. This inquiry-based remedial science education model could easily be expanded to provide intensive academic support at a large scale for students who fall behind. Because the tutors are local and the training is short, replicating the project would be relatively straightforward. However, to bring this remedial science education model to scale, we identify two important challenges. First, the remedial inquiry-based science education program did not significantly improve learning among girls. The achievement gains are concentrated among boys eligible for treatment, for whom gains are 0.22 SD. These differential gains are not explained by gender differences in treatment compliance or baseline achievement. Several factors could have contributed to the concentration of gains entirely among boys. One conjecture is that the absence of effect among girls could stem from an explicit or implicit preferential treatment of boys by tutors, which would be consistent with prior evidence documenting how teacher beliefs and effort may exacerbate gender gaps in beliefs and competence in scientific endeavors (Fenema, Peterson, Carpenter, & Lubinski, 1990; Mendick, 2006). Another conjecture is that the absence of effect among girls could arise from differential behavior and engagement of boys and girls in small-group tutorials (Beuermann et al., 2013).
Second, the overall effectiveness of the remedial education model was achieved despite less than perfect student attendance to remedial sessions. The effect could potentially be greater with increased compliance. Achieving greater student attendance may require clearer dissemination to and greater buy-in from parents, teachers, and students. Because many students are either economically active or provide help at home by taking care of younger siblings, a more flexible schedule could also help improve the participation rate.
In terms of directions for future research, four areas stand out. First, although our findings indicate that remedial inquiry-based science education can improve learning among low-performing students in early grades, additional research would be required to determine whether this approach is more effective than tutoring based on a direct instruction methodology. Second, future research could shed light on some of the posited conjectures about potential sources of observed gender heterogeneity. Third, future research could experimentally investigate the extent to which the inquiry-based approach to remedial science education is, in fact, more effectively conveyed through smaller student groups. Fourth, future research could try to experimentally estimate within-classroom remedial education spillovers through multilevel experimental designs.
Supplemental Material
DS_10.3102_0162373719867081 – Supplemental material for Remedial Inquiry-Based Science Education: Experimental Evidence From Peru
Supplemental material, DS_10.3102_0162373719867081 for Remedial Inquiry-Based Science Education: Experimental Evidence From Peru by Juan E. Saavedra, Emma Näslund-Hadley and Mariana Alfonso in Educational Evaluation and Policy Analysis
Footnotes
Acknowledgements
We thank Triana Yentzen who provided outstanding research assistance. We thank the editors, two anonymous referees, and Richard Murnane for extremely helpful comments and suggestions. We thank the Innovations for Poverty Action (IPA) Peru team, especially Andrea Cornejo and Adam Kemmis Betty, for their invaluable support on the field.
Authors’ Note
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: We acknowledge financial support from the Japan Poverty Fund of the Inter-American Development Bank.
Notes
Authors
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
