Abstract
Randomized experiments are a strong design for establishing impact evidence because the random assignment mechanism theoretically allows confidence in attributing group differences to the intervention. Growth of randomized experiments within educational studies has been widely documented. However, randomized experiments within education have received criticism for implementation challenges and for ignoring context. Additionally, limited guidance exists for programs that are tasked with both implementation and evaluation within the same funding period. This study draws on a research team's experiences and examines opportunities and challenges in conducting a multisite randomized evaluation of an internship program for teacher candidates. We discuss how problems were collaboratively addressed and adjusted to align with local realities and demonstrate how the research team, in consultation with local stakeholders, addressed methodological and program implementation problems in the field. Recommendations for future research are provided.
Keywords
Randomized experiments offer a strong design for assessing an intervention's impact. The random assignment mechanism allows confidence in attributing group differences to the intervention itself. An experimental evaluation design (e.g., randomized field trial) is a strong design for estimating programs’ impacts in that the design helps to rule out other plausible explanations for change. Evidence-based practice within education has seen an uptick since the late 1990s (Bridges et al., 2009; Hammersley, 2007) with additional emphasis in the early 2000s. It was at that time that the U.S. Department of Education's Institute of Education Sciences began prioritizing research that applied experimental designs, with prioritization of randomized experiments as the highest level of quality of evidence (Whitehurst, 2001). With the passing of the Foundations for Evidence-Based Policymaking Act of 2018 (P.L. 115–435), federal agencies are now required to submit a systematic plan for identifying, as well as addressing, policy questions. This plan must include, among other components, appropriate questions, data, and methods. Additionally, each federal agency is now required to designate an evaluation officer. In doing this, evaluation within federal agencies has been elevated.
We concur with Julnes (2020) and acknowledge that randomized experiments are appropriate for some evaluation contexts but not all. More specifically, the aim of this article is to promote the understanding of the context of randomized field trials and the extent to which they can be successfully implemented in applied settings, particularly when challenges arise. Where applicable, we discuss the degree to which these are situated within considerations in planning and implementing randomized field trials in evaluation as proposed in the white paper presented by American Evaluation Association's (AEA) Evaluation Policy Task Force (EPTF) (as summarized in Julnes, 2020). There is a need to contextualize design and implementation challenges in conducting randomized field trials and share strategies and lessons learned. We accomplish this using our own experience with implementing a multisite trial of an intervention for preservice teachers.
How Randomized Experiments Work
The distinguishing feature of a randomized experiment is that access to the treatments are assigned to experimental units by chance, such as tossing a coin or assignment based on a random number table (Shadish et al., 2002). The particular strength that sets experimentation apart from other research designs is in causal description (Shadish et al., 2002). Randomized experiments are a key component in the realm of evidence and research that relates to identifying effective interventions (Craig et al., 2008; Demby et al., 2020; Torgerson & Torgerson, 2001). The underlying properties of randomization are: (1) the average difference between the experimental and control units is an unbiased estimate of the causal effect for the typical matched trial; and (2) probability statements can be made regarding the extent to which differences between the experimental and control units are unusual given a specific hypothesis (Rubin, 1974). Causal inference is possible in experimental designs because it assumes that the intervention is the only element that has differed systematically between the individuals in the groups. Readers interested in learning more about experimental designs are referred to classic texts such as Shadish et al. (2002). In the context of experiments designed to address policies or programs that are conducted in real-life settings, Julnes (2020) refers to these as “randomized field trials.” Because the terminology of randomized experiments varies across disciplines, we have retained reference to our Project as a randomized field trial in this article.
Randomized Experiments Within Educational Settings
Growth of randomized experiments within educational studies has widely been documented, with over 760 unique studies published in English journals and reports between 2006 and 2016 (Connolly et al., 2018). Randomized experiments are valuable and enjoy growing use; however, randomized experiments within education have received criticism for implementation challenges and for ignoring contextual circumstances (Connolly et al., 2017). Indeed, volume 31, number 8 (2002) of Educational Researcher was a themed issue of response comments to the National Research Council's report (National Research Council, 2002) on what constitutes scientifically based research in education, and many of the articles in that issue addressed the problem of ignoring context (e.g., Berliner, 2002; Erickson & Gutierrez, 2002), as have more recent publications (e.g., Jacob et al., 2019; Kim, 2019; Siddiqui et al., 2018). Volume 60, issue 3 (2018) of Educational Research was a themed issue on randomized experiments, included among the papers were those discussing challenges and opportunities within randomized experiments (e.g., Dawson et al., 2018).
Randomized experiments are common in health-related fields, and thus there is a fair corpus of examination of challenges of randomized experiments within the health-related context (e.g., Campbell et al., 2000; Conlon et al., 2020; Donovan et al., 2002; Lawton et al., 2011; Moore & Moore Graham, 2011). There are also studies that have detailed challenges of randomized experiments within education and the social sciences (e.g., Bell & Peck, 2016; Berliner, 2002; Connolly et al., 2018; Dawson et al., 2018; Dixon et al., 2014; Kaplan et al., 2020; Macintyre & Petticrew, 2000; Odom, 2021; Phillips, 2006; Scriven, 2008) as well as alternatives to randomized experiments (e.g., Chatterji, 2007; Cook et al., 2010; Odom, 2021).
Despite challenges, randomized experiments are still a desirable design in many circumstances, and evaluators can benefit from lessons learned by researchers who have overcome these challenges. Randomized interventions that are part of complex evaluations (Miller-Young & Poth, in press) may exist in quite messy contexts where meeting the parameters of a randomized experiment are challenging (Dixon et al., 2014). Thus, recommendations for projects that are simultaneously developing and evaluating complex interventions which apply a randomized experiment are needed. This is because there is little guidance on strategies for successfully implementing randomized field trials (Demby et al., 2020).
Many randomized experiments are considered “efficacy” studies, where the intervention is implemented in ideal conditions (i.e., nonroutine practice) with likely more support and involvement of intervention developers (Institute of Education Sciences et al., 2013). In comparison, effectiveness studies evaluate interventions in real-world, typical, routine conditions. The recommendations herein may be less applicable for researchers in the former situation and/or researchers evaluating programs that are not in real-world conditions. This article focuses on the evaluation of interventions that are effectiveness studies implemented in real-world settings (Institute of Education Sciences et al., 2013), where the intervention may not yet be fully developed prior to the study (Flay, 1986).
Purpose of the Study
This study examines program implementation and evaluation design challenges in conducting a multisite randomized field trial of a complex intervention, specifically, an innovative internship experience program for teacher candidates. For each challenge, we present the primary issues, provide a description of how the evaluation plan was adjusted, and discuss the degree to which the adjustments were successful. Where applicable, we discuss the degree to which these are situated within considerations in planning and implementing randomized field trials in evaluation as proposed in the white paper presented by AEA’s EPTF (as summarized in Julnes, 2020).
Context and Evaluation Design
Project and Setting
For blinding purposes, we refer generically to the “Project,” which funds the delivery and evaluation of a university's five-year federally funded teacher preparation program designed to recruit, prepare, and sustain teachers in high-need schools. The project supports new teachers in high-need urban schools within one district in a Southeastern state. We will briefly offer context on the intervention (additional details on the intervention are provided later) to assist the reader in better understanding the challenges we faced. The intervention is testing a new model of supporting teacher candidates. During internship, a teacher candidate is assigned to a supervising teacher (a full-time teacher-of-record at the teacher candidate's assigned school) whose role is to be the primary mentor and a university coordinator that, in conjunction with the supervising teacher, conducts formal assessments of the teacher candidate. The intervention replaces the university coordinator with a Professor-in-Residence who is embedded at the teacher candidate's school and delivers enhanced coaching and support to the teacher candidate. It is hypothesized that this is a stronger model than the traditional model of teacher candidate supervision.
As of the Project's third year, 62 teacher candidates had been placed in four elementary schools, one K-8 school, and one high school as part of their student teaching requirement. All partnering schools receive Title I funding (i.e., they have a large proportion of economically disadvantaged students). However, the schools vary in student demographics, characteristics, and more, which influenced the evaluation design and is one reason why the unit of assignment was the teacher candidate rather than the school. For example, enrollment ranges from 309 to 668 at the four elementary schools. At six of the seven schools, Black students constitute two-thirds or greater enrollment; the remaining schools serve about 70% Hispanic students. English language learners range from less than 10% at three schools to more than 20% at two. Proficiency levels also vary with six schools demonstrating between 24% and 35% of students proficient in English-language arts while one school has 51%. Finally, per pupil expenditures range from less than $10 K at two schools to a high of more than $13 K.
Research Design
The Project is a two-group (intervention and business-as-usual) randomized block design (study participants randomized within site) with pretest, posttest, and multiyear follow-up. The unit of assignment is the teacher candidate (referred to as “study participant”). Study participants are in their last one or two semesters in their degree program and are required to complete a student teaching internship in a K-12 classroom. As such, each study participant is assigned a supervising teacher (i.e., the K-12 teacher-of-record for the study participant's assigned classroom, referred to as “intermediary”) along with a university internship coordinator. Study participants are eligible for the Project during either or both internship I (the semester immediately prior to graduation) or internship II (the semester in which they graduate). Internship I study participants receive less “dosage” of the intervention as they only report to their assigned schools twice per week while enrolled in nine credit hours of corequisite courses. They may lead instruction once or twice per semester. In comparison, internship II study participants are in their assigned classrooms daily and are expected to provide a substantial amount of instruction.
The “dosage” of the Project (i.e., intern I only, intern II only, or both intern I and II) is a mediator and is applicable to both intervention and BAU study participants. A mediator is a variable between the intervention and the outcome and helps explain the outcome. Within sites, study participants are randomly assigned to the intervention internship or to the business-as-usual internship (descriptions follow). Blocking is a research design technique that can assist with controlling nuisance factors by randomizing units within blocks (i.e., strata). Nuisance factors, or noise that exists because the groups vary in context, are unavoidable in all experiments and may affect the results. For example, in our Project, sites and internship level (i.e., dosage) were potential nuisance factors. Within blocks, the nuisance factor is held constant, allowing the factor of interest (e.g., intervention) to vary. In other words, there is a mini experiment within each block.
Business-as-Usual. The primary difference between BAU and intervention is the coaching provided. In the BAU model, each study participant is assigned to a university internship coordinator who performs a minimum of four observations each semester. The internship coordinator's observations are passive, lasting the duration of the lesson the student delivers, with the internship coordinator completing a college-provided rubric. After the observation, the internship coordinator meets with the teacher candidate and their supervising teacher to discuss results.
Intervention. Study participants in the intervention have the same program requirements as the standard model (e.g., required observations). Beyond the BAU, the study participants in the intervention receive additional support from a Professor-in-Residence (referred to as “interventionist”) who is embedded at each site and spends up to one full day per week working directly with the study participants.
Randomized Field Trial Challenges and Solutions
We identify implementation and/or evaluation challenges that arose during the randomized field trial: (1) evolving intervention; (2) treatment diffusion; (3) difficulties in measuring treatment fidelity; (4) inadequate numbers of and complexities and difficulties in placing study participants; and (5) procedural barriers for block randomization. These are situated within the conceptual framework proposed by Weiss et al. (2013) (see Figure 1). The conceptual framework assists in understanding how the pieces of the Project fit together, variation in project effects, and from where the variation may be stemming (Weiss et al., 2013). After the Project was proposed and funded, there were implementation and evaluation challenges. Implementation challenges (e.g., evolving intervention, external factors that resulted in the need for adapting project components) occurred during development, can be described in terms of components or processes in the project logic model, and may be a normal part of project development. Evaluation challenges (e.g., treatment diffusion; difficulties in measuring treatment fidelity) are factors which limit the statistical conclusions and internal validity of causal inferences. What happens in the interim between providing the intervention and receiving the intervention is reflected in take-up factors (e.g., the Project challenge of an evolving intervention) Bloom & Bloom, 1981.

Conceptual framework.
We document the challenges, the adjustments made by the team, and the extent to which these adjustments succeeded or failed. Adjustments revolved around the implementation of the randomized design and the need for a collaborative, iterative, adaptive evaluation design to account for professional concerns and contextual circumstances of the partnering schools and district.
Challenge 1: Evolving Intervention
Description of Challenge
Situated within the conceptual framework (Figure 1), the evolving intervention was an implementation challenge that directly affected take-up, and thus must be considered a mediator when assessing the project effect. The Project was funded in Fall 2018 as a design and development grant with the blueprint for the intervention and implementation represented in the far left of Figure 1 as proposed intervention (Weiss et al., 2013).
Adjustment and Degree of Success
As noted by Dixon et al. (2014), “the intervention must be defined as an entity in its own right, which may be present or absent. This may seem like a straightforward task when, for example, comparing two different medicines. However, when interventions are complex, namely containing multiple, interacting elements, the conceptual task is greater. …If evidence-based interventions are to be repeatable in different settings, they must be standardized, hence the development of intervention protocols or manuals” (2014, p. 1567). It was this situation that we found ourselves. The intervention seemed straightforward, but it took multiple years (beyond the pilot semester and into the experimental phase) to formalize the intervention, create a standardized intervention manual, and finalize the delivery. This implementation process (Figure 1) reflects adoption of the intervention to varying degrees based on the context and characteristics of the organizations involved in implementation (Weiss et al., 2013). It was not until the Fall of year 3 that the intervention was implemented with more stability. Given the multisite nature of the Project, there are multiple interventionists that must be trained and recalibrated, annually, or more often, to ensure the intervention was, and continues to be, understood. (Detailed later are challenges and adjustments associated with fidelity.)
The lack of a clearly defined, well-implemented, and stable intervention is one potential threat that can limit the value of a randomized evaluation design (Julnes, 2020). The Project recognized that revisions to the intervention were needed and worked to make adaptations to the intervention based on the context of the study (recognizing that program evolution is inherent to a study of this type) to ultimately stabilize the treatment such that minimal adjustments were made to the intervention. Unfortunately, the situation in which we found ourselves is probably like many large-scale applied experiments implemented in the context of grant programs—where development and evaluation of the intervention are co-occurring. The evolving intervention is a program challenge and specifically represents instability in program core components and inputs. This instability, however, is a normal part of program development. Researchers conducting experimental studies in contexts such as this should understand that, once implemented, even the best designed intervention may require local adjustments to work in a given context (Berliner, 2002; Bonell et al., 2012; Kaplan et al., 2020; Michie et al., 2017). This does not, necessarily, equate to scrapping the study in its entirety or becoming one of the studies that lend criticism to experimental evaluation. Rather, it gives rise to consider how to continue to move forward successfully, guided by needed adjustments (evidenced by formative evaluation results) and examination in the analyses. For this Project, we recognize that during the first three years of the experimental design, the intervention was evolving. Such mediators are being considered as the data are analyzed (e.g., assessing impact intermittently, building up to the contrast in years 3 and beyond).
Challenge 2: Treatment Diffusion
Description of Challenge
Treatment diffusion is an implementation challenge that directly affects take-up. As noted previously, the implementation process (Figure 1) reflects adoption of the intervention to varying degrees based on the context and characteristics of the organizations involved in implementation (Weiss et al., 2013). This then influences what and how services are offered and the intervention ultimately received (Weiss et al., 2013). The bottom-half of the intervention contrast (Figure 1) is the intervention, or lack of intervention, received by those in the BAU condition. Understanding project effects requires understanding both parts of the intervention contrast (i.e., intervention and BAU), including understanding variation in what should have been provided (or not provided) and what was actually received (Weiss et al., 2013). Mitigating treatment diffusion is a priority for ensuring Project fidelity so that statistical conclusions and internal validity are not limited. The complexity that arises due to the interaction of intervention and context has been cited by others (e.g., Datta & Petticrew, 2013). There are multiple avenues by which treatment may be diffused in this Project including: (1) interventionist; (2) intermediary; and (3) study participants.
Interventionist (i.e., Professors-in-Residence, PIR). Treatment diffusion resulted not from ill intentions but the opposite—the desire of the interventionists to provide support, which is amplified in high-needs study sites. This broader context of the school environment in which the Project operates may moderate the effects of the Project (see Figure 1). Evidence of treatment diffusion surfaced through interviews, project meetings, and informal conversations with the interventionists. This conflict in the roles and responsibilities of a researcher with requirements to follow a standardized protocol, along with the roles and responsibilities of a practitioner, has been recognized as a challenge in randomized experiments (Lawton et al., 2011). Lawton et al. (2011) found that research staff sometimes addressed conflict by revising clinical practice and adapting protocol to align with clinical practices and experiences—which led to not following protocol (i.e., adapted protocol was misaligned to original protocol). Additionally, given the substantial amount of time the interventionist spends at the study site, they are seen as a resource by the administration and staff. It is not uncommon for an administrator to request an interventionist to work further with a study participant—including BAU.
Intermediary (i.e., Supervising Teachers). The supervising teacher acts as an intermediary (Feser, 2023) in that they are present during the intervention, and thus reasonable that they pick up on the intervention and may use it with study participants. So, although the intermediary is not the target of the intervention, the intervention may be diffused to them, and the intermediary in turn may be using it. To decrease this possibility of diffusion, intermediaries must stay in whichever condition (intervention or BAU) their initial study participant is assigned throughout all future assignments. However, study participant placement decisions are based upon multiple factors driven by both the site needs and study participant preference and is ultimately outside of the control of the Project. Thus, the Project cannot ensure consistent intermediary placements until there is a change in the way they and study participants are assigned.
Study Participants (i.e., Teacher Candidates). Treatment diffusion through study participants has always been a potential threat. As noted by Scriven, “all subjects must be informed that they are part of an experiment, and it takes a subject with an IQ of 95 or more about two days to work out whether she or he is in the experimental or the control group” (Cook et al., 2010, p. 108). Through the interventionists, we became aware of BAU study participants expressing the need for more support—a direct result of their observations of support provided to intervention study participants. Secondary diffusion, where participants in the intervention and BAU groups interact during the study, was also witnessed. For example, during intern orientation, one study participant created a Facebook group for all study participants—unintentionally mixing study participants from both conditions. There is further potential for diffusion through study participants during their cocurricular class discussions or informally before or after class.
Adjustment and Degree of Success
Mitigating the conflicting roles and responsibilities of the interventionists to bolster treatment fidelity has been difficult. As noted by Julnes (2020, p. 488), a tension exists between maintaining fidelity and the need to adapt due to contexts. The interventionists understand the importance of treatment fidelity and the role they play in adhering to treatment fidelity. This is regularly reinforced to the interventionists from the evaluation team, and the Project has developed a comprehensive manual for the interventionists that define their roles and responsibilities. Additionally, there are multiple ways in which the evaluation team is collecting data for potential diffusion. Interventionists complete an activity log documenting the types and frequencies of intervention components engaged in and with treatment condition. Findings are shared monthly among all the interventionists so that they can see changes in trend and how their work compares to the others and realign if necessary. Interventionists also complete a questionnaire and interview at the conclusion of each semester to gauge estimated diffusion. The Project diligently collects data to aid in understanding mediators that may mitigate Project effect. This highlights the need for continued discussions between stakeholders in complex interventions and the need for each party to understand the reasoning behind requests made. We do not anticipate that we will achieve 100% treatment fidelity due to the complexity of the role of the interventionist in the larger landscape of the study sites. We do, however, continue to strive for it and believe interventionists can operate differently within local adaptations while still maintaining fidelity (Wiegand et al., 2015). Comingled with this, the relationship that the interventionists are building with the study sites is beginning to forge a path for the Project to play a role in matching study participants to intermediaries to ensure that conditions are aligned. Although randomization at the site level may have averted issues with treatment diffusion within the sites, it was not a feasible alternative for this Project given the small number of sites to be randomized and the very different contexts of sites.
In terms of the resentment of BAU study participants toward the support received by intervention study participants, this may partly be mitigated by a more comprehensive consent process which is now implemented. The Project has also implemented overall communications with all study participants assuring them that although they will notice differences between intern experiences, all study participants will receive support needed to be successful. Finally, with support of the College, the Project has requested that a Project team member be the instructor-of-record for the internship course completed by study participants during their internship semester. In doing so, the Project can separate study participants by condition in their course and alleviate treatment diffusion that may occur in this setting.
Challenge 3: Difficulties in Measuring Treatment Fidelity
Description of Challenge
Treatment fidelity, an implementation challenge, represents the degree to which the intervention was provided as designed and influences the project effect. Thus, measuring treatment fidelity allows the determination of whether the key aspects of the project were delivered as intended (e.g., Yeaton & Sechrest, 1981). Results from treatment fidelity provide the opportunity for formative correction (Rossi et al., 2018) and enable evaluators to contextualize observed effects (e.g., the extent to which adherence to treatment fidelity relates to positive, negative, or neutral intervention findings) (Bellg et al., 2004). Difficulties in measuring treatment fidelity may limit the internal validity of causal inferences. It was not until year 3 of the Project that the intervention was fully defined and implemented. Because of the lengthy development of the intervention, there was a lot of time to consider how to measure and conduct pilot fidelity testing.
Adjustment and Degree of Success
Given the complexity of the Project, we examined multiple angles for measuring treatment fidelity recognizing that between intervention receipt and intervention outcomes were potential mediators (see Figure 1) for which it was important to measure the dosage, quality, and program differentiation of the intervention components (e.g., type of coaching, coaching quantity, and number of intended weekly professional learnings received). For all measures, the evaluation team worked in partnership with the Project implementation team to construct the tools (e.g., logs, self-report, observation checklist) which resulted in a number of benefits (e.g., using the same terminology as the implementation team; collecting data that would both accurately assess fidelity and be useful for Project staff) (McCormick & Maier, 2018).
Activity Logs. Logs can assist in understanding variation in implementation and are particularly helpful with multisite projects (Balu, 2017). Interventionists are asked to self-report all Project activities in which they engage using an activity log tracker, a very brief online questionnaire. The log captures the types of activities in which they engage as well as with whom (i.e., treatment condition). In theory, the logs capture every activity the interventionists complete. In practice, activities are moderately captured through the logs due to limited time, lack of a device to log as activities happen, and the inability to remember what they engaged in when they log after-the-fact. To encourage the interventionists to log activities as they are conducted, the Project supplied each interventionist with a tablet computer and stylus so they could quickly click through the log. But this modification may have come too late as they were already well established in their own routines.
Self-report by study participants. Study participants are blind to the condition to which they are assigned; however, all study participants are asked to report on the interactions in which they engage with their interventionist through online questionnaires and interviews. This feedback assists in understanding, from the perspective of the study participant, the extent to which there was treatment fidelity.
Observations. With the help of the Project implementation team, a checklist was developed to measure fidelity of a sample of interactions between the interventionist and study participants from both the treatment and BAU conditions. The observations have been integral to our understanding of the service take-up, the link between the intervention provided and the intervention received.
Informal means. To the extent possible, the evaluation team attends Project meetings that may provide insight into fidelity including, for example, monthly meetings of the interventionists. The meeting topics are geared toward implementation; however, the discussions provide the opportunity to pick up on aspects of implementation that are related to treatment fidelity.
Challenge 4: Inadequate Numbers of and Difficulties in Placing Teacher Candidates
Description of Challenge
Another challenge was insufficient numbers of study participants relative to the number targeted in the funded proposal, an additional practical value threat of applying randomized field trials to experimental evaluation (Julnes, 2020). This is both an evaluation challenge in that it has the potential to limit the statistical conclusions, specifically reflected as a project moderator (see “Project Characteristic” in Figure 1) as well as an implementation challenge in determining how to secure sufficient study participants.
One source that fed into this challenge is the general teacher shortage which is linked to a decline in the number of undergraduates that choose teaching as a career (Sutcher et al., 2016). The funded University also experienced comparable declines in the number of students selecting to major in teacher education (which equates to declines in potential study participants). Additionally, high-poverty schools tend to experience this at a greater rate with lower interest in teaching at high-poverty schools (Simon & Johnson, 2015). Inadequate numbers also resulted from fewer study participants being placed in partner schools than the number requesting placement. This resulted from a combination of the District's desire to meet the needs of nonpartner schools who wanted teacher candidates as well as too few supervising teachers within the study sites. As an example, Figure 2 illustrates that of nearly 40 potential study participants from the University who expressed interest in placement at a partner school during Spring 2020, only 25 were placed at a partner school. Thus, about 40% of interested study participants were not placed with the Project. Additionally, the teacher candidate placement process is a complex multiagency process where multiple considerations had to be taken. This program challenge is situated within the conceptual framework (Figure 1) as a partner district characteristic, arising during Project development and implementation.

Spring 2020 random assignment flowchart.
Adjustment and Degree of Success
The Project was faced with a market (i.e., a teacher pipeline) that had declined considerably from the time the proposal was submitted to when it was awarded. As such, the Project ultimately had to address a system change involving multiple separate entities, including the College, the Project, and the District. The developing relationship between the Project and the District was critical in obtaining buy-in of those that can affect change in terms of recruitment components (i.e., recruitment into teacher education programs throughout the College and recruitment of teacher candidates and supervising teachers specifically into the respective high-needs schools partnering with the Project).
A systems approach (Jackson, 2007) to tackling the dwindling numbers of study participants cannot be understated as it was through this approach that we were able to co-construct viable measures to increase participation. Additionally, it shed light on a larger problem (i.e., rapidly declining numbers of teachers) and brought together leadership from multiple avenues that are likely to influence long-term change.
Challenge 5: Procedural Barriers for Block Randomization
Description of Challenge
Situated within the conceptual framework, procedural barriers within the partner District and sites created an evaluation challenge for block randomization. These arose during Project implementation and had the potential to moderate the relationship between the intervention provided and project outcomes (seen as the bottom bar in Figure 1). The idea behind randomized block is that homogenous blocks are created within which nuisance factors are held constant. An easy way to think of randomized block designs is to consider the blocks as a set or collection of randomized experiments, with each block being part of the larger experiment. Block randomization should assist in achieving balance between intervention and comparison groups within blocks (i.e., schools and other blocking factors) and reduce the potential for bias and confounding. As proposed, blocking factors included: study site (multisites); type of participant (i.e., study participant or intermediary); intervention “dosage” (i.e., internship I vs. internship II); intermediary (i.e., supervising teacher) characteristics (e.g., type of certification); and in years 2–5, blocking within condition previously served by the intermediary. Each of these are explained as follows:
Study site blocks (i.e., intervention location due to multisite design). Blocking by site was introduced to assist in removing contextual effects. Schools serve as blocks such that at each school both intervention and business-as-usual study participants are assigned.
Type of participant. Blocking on type of participant within the school—study participant or intermediary—was proposed given the different roles each had within the classroom.
Intervention dosage. Study participants within a site were blocked on whether they were enrolled in intern I or intern II given the different amount of time each spent in and responsibilities within the classroom.
Intermediary characteristic. Within sites, blocking on intermediary (i.e., supervising teacher) characteristics (e.g., type of certification) was proposed. Within those blocks (e.g., intermediary holding full certification versus alternative certification), the intermediaries would be randomly assigned to receive either an intervention or BAU study participant. This blocking was intended to assist in removing differences in instructional abilities due to different training received by the intermediary.
Blocking within condition previously served by the intermediary. In year 2 and forward, once study participants were randomly assigned to condition, they would then be randomly assigned to an intermediary who had previously served in that condition. Thus, intermediaries who had supervised a study participant in the intervention condition in year 1 would be randomly assigned study participants in this condition throughout the remainder of the grant (and likewise for BAU). This was to assist with preventing treatment diffusion.
Adjustment and Degree of Success
We have successfully blocked within study site and generally had success in blocking within intervention dosage as randomization within these factors is under our control. The exception to the latter is for when there are only two or three study participants within a site, which makes additional blocking within dosage impossible.
Where we have not successfully blocked is within intermediary characteristic or within the condition previously served by the intermediary. The number of study participants enrolled by semester is insufficient to have an additional block based on intermediary characteristics. Given the challenge in reaching the targeted numbers, we have come to the realization that blocking within intermediary characteristic may not come to fruition due to procedural barriers. The takeaway from this is to thoroughly understand the processes and what is, or is not, under the control of the evaluation team.
In theory, the solution to blocking within condition previously served by the intermediary is for the Project to take over this task from the administrator. In reality, this is a process and goes back to understanding the systems and having the right people at the table to help make modifications. Project personnel have established good working relationships with the partner administrators, and this is the first step in having input into matching study participants and intermediaries. But we do not foresee this process being completely turned over to the Project. Thus, the next best case, which has been implemented in sites where administrators are agreeable, is to provide the administrators with lists of study participants and intermediaries to which we request matches to be made. In doing this, we can at least ensure that intervention study participants are matched only with intermediaries who had previously supervised intervention study participants (and similar for BAU). Because the matches by the administrators are as good as random, this is a solution to the lack of block randomization within intermediary condition.
Conclusions and Recommendations
This study experienced many of the same challenges that have made scholars critical of randomized experiments in education (Connolly et al., 2017), but we worked collaboratively to adapt to local contexts and realities. The Project completed year 4 at the conclusion of Spring 2022. Although the number of year 4 study participants was less than proposed (37 served in Fall and 38 in Spring; proposed was 150 total), enrollment continues to increase each semester and the District partner continues to show strong support including recommending additional study sites (two were added in year 4) and communicating with the Project to identify new ways to support the Project. As an example of additional partner support, the District developed a signing bonus incentive that was released in Spring 2022. The Project anticipates this will serve to attract study participants to request placement at a study site so they, too, can become eligible for the incentive (which will further increase the number of study participants in the Project).
We stand with Macintyre and Petticrew (2000) in the belief that practical difficulties in conducting randomized experiments in real-world settings can be overcome, and there is value in experimental evaluation in applied settings. Aligned with working to understand when and how these experiments are appropriate and effective (Julnes, 2020), this paper outlines some of the more challenging experiences in conducting a multisite randomized field trial and demonstrates adjustments to practical barriers to randomized field trial implementation. These may be useful to others conducting randomized experiments in educational settings, particularly those that may be in funded projects where development, implementation, and evaluation are occurring simultaneously.
Given our experience in conducting this multisite randomized field trial, the program and design challenges encountered in implementation, and the solutions—and sometimes failures—in addressing these challenges, there are several recommendations offered to researchers planning similar designs. In proposing an efficacy study, researchers should address most recommendations prior to applying for funding or at least prior to beginning the randomized field trial. For example, a strong partnership with partnering entities and a clear understanding of the procedures and policies that will impact implementation of the intervention should be in place before beginning the randomized field trial so that challenges are mitigated.
First, consider whether your project has sufficient conditions to support a randomized field trial and the extent to which potential threats may pose a design challenge (Julnes, 2020). Recommendations from the medical community should be heeded: There are many research designs available, with different designs appropriate for different questions and situations (Craig et al., 2008, p. 981). Even if a project meets all the conditions, there may be gaps when implemented. Rather than ditching the study, adapt to the realities, and document the challenges so that context to the findings can be provided. This may mean thinking creatively. There are generally ways to salvage a project, including design and analytic strategies not only to meet the challenges of a complex situation but we also recognize that not every situation is salvageable (i.e., there are sometimes fatal flaws). Strategies may include, for example, demoting the study to a quasi-experimental design, applying propensity score methods to match study participants in condition (even if the randomized design holds) or disaggregating participants in conditions to represent varying levels of dosage or treatment fidelity.
Next, know going into the project that the research design may not be implemented with textbook perfection. Give due diligence to considerations that will support an effective randomized experiment (Julnes, 2020), anticipate the need to assess and address hiccups in implementation, and understand issues that may be fatal flaws. This requires expertise of a team member in research design—someone who can design a rigorous randomized field trial and redirect as needed to stay on track and/or identify problems that escalate to fatal flaw. Integrating frequent monitoring and feedback systems (both qualitative and quantitative) into the study's implementation protocol from the start can also help quickly identify and address issues.
Then document all aspects of the project in relation to the research. As noted by Chatterji (2007, p. 244), “Without adequate and supporting data on context variables, implementation inputs, delivery processes, and other factors carrying the potential to affect outcomes simultaneously, outcome results are difficult to link with the new intervention, as well as hard to understand, explain, and defend.” This is even more important with large-scale and/or multisite projects as challenges can grow exponentially (Chatterji, 2007). As a project changes (or unravels), claims of validity require clear accounting of decisions made. Preregistering a study plan (e.g., Registry of Efficacy and Effectiveness Studies) allows an objective review of departures from what was intended and explanation for why changes were needed (e.g., O'Donohue et al., 2022; Toth et al., 2021). This then allows independent review of whether changes invalidate the results.
Also, anticipate structural and procedural roadblocks and understand that these may not be completely evident until you are well into the project. The difficulty is that this requires time to think through structures and processes, and in the rush to hit the ground running in a funded project, this step is easy to overlook. However, conversations about processes, procedures, and logistics prior to implementation beginning may shed light on potential issues. For our Project, we should have better understood the study participant placement process. One way we may have anticipated this roadblock would have been to comprehensively understand the placement process at the onset by working with the College and District placement offices and partner administrators. Understanding the District's placement flow was critical for the Project to determine how the College internship placement procedures fit within the partner's processes.
Additionally, develop strong partnerships with project partners as this can assist in circumventing challenges. But note that a cooperative working relationship takes time, so continue building on the momentum developed during the proposal stage, while you wait to hear about funding. The coordination of multiple partners in effectively implementing randomized experiments in practice has been noted by others (e.g., Demby et al., 2020; Moore & Moore Graham, 2011). Engaging partners and maintaining excitement about the project initially and over the course of the study is important (Demby et al., 2020). Consistent communication with study sites and District partners has allowed the Project to maintain and gain enthusiasm. Being proactive, engaging in conversations early so that relationships can be built with partners is critical (Demby et al., 2020), and we have been diligent in this aspect. All this said, recognize there will be limitations of the partner in structure and process for implementation of the randomized field trial.
Budget additional time for development of the intervention should your project be developing a complex intervention. In our Project, the intervention was not fully developed until year 3.
Remember, educational evaluations operate in complex systems and under changing conditions (Miller-Young & Poth, in press). Understand the realities of the context in which you are working and be ready to adapt to them, while at the same time understand the implications for study design. In our case, for example, this included working with a smaller pool of study participants from which to recruit as well as a smaller pool of intermediaries with which study participants could be placed. This may translate to insufficient power to detect results if the trend continues. We are investing resources into targeted recruitment of study participants as well as training future intermediaries. At the same time, larger samples don’t guarantee adequate power if implementation is poor and impact is compromised.
Understand that implementation variation across sites is natural (McCormick, 2019). Be flexible and think creatively in terms of measuring treatment fidelity. Particularly in complex multisite projects, there is not a one-size-fits-all approach to assessing fidelity of implementation. Take-up (see Figure 1), the link between the intervention offered and intervention received, is not dichotomous. Rather, it is on a continuum, varying in quantity and quality of services made available and received, and this needs to be measured (Weiss et al., 2013).
To improve accuracy of data collected, offer multiple modalities for collecting and submitting data and have continuous recalibration meetings to ensure understanding and to catch interventionist drift. As an example, the self-reported interventionist logs were a critical measure for allowing examination of both dosage and diffusion. However, it took regular recalibration meetings to ensure they were being used as intended.
Footnotes
Acknowledgments
This research is sponsored by the U.S. Department of Education (USDOE), Teacher Quality Partnership Grant Program, a Department of Education grant within the Effective Educator Development Programs in the U.S. Department of Education, #U336S180044. Additional support is provided by the University and School District partner. The contents of this paper reflect the views of the authors, and readers should not assume endorsement by the U.S. Department of Education or any other partnering entity.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by the U.S. Department of Education, (grant number U336S180044).
