Abstract
Evidenced-based mathematics interventions are critical for supporting students with mathematics difficulties. In research and practice, collecting implementation fidelity is important for ensuring that all the core components of the intervention are implemented as designed. Historically, implementation fidelity has been defined as multifaceted, including examinations of adherence, instructional quality, and student engagement, though mathematics intervention studies rarely report on fidelity components outside of adherence. The current study examined the relationships between these different components of fidelity and whether they are associated with student mathematics outcomes and intervention group size within the context of a first-grade mathematics intervention. Findings revealed relationships between components of fidelity with student’s initial mathematics skill; however, no relationship was observed between fidelity components and student mathematics growth. Findings for group size were mixed. Limitations, implications for research and practice, and future directions are discussed.
Keywords
Early Mathematics Intervention
Mathematics intervention is a core component within a multitiered system of support (MTSS) in which students receive additional mathematics instruction based on their need. Although there are recommendations for providing intervention within MTSS frameworks (Fuchs et al., 2021) and evidence supporting the use of specific strategies and instructional elements (Baker et al., 2002; Gersten et al., 2009), more work is needed to determine the effectiveness of specific intervention programs with a focus on the core components of the intervention that contribute to its effectiveness. Importantly, researchers and educators also need to consider the implementation of core intervention components.
Implementation Fidelity
Implementation fidelity (IF) has generally been defined as “the determination of how well an intervention is implemented in comparison with the original program design during an efficacy and/or effectiveness study” (O’Donnell, 2008, p. 33). Although this definition provides a summary of how IF is typically described, there have been a wide variety of definitions and conceptualizations across fields and authors, hence making it difficult to pinpoint a universally accepted definition (Sanetti & Kratochwill, 2009). Measures of IF are imperative within research studies to ensure that the intervention was implemented as designed, thus strengthening the integrity and validity of the study (Stains & Vickrey, 2017).
When it comes to measuring IF, many different frameworks and measures have been used across studies. For example, Carroll and colleagues identified a conceptual framework for IF that includes the assessment of program adherence and a focus on intervention complexity, facilitation strategies, quality of delivery, and participant responsiveness (Carroll et al., 2007). Similar elements have been identified across other proposed models of IF, including adherence, exposure (dosage), quality, participant responsiveness, and program differentiation (Sanetti & Kratochwill, 2009).
Adherence
The most common component of IF that researchers measure is adherence (Bos et al., 2022). Because IF measures the core components of an intervention, it has been noted that “fidelity assessments are inherently unique to each intervention and thus rely primarily on guidance from developers” (Abry et al., 2015). Typically, IF measures are created by researchers based on the core components of the intervention as Abry suggested. These measures most often assess adherence fidelity which refers to the extent to which an interventionist implements the intervention as intended. For example, in their examination of a reading comprehension intervention, Fogarty and colleagues (2014) discuss an IF assessment procedure in which independent observers recorded adherence data based on the presence of key components of the intervention. Using this method, researchers can determine if the program was implemented as written or designed.
Quality
Although adherence is the most commonly assessed, researchers have also examined different elements of IF beyond adherence. For example, the quality of intervention delivery is often included in intervention theories of change and has been examined in various ways. Measures of quality typically examine how well the steps of the intervention were delivered and in the context of academic interventions may also refer to the quality of instruction. In their examination of a first-grade mathematics intervention, Clarke and colleagues observed intervention sessions and rated the quality of instructional interactions occurring between interventionists and students (Clarke et al., 2014). Other researchers have reported assessing the presence of features of explicit and systematic instruction, and the interventionist’s ability to manage students’ behavior as the quality of implementation indicators (Bryant et al., 2021).
Student Engagement
Another component identified in various IF models is participant (student) responsiveness (Sanetti & Kratochwill, 2009). Student responsiveness has historically been defined as “a measure of participant response to program sessions, which may include indicators such as level of participation and enthusiasm” (Dane & Schneider, 1998, p. 45). Between measures of adherence, quality, and student responsiveness, previous reviews of the literature have found that student responsiveness is rarely reported in published articles (Bos et al., 2022; Dane & Schneider, 1998). A recent study from Doabler and colleagues operationalized student responsiveness within the context of a first-grade mathematics intervention via direct observation (Doabler et al., 2021b) as individualized practice and group practice opportunities. The study findings revealed that student mathematics gains from pretest to posttest were associated with the rate of group practice opportunities, suggesting that student’s responsiveness relates to the intervention outcomes. As an indicator of student participation and enthusiasm, student’s responsiveness may also be described as behavioral engagement. Behavioral engagement has been described as students’ attention and participation in instruction (Fredricks et al., 2004). This definition aligns most closely to the definitions of student responsiveness as a component of IF. Throughout the current study, student responsiveness is defined as student engagement, referring to behavioral engagement during instruction. Ratings of student engagement provide a measure of the level of student participation and enthusiasm in accordance with this definition.
Variables Affecting Fidelity
It is also important for researchers to measure IF because differences in intervention implementation between interventionists occur for a range of reasons. In the context of reading intervention, factors such as classroom management, pretest performance, English learner status, special education status, and gender have been examined as moderators of adherence fidelity and quality of instruction (Capin et al., 2022). In the context of mathematics intervention, intervention group size has been previously examined in relation to intervention effectiveness and IF. For example, in the context of a first-grade mathematics intervention with students grouped in large group (five students: one interventionist) and small groups (two students: one interventionist), statistically significant differences were detected in some measures of student mathematics gains, favoring students in the smaller groups (Clarke et al., 2022). Additionally, researchers examined differences in implementation quality between large and small intervention groups. Results revealed statistically significant differences in the number of individual practice opportunities, such that students in small groups received more individual practice opportunities than those in large groups (Clarke et al., 2022; Doabler et al., 2019). In the most recent examination, Clarke and colleagues also found that interventionists teaching the same content to smaller groups taught more activities, met more instructional objectives, followed teacher scripting more closely, used more prescribed mathematical models, and had overall higher total fidelity scores than those leading larger groups. These results suggest that intervention group size may affect an interventionist’s implementation of an intervention.
Implementation in Mathematics Intervention Research
A recent review of the mathematics intervention literature from 1990 to 2018 examined if mathematics intervention studies included measures of fidelity, and if so, what types of measures (Bos et al., 2022). Based on common IF frameworks, the authors coded studies for components of IF including adherence, quality, and student engagement. Among the 99 included studies from 1990 to 2018, Bos and colleagues found that a large portion (75%) reported collecting quantitative adherence fidelity data. Of the six total studies that reported quality of intervention data, only four reported quantitative quality data. Some studies reported a combined fidelity score composed of both adherence and quality but few reported collecting quality data alone. Additionally, only 36% of studies included a report of student engagement, though most were strictly narrative. Bos and colleagues noted that beyond adherence, both quality and student engagement data will be important for researchers to consider when creating the theories of change and evaluating intervention implementation (Bos et al., 2022). The findings from this review demonstrate that within the mathematics intervention literature, few researchers are collecting quantitative fidelity data across difference components of IF.
Recent work in mathematics intervention and implementation from Nelson and colleagues have further explored the relationships between adherence, quality, engagement, and student mathematics outcomes within the context of a mathematics intervention (Nelson et al., 2020). This study examined adherence, quality, and engagement within the context of a scripted whole and rational number intervention (Math Corps) for students in Grades 5–8. Using multilevel regression models, the authors found that in a model containing free and reduced-price lunch status, adherence fidelity, quality, and student engagement, only free and reduced-price lunch status and student engagement were found to have a statistically significant association with student outcomes. Results from this study add to the literature by identifying another component of IF (student engagement) that may be significantly related to student positive response to mathematics intervention.
Purpose of the Current Study
The purpose of the current study is to explore the relationships between IF measures of adherence, quality, student engagement, and student mathematics outcomes within the context of a first-grade mathematics intervention. This is a retrospective study that includes data from multiple studies of the Fusion intervention across four cohorts with the aim of taking a closer look at IF. Adherence, quality, and engagement were chosen as the focus of the current examination as they are often described as core components of implementation fidelity frameworks; however, they are all rarely examined within the context of mathematics interventions. Furthermore, this study will add to the current literature by building on the framework outlined in the work by Nelson and colleagues (2020) which used a multilevel modeling approach to determine the amount of variance in mathematics scores explained by the different components of IF. Although Nelson and colleagues explored the unique variance in mathematics outcomes explained by adherence, quality, and engagement, the current investigation aims to explore the total effects of each IF component on student outcomes. The current study will address one primary research question associated with student outcomes and one exploratory research question examining impacts on components of IF:
To what extent are components of implementation fidelity (adherence, quality, and engagement) associated with gains in student outcomes?
How does intervention group size relate to each component of implementation fidelity at the group level?
Results from the current investigation will add to the literature by providing further insight into the relationship between different IF components and mathematics intervention outcomes. Although measures of IF are often included as a piece of mathematics intervention studies, relationships between IF and student mathematics gains are less often examined as a part of those studies. Additionally, results from both research questions may lead to implications for practitioners and educators regarding providing quality mathematics intervention. When planning for mathematics intervention, educators will need to consider the importance of the different IF components for enhancing student outcomes. Findings from the exploratory research question regarding group size may also help educators make informed logistical decisions about the ideal intervention group size for supporting student mathematics skill growth.
Method
The current study analyzed the data collected from the Fusion Efficacy Project (Clarke et al., 2016–2020), a multi-year research project funded through the Institute of Education Sciences (IES). The Fusion Efficacy Project included four independent cohorts and used a partially nested randomized control trial design blocking on classrooms across cohorts. Three cohorts of participating students were located in Oregon during the 2016–2017, 2017–2018, and 2018–2019 school years and one cohort of participating students was located in Massachusetts. With this design, 970 students were randomly assigned within classrooms to either (a) receive the Fusion intervention in a small group (two students), (b) receive the Fusion intervention in a large group (five students), or (c) a business-as-usual control condition in which students did not receive the Fusion intervention. In total, 194 students making up 97 groups were assigned to the small group intervention condition, 485 students making up 97 groups were assigned to the large group intervention condition, and 291 students were assigned to the business-as-usual control condition. All students were identified as experiencing mathematics difficulty based on their scores on a screening measure. Students included in the treatment conditions received the Fusion intervention in addition to their business-as-usual mathematics instruction. For the purpose of the current study, only data from students in the treatment conditions were analyzed.
Participants
Data included in the current study were collected from 26 elementary schools representing six school districts in Oregon and Massachusetts. Of the six school districts, two were in large suburban areas in Massachusetts and four were in small- and medium-sized cities in Oregon. Student enrollment across the participating districts ranged from 5,492 to 40,495 students. Within the participating schools, between 12% and 19% of students had disabilities, 4% and 38% were English learners, and 19% and 65% were eligible for free or reduced-price lunch. Additionally, between 1% and less than 1% were American Indian or Alaskan Native, 1% and 16% Asian, 1% and 5% Black, 9% and 87% Hispanic, between less than 1% and 2% were Native Hawaiian or Pacific Islander, 7% and 73% were White, and 1% and 8% were more than one race.
Students
Parental consent was obtained for all participating students. All participating students (2,304 in total) were screened in the fall of their first-grade year. Four measures from the Assessing Student Proficiency in Early Number Sense battery (ASPENS; Clarke et al., 2012), including the Magnitude Comparison, Missing Number, Basic Arithmetic Facts, and Base-10 were administered during the screening process. Students were considered eligible for the Fusion intervention if they had an ASPENS composite score in the Strategic (raw score between 13 and 26) or Intensive (below 13) categories based on winter benchmarks. Students who score at or below the Strategic category have less than a 50% chance of meeting end-of-year grade-level expectations in mathematics (Clarke et al., 2012).
Students who were found eligible for Fusion were ranked in each participating classroom by an independent evaluator. The 10 students with the lowest ASPENS composite scores were then randomly assigned into one of the study conditions: (a) small-group Fusion intervention, (b) large-group Fusion intervention, or (c) a business-as-usual control condition. Of the 2,304 students screened for eligibility, 1,455 met eligibility criteria. Randomization blocks consisted of the 10 students in each participating classroom with the 10 lowest ASPENS scores. If a classroom had fewer than 10 students eligible for Fusion, classrooms were combined to form virtual randomization blocks. In total, 97 classrooms (including virtual classrooms) were formed containing 10 students each. Students were then randomly assigned within classrooms to the Fusion small group (n = 194), Fusion large group (n = 485), or the control condition (n = 291). Students assigned to the control condition were not further divided into groups. For the current study, only data from the two treatment conditions (large and small Fusion groups) were included in the analysis due to the focus on intervention IF.
Interventionists
Fusion intervention groups were taught by interventionists that were either district-employed instructional assistants or hired specifically for the study. A total of 87 interventionists participated across the cohorts. Among the interventionists, a majority identified as female (94.4%) and White (77.5%), with 3.4% identifying as Hispanic, 6.7% two or more races, 2.2% African American, 1.1% Asian American/Pacific Islander, and the remaining 8.9% identified as another race or ethnicity or declined to respond. Many interventionists had previous experience with teaching small groups (90.1%) and with mathematics instruction (62%). On average, interventionists had 7.3 years of teaching experience (SD = 9.4). Of the 89 interventionists, 15.7% had a current teaching license; and 76.3% had taken an advanced mathematics course, such as calculus, algebra, and statistics at the college level.
Procedures
Fusion
The Fusion intervention is a Tier-2 first-grade mathematics intervention focused on teaching whole number concepts and skills. The Fusion intervention is highly scripted and scaffolded for interventionists in that it includes built-in teacher scripting, models, opportunities for practice, and corrective feedback. Fusion comprised 60 lessons and can be conceptualized within a three-component framework: (a) understanding of whole number concepts and skills, (b) principles of instructional design and delivery, and (c) high-quality instructional interactions. Fusion content is aligned to the Common Core State Standards Initiative (2010) and includes Base 10, place value, number to 100, basic number combinations, operations with two-digit numbers, story problems, and number properties. The Fusion program also prescribes the use of systematic and explicit instruction, including teacher modeling, scaffolding of content, and opportunities for student feedback. Across all four cohorts included in the current study, the Fusion intervention was delivered to groups of students 5 days per week for approximately 12 weeks. Each intervention session was approximately 30-min and occurred outside of Tier-1 mathematics instruction. Intervention began for all students in the early winter and ended in the spring to allow students time to respond to the core instruction.
Professional Development and Coaching
Two 4-hr professional development training workshops were delivered to all interventionists by project staff. During both workshops, project staff explicitly modeled instructional practices including using group response signals, correcting student errors, and pacing of activities within lessons. Interventionists were provided with opportunities to practice implementing Fusion lessons with feedback from project staff in both workshops. Project staff also provided coaching support throughout Fusion implementation. Each interventionist received two coaching visits which each consisted of an observation of a lesson followed by feedback.
Measures
Test of Early Mathematics Ability–Third Edition
The Test of Early Mathematics Ability–Third Edition (TEMA-3) was the primary distal mathematics outcome measure for the study. All students were individually administered the TEMA-3 at pretest in the winter of first grade and at posttest in the spring of first grade. TEMA-3 (Ginsburg & Baroody, 2003) is a standardized, norm-referenced assessment that measures mathematics ability in children ages 3–8 years 11 months. Content on the TEMA-3 includes numbering skills, number comparison, numeral literacy, mastery of number facts, calculation skills, and understanding of concepts. Alternate form and test–retest reliabilities are reported at .97 and .82–.93, respectively. Concurrent validity with other early mathematics assessments ranged from .54 to .91.
Fidelity—Adherence
Implementation measures were collected via observation. Each Fusion intervention group was observed approximately three times in person with approximately 3 weeks separating each observation. Observers comprised former educators, doctoral students, faculty members, and other experienced data collectors. All observers received approximately 10 hr of training which included a focus on direct observation procedures and the use of The Quality of Explicit Mathematics Instruction (QEMI; Doabler & Clarke, 2012) observation instrument. Observers were required to complete two practice observations and meet an interobserver agreement of .85 or higher before beginning observations for the study. In total, 672 observations were completed and 35.1% were coded by an additional independent observer. The average observation lasted approximately 25 min. The adherence fidelity measure was researcher-developed and measured the interventionists’ implementation of the intervention as intended. The adherence fidelity measure is provided in Appendix A. Observers rated adherence fidelity on a four-point scale (4 = all, 3 = most, 2 = some, 1 = none). Using this scale, observers rated the extent to which the interventionist (a) met the lesson’s instructional objectives, (b) followed the teacher scripting, and (c) used the lesson’s prescribed mathematical models. A total adherence fidelity score was computed by calculating the mean score of the three items above. Cronbach’s alpha was calculated at .81 for this measure. Interobserver agreement was calculated via intraclass correlation coefficients (ICCs) at .95. A stability ICC was also calculated to describe the proportion of variance in adherence between-groups versus within-intervention groups. Stability was calculated at .29.
Fidelity—Quality
The quality dimension of fidelity was measured using six items from the QEMI. Observers were asked to rate the extent to which the interventionist applied the following six instructional strategies: (a) pacing, (b) interventionist modeling, (c) providing group practice opportunities, (d) providing individual practice opportunities, (e) providing academic feedback, and (f) instructional scaffolding. Post-observation, observers rated each item on a one to four scale with a score of 1 representing that the item was not present, a score of 2 representing that the item was somewhat present, a score of 3 representing that the item was present, and a score of 4 representing that the item was highly present. A total quality score was derived by calculating the average score across the six items. ICCs were calculated to estimate interobserver agreement. Interobserver agreement for the six-item version of the QEMI was calculated at .97. Additionally, a stability ICC was also calculated to describe the proportion of variance in instruction quality between-groups versus within-intervention groups. The stability ICC was calculated at .53. Internal consistency of the six-item measure used in this study was calculated at .95 (coefficient alpha).
Fidelity—Engagement
Student engagement was measured via one item from the QEMI. Unlike the other six items within the QEMI, engagement was measured based on student behavior rather than interventionist behavior and was therefore pulled out to be analyzed separately from the other interventionist-focused items. The single item was scored by independent observers via the procedures outlined above. Observers rated the level of student participation and engagement based on students’ active involvement in the intervention, their compliance with the interventionist, and their completion of work during the lesson. Engagement was rated on a four-point scale with a score of 1 representing that student engagement was not present, a score of 2 representing that student engagement was somewhat present, a score of 3 representing that student engagement was present, and a score of 4 representing that student engagement was highly present. Interobserver agreement and stability were calculated using ICCs. Interobserver agreement was calculated at .88, and the stability was calculated at .37.
Statistical Analysis
Multilevel modeling was used to examine whether each component of implementation fidelity (adherence, quality, and engagement) was associated with gains in students’ TEMA scores (Research Question 1). The models accounted for pretest and posttest TEMA scores nested with students and students nested within Fusion intervention groups, and are represented by the following set of equations:
Ytij represents a TEMA score for assessment occasion t on student i in the Fusion group j. The model included three predictors: time at Level 1, Timetij (coded 0 at pretest, 1 at posttest); an IF component score at Level 3, IFComponentj (grand-mean centered); and their cross-level interaction. The model produced estimates of (a) the pretest TEMA score for students in Fusion groups with the mean level of IF, γ000; (b) the association between IF and pretest TEMA scores, γ001; (c) the pretest to posttest gain in TEMA scores for students in Fusion groups with the mean level of IF, γ100; and (d) the association between IF and pretest to posttest gains in TEMA scores, γ101. Full information maximum likelihood estimation was used for each model and
Grouped t-tests were used to test for differences in IF component mean scores between small Fusion groups (2:1) and large Fusion groups (5:1) (Research Question 2). Assumptions of normality and homogeneity of variance were assessed prior to running the t-tests. All analyses were conducted using R Studio software.
Results
Prior to running analyses, descriptive statistics were computed for each IF component, and pretest and posttest TEMA scores across treatment conditions. Descriptive statistics for student- and group-level variables are displayed in Table 1. IF component descriptives are based on mean ratings aggregated across three observations per group. Generally, based on the mean and median values of each component which were all rated on a four-point scale, ratings were typically moderate–high across observations. This is especially true for adherence fidelity with a mean score of 3.4 and a median of 3.4. The mean score for quality of instruction by group was 3.1 and the median was 3.0. Similarly, the mean score for engagement was 3.1 and the median score was 3.0.
Individual and Group-Level Descriptive Statistics Across Treatment Conditions.
Note. TEMA = test of early mathematics achievement.
To assess the extent of correlation between the different components of IF, parametric Pearson’s r correlations were analyzed. Each IF component’s skewness and kurtosis values were within ± 2 suggesting adequate distributions for correlation analyses. Overall, the correlation coefficients ranged from .60 to .79, suggesting strong correlations between all IF components (Cohen, 1992). There was a strong significant correlation between group adherence fidelity and group instructional quality (r = .73, p < .001). The correlation between group adherence fidelity and group student engagement ratings was calculated at r = .60 (p < .001). Lastly, the strongest correlation was calculated at r = .79 (p < .001) between group instructional quality and group student engagement.
Each component of implementation fidelity was tested in a separate multilevel model to determine their total associations with student outcomes. Table 2 displays a summary of results for each model. To check for differences between study cohorts, a sensitivity analysis was conducted between the original models and models that included dummy-coded variables for the study cohort. Results from this analysis revealed similar fixed effects across the different models and no differences in statistical significance; therefore, the study cohort was not included as a fixed effect in the final models.
Multilevel Analysis Results.
Note. Standard errors in parentheses. TEMA = test of early mathematics achievement; df = degrees of freedom; ICC = intraclass correlation coefficient.
p < .001.
Model 1 included adherence fidelity as the predictor variable. Results demonstrate that time was a significant predictor of TEMA scores (p < .001). Students grew an average of 7.2 points on the TEMA from pretest to posttest in groups with mean adherence fidelity. The effect of adherence fidelity was also significant (p = .001), indicating a positive association between adherence fidelity and group-level pretest TEMA scores (
Model 2 included instructional quality as the predictor variable. Similar to Model 1, time was a significant predictor of TEMA scores (p < .001) such that students grew an average of 7.2 points on the TEMA from pretest to posttest. The effect of instructional quality was also significant (p < .001), indicating that groups with higher ratings of instructional quality included students with higher pretest TEMA scores (
Model 3 included student engagement as the predictor variable. Like in previous models, time was a significant predictor of student TEMA scores (p < .001). The effect of student engagement was also significant (p = .001), indicating that groups with higher ratings of student engagement included students with higher pretest TEMA scores (
Group-level means, standard deviations, and t-tests for each IF component by Fusion group size are presented in Table 3. Results indicated that large and small Fusion groups did not statistically significantly differ on the ratings of adherence fidelity, t(186) = −1.82, p = .07, or instructional quality, t(186) = −0.24, p = .81. The grouped t-test comparing student engagement ratings between large and small Fusion groups was statistically significant, t(186) = −1.99, p = .05, with small groups receiving higher student engagement ratings than large groups (Hedges’ g = 0.31).
t-Test Results Across IF Components Between Large and Small Groups.
Discussion
The current study examined the relationships between different components of IF and student outcomes within the context of a highly scaffolded and supported first-grade mathematics intervention. Additionally, the current study explored the relationship between group size and intervention implementation across the different components of IF. Importantly, this study included measures of IF beyond adherence, including quality of instruction and student engagement. These components have been historically identified in conceptualizations of IF, but recent reviews have shown that they are not often evaluated within mathematics intervention research (Bos et al., 2022; Dane & Schneider, 1998). Results indicated that IF components were strongly correlated with each other. IF components were not significantly related to student mathematics growth. However, the findings revealed that fidelity scores across IF components were significantly higher for groups with students who had higher initial mathematics skills based on pretest TEMA score. Results for the relationship between group size and IF components were mixed with small groups receiving higher ratings of student engagement compared with large groups but little difference in ratings of adherence fidelity and instructional quality.
Implementation Fidelity Components
Descriptive statistics for each IF component revealed that ratings of fidelity were moderately strong. Each IF component was rated on a scale from one to four, and mean scores were all above 3.0 illustrating moderate–high fidelity across components. Across IF scales, a score of 3 represented that a component was “present.” A mean score above 3.0 for all components suggests that on average, all core components of the intervention were present across all observations, although there is still room for improvement as items were rated on a one to four scale. It is important to consider the context in which these components were rated when interpreting their values. The Fusion intervention program has built-in supports including organized lesson structures and scripting. Teacher scripting that is built into the program was designed to ensure a minimum level of fidelity and quality as these components are a large part of Fusion’s theory of change and conceptual framework. Alongside the supports built into Fusion, the current study provided interventionists with a total of 8 hr of professional development which outlined the core components of the intervention program and provided interventionists with the opportunity to receive feedback on their delivery of the program. Interventionists that provided feedback on these trainings rated their ability to implement the Fusion intervention with fidelity strongly. Additionally, interventionists were also provided with multiple coaching sessions during implementation to further support their delivery of the program. It is likely that these built-in and additional supports contributed to the fidelity scores observed across adherence, quality, and student engagement. Although not touched on in the current study, future research may also consider how the amount and quality of professional development and coaching supports may have on different components of IF.
Results from the current study also illustrated strong correlations between adherence fidelity, quality of instruction, and student engagement. These results were consistent with previous findings showing significant correlations between measures of fidelity (Abry et al., 2015). The strong correlations between IF components provide additional information regarding the conceptual relationships between different components. In part, these results are expected based on the design of the Fusion intervention program. For example, if an interventionist has high adherence fidelity that indicates that they are following the lesson scripting and design as written and because lessons were written with instructional quality and student engagement in mind, it follows that high adherence to the program would also result in high quality and student engagement.
Fidelity Components and Mathematics Outcomes
Results from the multilevel models examining the relationship between IF components and student mathematics outcomes revealed that the ratings of IF components did not predict student growth in mathematics skills from pretest to posttest. The r-squared equivalents calculated for each time × predictor intervention term revealed little effect of each IF component on student mathematics growth. Findings from these analyses align well with results from the work by Nelson and colleagues (2020) which also showed non-significant relationships between student outcomes, and adherence fidelity and quality of instruction. Notably, while Nelson and colleagues found that student engagement predicted posttest mathematics scores, the current investigation did not find that student engagement was a significant predictor of mathematics gains. These results suggest that students can make gains in mathematics skills and knowledge when provided with an evidence-based mathematics intervention and that these gains may not be affected by variations in implementation. As detailed previously, there was little variation across implementation components and scores were generally high which may have limited the current study’s capacity to investigate this research question. Studies of intervention programs with a greater range of IF scores would potentially enable a more thorough investigation of the relationship between IF and student mathematics outcomes.
Results from the multilevel models also revealed a significant relationship between IF components and pretest mathematics scores. This finding demonstrates that intervention groups with students who had higher initial mathematics scores (as demonstrated by higher pretest scores) were in intervention groups with higher adherence, quality, and engagement ratings. This finding aligns well with the previous examinations which found that intervention groups with higher initial skill received more practice opportunities (Doabler et al., 2021b). A reasonable hypothesis is that students with higher initial mathematics scores were more likely to be successfully acquiring the skills taught in the intervention and thus more highly engaged during instruction. This level of engagement could have made it easier for interventionists to implement the intervention with adherence and quality. Additional inquiry into the impact of student initial skill on IF is needed to further quantify and understand this relationship.
Group Size and Implementation
Alongside the examinations of student outcomes, the current study also included a focus on factors that may relate to intervention implementation including intervention group size. It was initially hypothesized that the small intervention groups would have higher ratings of IF across components; however, results were mixed. Specifically, small groups with two students had significantly higher student engagement ratings than large groups with five students. Additionally, small groups had higher ratings of adherence fidelity compared with large groups; however, this difference was not statistically significant. The Hedge’s g effect size of .27 for this comparison suggests potential for clinical or practical significance and is similar to previous findings (Clarke et al., 2022). Differences in instructional quality between small and large groups were minimal and non-significant, suggesting that group size was not related to the quality of instruction. These results were consistent with previous findings that the quality of instruction did not differ between small and large intervention groups (Clarke et al., 2022).
The mixed results described above may be in part due to moderate–high fidelity scores across large and small groups. Even so, these results suggest some variation in IF based on group size. It may be that only having two students in a group allows the interventionist more time to complete more intervention activities thus contributing to slightly higher adherence scores in the small groups compared with the large groups. It may also be easier for interventionists to engage smaller groups of students with more individual opportunities to respond and practice. For example, within their examination of differences in Fusion intervention outcomes by group size, Clarke et al. (2022) found that students in small groups had greater gains as measured by the TEMA than those in large groups. They also found a statistically significant difference in the number of independent practice rates favoring the small intervention groups over the large groups, but no differences between the overall quality of instruction. These findings along with those illustrated by the current study support further inquiry into the role that group size plays in intervention delivery and student outcomes.
Limitations
Several limitations should be considered when interpreting findings from the current study. Firstly, ratings of IF were moderate–high across components. This limitation is due in part to the high level of interventionist supports that were built into the Fusion intervention, provided through professional development and provided via ongoing coaching that then aided interventionists in delivering the intervention with fidelity. Although these high ratings are promising when considering the feasibility of the Fusion intervention, they make it difficult to draw conclusions regarding associations between IF components and student outcomes. It was hypothesized that students in intervention groups with high IF ratings would experience greater gains from the Fusion intervention. However, with little variation in IF scores, the current study was limited in investigating the relationship between IF and student outcomes. This limitation highlights the need for more sensitive measures of IF to detect any additional variation in scores.
Secondly, stability ICCs were low across IF component ratings suggesting little stability in IF ratings across observation sessions. This lack of stability may be in part due to variation in implementation across intervention sessions. The low-stability ICCs make it difficult to conclude that each IF component rating is a strong representation of implementation within a group across the course of the intervention. Additional observations of intervention groups could result in higher stability across IF component ratings; however, additional observations would also require additional resources from the research team.
Lastly, measures were limited to those collected in the original Fusion efficacy trials. This limitation did not allow for more in-depth measures of IF across components. For example, the engagement measure consisted of a single item rated on a four-point scale. Additional items assessing student engagement could result in more accurate ratings and could allow for assessment across forms of student engagement (including emotional and cognitive engagement). A measure of academic engagement that includes ratings of behavioral, emotional, and cognitive engagement has been called for in previous literature (Fredricks et al., 2004). Within the context of early mathematics intervention research, measuring across forms of engagement would require ratings of students’ feelings, interests, and attitudes toward mathematics (emotional), their investment in learning and self-efficacy in mathematics (cognitive), and their attention and participation in mathematics instruction (behaviors; Fredricks et al., 2004; Kwan Lo & Foon Hew, 2021).
Future Directions
Several future directions are recommended based on the findings and limitations of the current study. To start, additional examinations of IF components within different settings are needed to further examine the relationships between IF and student outcomes. Specifically, recording IF components within more naturalistic environments that do not provide the same level of interventionist supports found within an efficacy trial where the primary goal is to investigate impact may result in more variation in IF component scores. This variation would allow for a more nuanced analysis of the relationship between IF components and student mathematics gains that more closely aligns with the real-world contexts.
Future studies may also include different types of mathematics programs as findings from the current study are also limited to the Fusion intervention specifically. Examining these relationships with interventions across grade levels and complexity of mathematics content is important for identifying intervention characteristics related to IF. Because the Fusion intervention covers early mathematics content that many interventionists are likely more comfortable and familiar with than more advanced mathematics content, lesson delivery may be easier for interventionists. Further examination of IF components within more complex mathematics interventions is needed to better understand if complexity level has any additional association with implementation. Similarly, examining these relationships within the context of non-scripted programs would also be important. As described earlier, the Fusion intervention is a heavily scripted program, and this level of scripting may assist interventionists in implementation across components. It could be hypothesized that when examined within the context of an early mathematics program that has less built-in support, implementation ratings would be more variable and IF components may have larger associations with student outcomes. Future inquiries into IF components should consider examining thresholds for which IF predicts student outcomes. For example, it is possible that once an interventionist has achieved a certain level of IF, there is no additional association with student outcomes.
Results from the current study also suggest that more work is needed to develop measures of IF that are more sensitive, include multiple IF components, and are feasible. Ratings of adherence fidelity within the current study were collected via a measure that was separate from the measure used for quality and engagement. Results from previous reviews have shown that few mathematics intervention studies report quality and engagement data while relatively more report adherence fidelity data (Bos et al., 2022). Development of new tools that incorporate multiple components of IF would provide researchers with opportunities for collecting, analyzing, and reporting IF data across components. Within their review, Bos and colleagues (2022) suggest that researchers reflect on their theories of change to identify which key fidelity components need to be measured. Bos and colleagues also call for future mathematics intervention research to include multiple measures of IF, including quality and engagement. The current examination included a measure of student engagement that aligns with the definition of behavioral engagement including involvement in learning tasks, effort, persistence, concentration, and attention (Fredricks et al., 2004). As noted above, there have been previous calls for measures of academic engagement to include multiple forms, including behavioral, emotional, and cognitive (Fredricks et al., 2004). Including items that address each of these forms within a measure of IF would provide a more diverse and comprehensive rating of student engagement. Based on the low-stability ICCs reported within the current study, future research on the development of IF measures should also consider the feasibility of these measures. Although Bos et al.’s recent review found that a majority of studies reporting adherence and quality fidelity data used live observation as their method for data collection, live observation often can be resource-intensive. Live observation often requires a large team of personnel to be trained and requires an acceptable level of inter-rater reliability to be met. When developing new measures of IF, researchers need to consider balancing the inclusion of various components with the feasibility of data collection. Alternative methods of data collection, such as video recording, audio recording, self-report, and interventionist report, should also be considered to maximize utility and feasibility.
Lastly, future research should continue to explore what factors are associated with IF including teacher mathematics knowledge (Sutherland et al., 2022), interventionist behavior management skill (Capin et al., 2022; Lekwa et al., 2019; Van Dijk et al., 2019), student initial skill (Capin et al., 2022), and intervention group composition (Doabler et al., 2021a). Identifying factors that influence interventionist’s implementation will provide further assistance to educators implementing mathematics interventions. Future IF research should be conducted in different contexts and with different curricula to further investigate the relationship between IF and factors that may relate to IF outside of the current context.
Conclusion
Within both research and practice, it is imperative that IF is measured and used in the interpretation of mathematics intervention effects. Although a majority of mathematics intervention studies only report adherence fidelity data (Bos et al., 2022), conceptual models of IF have historically included components, such as instructional quality and student responsiveness (Dane & Schneider, 1998; Nelson et al., 2020). The current study examined the relationships between IF across multiple components (adherence, quality, and student engagement) and student mathematics gains within the context of a highly supported and scripted evidence-based mathematics intervention (Fusion; Clarke et al., 2014). The current study also examined the relationships between intervention group size and IF. Findings revealed that ratings of adherence, quality, and student engagement are highly correlated with each other. Results also illustrated that while adherence, quality, and engagement scores were all positively related to student mathematics skills at pretest, they were not related to student gains in mathematics skills and knowledge from pretest to posttest. Additionally, there were differences in IF component ratings favoring small intervention groups over large groups. Based on previous reviews (Bos et al., 2022), inquiries (Nelson et al., 2020), and the results from the current study, there is a need for the development and use of IF measures that include multiple components within mathematics intervention research. The field would also benefit from additional explorations of IF components and student outcomes across various contexts, including more naturalistic contexts where implementation support is more limited or variable. Further exploration of relationships between IF components, student outcomes, and other environmental characteristics will aid educators in continuing to support students struggling with mathematics.
Supplemental Material
sj-docx-1-ldx-10.1177_00222194251315191 – Supplemental material for An Examination of Implementation Fidelity Within the Context of a Tier 2 Mathematics Intervention
Supplemental material, sj-docx-1-ldx-10.1177_00222194251315191 for An Examination of Implementation Fidelity Within the Context of a Tier 2 Mathematics Intervention by Cayla Lussier, Ben Clarke, Derek Kosty, Geovanna Rodriguez, Kathleen Scalise, Christian Doabler and Jessica Turtura in Journal of Learning Disabilities
Footnotes
Declaration of Conflicting Interest
The author(s) declared the following potential conflicts of interest with respect to the research, authorship, and/or publication of this article: B.C. and C.D. are eligible to receive a portion of royalties from the University of Oregon’s distribution and licensing of certain FUSION-based works. Potential conflicts of interest are managed through the University of Oregon’s Research Compliance Services.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This investigation was supported by the Institute of Education Sciences, U.S. Department of Education, through grants R324A090341 and R324A160046 to the Center on Teaching and Learning at the University of Oregon.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
