Abstract
In this systematic literature review, we examine the corpus of empirical studies in education that use administrative data (i.e., population-level data) to describe and estimate the impacts of service delivery models for specially designed instruction on outcomes for students identified with special education needs. We focus on studies that use quantitative data analysis—either descriptive or causal—to answer questions about the relationship between special education service delivery models and student outcomes. We analyze seven studies, each of which finds a positive relationship between more time spent in general education classrooms and outcomes for students with disabilities (SWDs). In our analysis, we discuss the affordances and limitations of this type of analysis and opportunities for the field to expand data collection and analysis of population-level data in a way that better illuminates the state of special education services, both in the present and longitudinally, for SWDs.
Where and how should students with disabilities (SWDs) be educated? This is one of our field’s oldest questions (Dunn, 1968) and one that has produced substantial commentary and scholarship (e.g., Suter & Giangreco, 2009; Zigmond, 2003). It is also a perennially pressing question for school personnel, who must make annual decisions about the placements that will best meet their students’ individual needs. Yet, nearly 50 years following the first federal legislation guaranteeing a free and appropriate education for SWDs (Public Law 94–142), our field still does not share consensus on what kinds of service delivery models—how and where students are educated—are most likely to lead to improved student outcomes. Central to the challenge of building consensus on this topic is that a defining feature of special education is its individualization; decisions about services and the settings in which students should receive them are made on a case-by-case basis. In addition, schools expect special education to address many different academic, behavioral, and adaptive goals, often requiring a range of personnel to meet these goals. Furthermore, even though federal legislation demands that students receive instruction in the least restrictive environment (LRE)—meaning that students be taught alongside their peers without disabilities to the maximum extent that is appropriate—schools use a range of models to meet these requirements.
As we argue in this article, an additional reason driving a lack of consensus around service delivery models in special education is that the field lacks sufficient research exploring the relative effectiveness of any one of these different approaches at scale. As we describe below, research on service delivery models has, with few exceptions, fallen into one of a couple of categories: qualitative research or experimental evidence from one or a small number of schools. There is a need for additional research drawing on large-scale, longitudinal data sets that would facilitate analyses that are both descriptive (i.e., no cause-and-effect claims) and causal (i.e., seeking to credibly estimate a cause and effect) and are focused on the implementation and effectiveness of different service delivery models. These types of data, which we define in more depth in the “Method” section, are often termed “administrative data” and are particularly useful in providing data on the “natural” conditions of how service systems, such as schools, function. The goal of this literature review is to better understand the extent to which administrative data have been used to explore the effectiveness of service delivery models in special education and to identify ways in which the use of administrative data might provide a fuller picture that informs practical decision-making about where and how students should receive services in schools.
Service Delivery Models
The setting in which students receive services is central to service delivery models; this setting is denoted in a student’s Individual Education Program (IEP) and determined based on the recommendation of a student’s educational decision-making team. Some of the most common service delivery models include support through “push-in” or co-teaching methods in general education classrooms; supports supplemental to, but outside of general education classrooms, such as instruction received in a resource room (i.e., “pull-out” services); or supports for the majority of the school day provided in a substantially separate, “self-contained” classroom (Epler & Ross, 2015; Florida Inclusion Network, 2012). Given their interrelated nature, it is challenging to parse out service delivery models from their classroom settings. For example, while studies of co-teaching models frequently refer to co-teaching in general education classrooms, the model could also be implemented in a self-contained classroom setting. In this review, we recognize that time spent included in general education classrooms—often termed “inclusion”—is often used as proxy for access to different service delivery models, both in research and in federal reporting. Although using classroom settings as a proxy for service delivery models lacks some desired level of precision, we posit that this proxy nonetheless is additive in that it furthers our ability to narrow in on how and where SWDs are being educated at scale.
Educational decisions about the ways in which students access their instructional services requires consideration of a variety of aspects including experience and credentials of the instructors, sociohistorical contexts including segregation and inclusion, and the alignment of services with the general education curricula (Artiles et al., 2006; Rea et al., 2002). Furthermore, the rights of SWDs to be included with their general education peers must be considered not only as a legal right, but also as an issue of civil rights and equity (Voulgarides & Tefera, 2017). Given the significance of these various factors, it is essential to increase empirical evidence regarding professional recommendations of service delivery. Below, we focus on three high-frequency classroom settings and some service delivery models within each.
General Education Classrooms: Inclusion, Co-Teaching, and Push-In Services
General education classrooms can be defined as the classroom setting in which the student would be instructed without the provision of an IEP and special education services. One way that SWDs may receive instruction in inclusive general education classrooms is through a single-teacher instructional model. In this model, a general education teacher collaborates and consults with the student’s case manager to provide additional instructional supports either within or outside of the classroom. Another common model for serving SWDs in general education is co-teaching, which is broadly defined in the literature as two staff members sharing responsibility for instructing a group of students (Friend, 2008). Co-teaching can look like small group instruction within larger whole-group instruction, or back-and-forth “team” teaching. With this model, the service provider serving as a “co-teacher” brings services into the general education space, which is considered more inclusive than “pull-out” or segregated classroom settings. While co-teaching is an increasingly prevalent form of service delivery, there is little large-scale empirical evidence for co-teaching, especially related to academic impacts, to support its use broadly (Jones & Winters, 2022; King-Sears et al., 2021; Murawski & Lee Swanson, 2001; Scruggs et al., 2007; Solis et al., 2012).
Resource Room Classrooms: Pull-Out Services
Resource room classrooms are exclusively designed to support SWDs receiving special education, in which they receive additional instruction and intervention, which is often termed “pull-out” services. Students may receive resource room instruction in conjunction with inclusive general education instruction, but their interventions are delivered in a separate physical space. Approximately, 16% of students identified with disabilities were educated with this model in 2022 (National Center for Education Statistics, 2022). In a “resource room” model, a special educator provides specially designed instruction necessary for a student to meet their educational goals outlined in their IEP. Research investigating the pull-out model has frequently been qualitative, and often investigates the social impacts of pull-out models on older elementary students (Hannes et al., 2012). Other studies have focused on small sample sizes and have not used methods that investigate causality (Hannes et al., 2012; Rea et al., 2002).
Self-Contained Classrooms: Substantially Separate Services
Finally, receiving specially designed instruction in a “substantially-separate” or “self-contained” classroom is particularly common for students with significant emotional, behavioral, or intellectual disabilities (Epler & Ross, 2015). In such a model, students are educated in classrooms serving only SWDs. IEP teams typically select this service delivery model when students are seen as needing separate placement due to intellectual or behavioral challenges that greatly outweigh their ability to engage in classroom instruction with their peers in the general education setting. Often, a special educator provides primary services in these classrooms and relies on paraeducators to support various parts of their students’ school day.
Service Delivery Models and Time in General Education
Each student’s IEP specifies a proportion of time per day spent in general educational settings, and service delivery models may vary within each of these settings. In practice, students may engage with any combination of service delivery models, including but not limited to those described above. In federal reports and in many administrative data sets, the time that a student is assigned to be included in general education is segmented in the following ways: over 80%, 40% to 79%, or less than 40% of a students’ time in school is spent in included in general education classrooms. Thus, in research and practice, these proportions of time are often used as a proxy for delineations of different service delivery models: inclusive general education, resource room, and self-contained settings and instruction, respectively. This proxy undoubtedly lacks detail enough to answer specific, necessary questions about service delivery models, such as: in which classroom settings does co-teaching work best as a service delivery model, and for whom? We argue, however, that empirical investigations of even these broadly defined categories of time spent in general education do get us closer to establishing a baseline understanding of the effectiveness of different strategies and settings for educating SWDs at scale. This work can serve as a necessary complement to other types of analyses that provide finer-grained levels of nuance, such as qualitative studies. Therefore, it is important to understand and expand the landscape of research on service delivery models and the ways in which they shape outcomes for SWDs, so that, practitioners can make evidence-based decisions about where and how to educate SWDs, and how their classroom placements may be adjusted to better serve students’ needs.
Current Research
Although significant qualitative, quantitative, and mixed-methods research explores the conceptualization and use of a variety of service delivery models, research that assesses their effectiveness at scale (i.e., using state- or district-level population data, which include not just a selected sample of students in a particular group but the entire group itself) remains limited. It is challenging to design generalizable studies of highly individualized services; however, there is some research that examines the relationship between instructional setting and student outcomes, but it is primarily done within the context of studying effectiveness of highly specific interventions (Bottge et al., 2018; Fuchs et al., 2015; King-Sears et al., 2020). For one example, Fuchs et al. (2015) synthesize the effects of math fraction interventions tested in larger randomized control trials when delivered in more inclusive settings. Although it is important to best understand the contexts in which interventions are most effective, these studies do not necessarily replicate a “natural” condition or allow us to disentangle effectiveness of the intervention itself from effectiveness of a service delivery model or placement.
One way to address this gap is to conduct studies that include population-level data and look at service delivery models that vary by context and population, not just by time spent in a particular setting. By population-level data, we mean that which includes an entire group being studied rather than a selected sample. This can be done at multiple levels, such as a state, district, or school. For example, a state-wide longitudinal data set on personnel may include data on all teachers employed within the state, while national surveys like the National Teacher and Principal Survey (NTPS) provides a snapshot of a selected sample of teachers. While both have distinct benefits, population-level data allow one to study all cases of a certain group, including those which may not fall into a survey group (e.g., individuals at the margins; Figlio et al., 2016). This is particularly useful in special education, given that some groups of students—such as those with lower-incidence disabilities—represent a smaller proportion of students overall and thus it is more challenging to conduct inferential analyses about these groups, given considerations of statistical power. Using population-level data provides a key opportunity for understanding the full scope of a group. Although sample sizes are often, but not always, large in administrative data sets, they provide an opportunity to observe all members of a group.
Why Focus on Administrative Data?
Numerous studies focus on inclusion and the features of specific service delivery models, such as co-teaching (see King-Sears et al., 2021 for a review). Broadening research methods and data sources may provide a richer starting point to inform decision-making around a student’s LRE, and administrative data provide some key advantages for studying the trade-offs of different service delivery models. Figlio et al. (2016) argue that the use of administrative data in educational research has several benefits, including: (a) increasing the ability and precision in identifying meaningful relationships, (b) increasing opportunities to identify rare events (e.g., in the case of service delivery models, perhaps it is the case that some settings and models work well for students with a particular set of circumstances or diagnoses, including those with smaller populations), and (c) gathering generational, longitudinal, and population-specific information and patterns that are necessary to inform policy, including district and state-level policies. Note that we distinguish administrative data from data sourced from nationally representative surveys because surveys do not include the full universe of the student and school personnel population in sampling. Although the claim could not be made that the analysis of administrative data paints a full picture of service delivery effectiveness alone, effective analysis of administrative data can certainly inform some existing research gaps. We believe that these data provide a complementary way to explore service delivery models systematically and broadly, motivated by the fact that studies investigating the effectiveness of such models is a federal law and state policy concern. In fact, Part B of Individuals with Disabilities Education Act (IDEA) mandates that states collect certain administrative data, such as child count and discipline data, as part of regular reporting requirements.
Yet, as with any methodology or data type, administrative data are not without their drawbacks. One of the major hindrances to using administrative data is that it can be challenging to acquire population-level data that address the level of granularity a researcher may want to conduct certain analyses. Some population-level data are publicly available, ranging in accessibility from freely downloadable to accessible through applications for state data systems (e.g., North Carolina). Often, however, data are accessed by developing trusting research practice partnerships that provide researchers with the data access necessary to conduct analyses. Thus, a push for additional analyses of special education services using administrative data must be considered carefully in the context of feasibility. While, again, there are some administrative data sets that are publicly available, the level of granularity of publicly available data often does not lend itself to studies that could examine the impact of a specific different service delivery model on individual students. For example, publicly available Integrated Postsecondary Education Data System (IPEDS) data are aggregated at the school-wide level, which are useful for some analyses, but do limit some others. More still is available through applications, but there are nonetheless considerations of feasibility in that regard as well. Thus, to gain access to more fine-grain—and more protected—data, researchers must likely establish or leverage research–practice partnerships, which are challenging and timely to develop (Henrick et al., 2017). These kinds of partnerships also often come with strict data-sharing agreements, which can limit the extent to which researchers can share their data beyond the research team (Conaway et al., 2015). If we are to encourage states to more frequently partner with researchers, we need to acknowledge the challenges around data sharing and identify strategies that may either protect participant privacy or describe the benefits that come with making data accessible to a broader range of researchers.
Researchers emphasizing the strengths and possibilities of Open Science point to its opportunities to limit bias, encourage replication and reproducibility, and add valuable null results to the literature (e.g., Lombardi et al., 2023). With research practice partnerships, these benefits may be especially appealing in building trust with partners and making them a part of the research process. Thus, administrative data analysis both offers an opportunity to enrich our understanding of topics, such as the effect of different service delivery models on student outcomes in necessary ways that complement current methodologies, and also provides an opportunity to expand the use of Open Science practices in special education research for the benefit of increasing transparency and expanding possibilities for replication.
Research Questions
With consideration for the need for additional evidence for decision-making surrounding service delivery models and the possibilities presented by administrative data, the present study seeks to explore the following research questions:
Method
Key Definitions
Prior to engaging in our search and review process, we established working definitions of several key terms. These definitions were iteratively refined throughout the review process.
Service delivery models
Considering that this review is focused on service delivery models for SWDs, we wanted to establish a clear definition of how we conceptualize service delivery models. Primarily, we defined service delivery models as modes of instruction that would be specified in a student’s IEP. Such models include, but are not limited to, general education inclusion, resource rooms, and self-contained classrooms. In addition, we classified classroom staffing models (e.g., co-teaching) as service delivery models given that specialized instructional support may be delivered by a special education co-teacher or a paraeducator assigned to an individual student (i.e., a dedicated aide) or a whole classroom. In such a case, those services would most likely be specified in a student’s IEP. Importantly, we only included such staffing models if the study explicitly focused on the assignment of SWDs to such a model (i.e., studies focused generally on teacher or paraeducator staffing would be excluded).
Positive Behavior Intervention Systems (PBIS) and Response-to-Intervention (RTI) were not classified as service delivery models because they are designed as whole-school systems that include all students. Our review is focused on service delivery models at the individual classroom level that are intended to support students already identified with disabilities. Thus, even though there have been some high-quality studies on PBIS and RTI using statewide administrative data, these are outside of the scope of this review. In addition, we did not classify specialized programs (e.g., Career and Technical Education [CTE], gifted education, charter schools) as service delivery models because our focus is on specialized instructional services.
Administrative data
We define administrative data in accordance with extant literature on administrative data use in education and public policy analysis. Administrative data can be defined as “data collected by government entities for program administration, regulatory, or law enforcement purposes . . . usually collected for the full universe of individuals, businesses, or communities affected by a particular program or regulation” (Office of Management and Budget [OMB], 2016). Importantly, we define administrative data as population-level data (i.e., data “collected for the full universe of individuals”). This necessarily excludes other important, but different data, such as data collected via a nationally representative survey (e.g., data from the National Longitudinal Transition Study-2 [NLTS2] or the Youth Risk Behavior Surveillance System [YRBSS]) or for a specific study purpose (e.g., data collected from a purposive, but small, sample of schools or classrooms).
Student outcomes
Given that student outcomes can be defined broadly, we were inclusive of numerous domains in our review, including academic, postsecondary, and social–emotional (e.g., sense of belonging, growth mind-set) measures. In our definition, student outcomes did not, however, include the classification, reclassification, or declassification of a student for special education services. While these outcomes are at the core of service provision for SWDs, our review was focused on service delivery models that are used after determination that a student is eligible for services. Thus, studies on identification and eligibility for services or those focused on placement as an outcome in and of itself were outside the scope of this review.
Search and Screening Procedures
The author team engaged in a multistep process of article identification and analysis. We first identified a set of potential articles via electronic database searches and additional search efforts. Next, we screened the titles and abstracts of those articles for inclusion in an ancestry and progeny search. Finally, we conducted a full-text review of articles for inclusion and analysis (see Figure 1).

Inclusion and Exclusion Criteria (PRISMA Diagram).
Identification
Electronic database search
The search process for this review began with an electronic database search of Web of Science, EBSCO, ERIC, and Education Database. In each of these databases, the same search terms were used (see Table 1). To identify studies with a clear focus on using administrative data to examine service delivery models for SWDs, we specified that the search terms needed to be found in either the title or the abstract of a study.
Database Search Terms and Yield.
Additional search efforts
In addition to the database reviews, the authors conducted a search of three working paper repositories: NBER, CALDER, and EdWorkingPapers. These repositories were included given their particular focus on educational research that involves quantitative analysis using administrative data. NBER and EdWorkingPapers both had search capabilities, so we searched broad key terms (i.e., “students with disabilities,” “service delivery”). CALDER’s repository did not have a search function, so each CALDER working paper was individually reviewed by title. A total of 37 working papers were reviewed by at least one author, and 10 were reviewed by the full team. While we structured our search of electronic databases and working paper repositories to be as rigorous and replicable as possible, the author team was made aware of some articles that met our criteria but were not identified in our searches (i.e., Barrett et al., 2020; Jones & Winters, 2022; Schifter, 2016). Those articles were included in the ancestry and progeny search, and the journals in which they were published were hand searched.
Screening
After a preliminary search of articles with each method, one author read through each title and abstract (n = 263) and identified articles that met the following three criteria: quantitatively analyzed a large data set (n > 50), focused on classrooms serving SWDs (including general education inclusion), and included student outcomes as a variable for analysis. This generated a total pool of 51 records, which were reviewed with a second author and consolidated to a pool of 37 records.
Inclusion criteria
After identifying 37 studies, the full author team reviewed each title and abstract and coded each article using the seven inclusion criteria. Included articles must: (a) use administrative data, (b) use data from the United States, (c) apply quantitative methodology, (d) focus on how special education services are organized and delivered, (e) focus on specific classroom types (e.g., inclusion, resource room, self-contained classroom) and/or classroom staffing organizations (e.g., co-teaching), (f) focus on students already receiving special education services, and (g) include individual student-level outcomes in analysis. See Table 2 for specifications of each inclusion criteria, with examples.
Inclusion Criteria Definitions With Examples.
Ancestry and progeny search
After coding for inclusion and coming to consensus as a team, one author conducted an ancestry and progeny search of three records (i.e., Clotfelter et al., 2016; Jones & Winters, 2022; Theobald et al., 2019). The ancestry search was completed by searching the references page of each included article. Only the titles were screened in the initial round of this process, and abstracts were screened if needed. To complete the progeny search, the author conducting the search used Google Scholar to examine all articles that cited each included article since its publication. In the progeny search, both titles and abstracts were screened. Two additional articles were found in this search (i.e., Cole et al., 2021, 2023), along with the peer-reviewed version of one working paper (Theobald et al., 2019). The ancestry and progeny searches were repeated for each of those articles and no additional articles were identified as eligible for inclusion.
Hand search
Finally, after identifying five articles for inclusion, we conducted a hand search of the journals in which each of the five articles was published (i.e., Exceptional Children, Journal of Special Education, Journal of Learning Disabilities, and Journal of Human Resources). In addition, we hand-searched Remedial and Special Education and Exceptionality, given their focus on special education and their relevance to the field. A hand search of each journal was conducted of all issues published from 2013 to the present. In 2013, a seminal article by Feng and Sass (2013) was published that used administrative data to examine issues in special education. Thus, beginning our search in 2013 allowed us to examine 10 years of articles beginning in a year in which a foundational article was published, which may have spurred similar research in the fields of special education, and educational studies more broadly. Two articles were found eligible for inclusion after this search (Barrett et al., 2020; Kleinert et al., 2015).
Included articles
Ultimately, seven studies were included in the review. All three authors agreed on their inclusion based on multiple reviews of the title, abstract, and full text of each article. They also made final decisions to exclude a handful of final articles. All such decisions were based on adherence to specified inclusion and exclusion criteria (see Table 2). For example, papers by Hemelt et al. (2021) and Hemelt & Ladd (2017) focused on the impact of paraeducators on student outcomes, yet neither paper specifically focused on paraeducators as a part of specialized instructional delivery, nor did they tease apart findings for students with and without disabilities. In addition, a study by Rabren et al. (2003) was excluded because although it used administrative data from the state’s school system and Department of Rehabilitation Services, they limited their sample to individuals who participated in data collection explicitly designed for research purposes, which significantly restricted their sample. In contrast, while Kleinert et al. (2015) restricted their sample to students eligible to take alternate assessments, their sample across multiple states reflected 100% of the students in that group, thus reflecting a population-level administrative data set focused on SWDs.
Analysis
Once the study team came to consensus on the seven articles included in the review, each author individually reviewed the study and took note of the basic characteristics of each article (e.g., sample, methods, findings) as well as overall themes, questions, and implications of each study related to the research questions. The three authors met to discuss each of their individual findings and iteratively came to consensus on the study characteristics, themes, and conclusions.
Results
Overview of Studies
Table 3 provides a summary of the seven studies included in the review. Four of the included studies (Cole et al., 2021, 2023; Schifter, 2016; Theobald et al., 2019) examine the relationship between student outcomes and three distinct categories of time included in general education: 80% of the school day, 40% to 79%, and under 40%. One study (Barrett et al., 2020) uses a continuous measure of percent of time spent in general education summed across three consecutive years. Kleinert et al. (2015) examines not only three categories of inclusive placement, but five: separate school; self-contained classroom; majority self-contained classroom with inclusion in general education classrooms for less than 40% of the day; majority resource classroom with inclusion in general education classrooms for less than 40% of the day; and inclusion in general education classrooms for more than 80% of the day. Finally, only one study, Jones and Winters (2022), examines a specific service delivery model within a specific setting: co-teaching in the general education setting. The authors use administrative data from the Massachusetts state department of education to estimate the impact of co-teaching on students with and without disabilities. Jones and Winters (2022) specify that their definition of co-teaching occurs in the general education setting, and thus their assessment of the model is limited to the ways in which it is implemented in inclusive general education settings, although it may also be a service delivery model implemented in more restrictive settings.
Summary of Included Studies.
Note. ELA = English language arts; AAC = augmentative and alternative communication.
Both the strengths and limitations of administrative data are documented through the seven studies’ research designs. In all of the included studies, authors leverage population-level data to answer research questions using either descriptive (Barrett et al., 2020; Kleinert et al., 2015; Schifter, 2016; Theobald et al., 2019) or quasi-experimental designs (Cole et al., 2021, 2023; Jones & Winters, 2022). Descriptive studies can reflect the correlational relationship between two components (e.g., a service delivery model and student outcomes), which provide important context, especially from a policy perspective, on the state of education (Loeb et al., 2017). Causal studies, however, leverage variation through experimental or quasi-experimental designs to estimate the effect of one element on another (e.g., the effect of a particular service delivery model and student outcomes). It is important to take a rigorous approach to assessing causality in such studies, however, as there are key design considerations that researchers must address to make valid causal claims. Three studies in this review use quasi-experimental designs and the authors make causal claims, which should be critically examined, about the impacts of different service delivery models. The two studies by Cole et al. (2021, 2023) leverage propensity score matching (PSM) to compare students with similar profiles—as measured by available covariates—in more and less inclusive settings. Jones and Winters (2022) use a series of student and school fixed effects to compare students’ academic performance in years in which they are in co-taught classrooms in comparison with years when they are not. Across all seven studies, authors combine student profile data (e.g., their demographic and disability status data) with detailed data on their academic and other outcomes to investigate, with sufficient sample sizes, particular patterns by student disability category. This strength is particularly salient not only in the Theobald et al. (2022) study, in which the focal sample is students with learning disabilities, but also in the study by Kleinert et al. (2015), which focuses on students who are eligible to take alternative assessments. The study by Kleinert et al. (2015) is a unique contribution to the literature because most studies using administrative data analysis focus on students who are eligible for and take standardized rather than alternative assessments.
Service Delivery Models
One of the main findings that we can conclude from the literature in these studies is twofold: first, that inclusion (as defined by time spent in general education classroom settings) is an increasingly prevalent service delivery model for the majority of SWDs, and second, that increased inclusion is associated with positive student outcomes. This finding parallels extant literature on historical patterns in service delivery (Williamson et al., 2020), as well as that which highlights the benefits of inclusion for SWDs (Hehir et al., 2016). Among the studies in this review, the study by Jones and Winters (2022) was unique in that, that it was the only study that examined a specific service delivery model without relying on a proxy of time spent in general education as a marker of inclusive, resource room, or self-contained service delivery models. Studies by Kleinert et al. (2015) and Barrett et al. (2020) were additionally unique; Kleinert et al. (2015) examined five models of service delivery aligned to time spent in general education, while Barrett et al. (2020) studied cumulative time in inclusive settings across multiple years.
Student Outcomes
The studies included in our review covered a range of student-level outcomes, including test and non-test measures as well as outcomes within and beyond the K-12 school setting. These included mathematics and English Language Arts (ELA) test scores (Barrett et al., 2020; Cole et al., 2021, 2023; Jones & Winters, 2022), expressive communication, mathematics and reading skills, alternative communication systems (AAC) use (Kleinert et al., 2015), high school graduation (Cole et al., 2023; Schifter, 2016; Theobald et al., 2019), high school absences (Theobald et al., 2019), and postsecondary employment (Theobald et al., 2019). While some of the outcomes used in these studies are shared across studies (e.g., five studies examined math and ELA outcomes, and three studies examined high school graduation), there is not yet a substantial body of literature across any one of these outcomes to make generalizable claims across contexts. Furthermore, this list inarguably leaves out numerous other outcomes of interest (e.g., discipline rates) that prior qualitative and small-scale quantitative studies have examined. See Hehir et al. (2016) for a review of some recent research and case studies of the positive impacts of inclusion, including improved academic, social–emotional, and employment outcomes, which administrative data analyses using similar outcomes could complement by providing evidence of the impacts of inclusion on student outcomes at scale.
Study Findings
There was a clear pattern in overall findings across the seven studies. The studies found positive relationships between inclusion and student outcomes, leveraging a range of data sets, analytic strategies, and outcome variables. In the Cole et al. (2021, 2023) studies, students who were placed in general education for greater than 80% of their school days performed higher on standardized tests in mathematics and ELA. They were also more likely to receive diplomas representing more rigorous course taking patterns in high school. The study by Theobald and colleagues (2019) reflected similar patterns: students who spent more time in general education exhibited higher rates of on-time graduation, college attendance, and employment than peers who spent less time in general education. Schifter (2016) examined the likelihood of on-time graduation across several disability categories, specifically investigating the difference in graduation rates among students who were fully included and their peers. Across all disability sub-categories, students who were fully included were more likely to graduate on time. Barrett et al. (2020) found that students who are included in general education classrooms for more cumulative time over three consecutive years were more likely to demonstrate higher standardized test scores in mathematics and ELA. Similarly, Kleinert et al. (2015) demonstrated that this pattern holds for students eligible for alternative assessments, who are most likely to be students with higher support needs. The expressive language and mathematics and reading skills of students in their sample were higher for students in more inclusive settings, while use of AAC was lower for students in more inclusive settings. Finally, Jones and Winters (2022) saw small positive effects of being placed in a co-taught class, by about 0.016σ in ELA and about 0.026σ in mathematics. Their results were robust to the use of both general education classes and self-contained special education classes as the counterfactual condition (i.e., co-teaching was more effective than both other conditions, when comparing students with themselves when enrolled in classrooms with other service delivery models).
Discussion
The goal of this systematic review was to summarize existing literature on the impact of service delivery models on student outcomes, as measured in studies that leveraged state administrative data sets. Our hope was that our synthesis would lead to some clearer answers about whether and how service delivery models differ in their effectiveness. At the same time, leveraging administrative data to answer questions about service delivery models is not a silver bullet. There are constraints in these data that become apparent when considering how the findings from these studies might be taken up in practice. First and foremost, these data provide little detailed information about specific service delivery models and what the trade-offs of each might look like. The data do not provide guidance surrounding variation within a specific service delivery model, such as co-teaching, wherein a number of different approaches have been proposed in the literature. It is clear from reviewing these studies that what administrative data can tell us about service delivery models is shaped by the variables that states include in their data sets. In most cases, details related to special education services are quite limited, a point we return to in the “Implications” section. Throughout the following section, we elaborate upon the results and posit some key contributions and next steps for research in this area.
One of the major points of consideration given the results of this review is around the methodologies used across each study. The majority of the studies included in this review leverage descriptive quantitative methods (e.g., correlations, regression, analysis of variance [ANOVA]) and do not estimate causal effect sizes. While descriptive research in education is much needed and is significantly additive to both research and practice in the field (Loeb et al., 2017), there are limitations to the conclusions that can be drawn from such analyses. For example, identifying a linear relationship between time spent in inclusive classroom settings and outcome variables related to normative academic performance (e.g., mathematics and ELA scores, expressive language), as is done in studies by Barrett et al. (2020) and Kleinert et al. (2015), is unsurprising given that IEP teams consider such aspects of students’ in-school performance in making their placement and service decisions. As such, we are limited in making claims about the impact of placing students in one setting or another. To address this limitation, Cole et al. (2021, 2023) and Jones and Winters (2022) use quasi-experimental methods to develop what they claim are causal estimates. It is important to note, however, that the methods used in these studies require very specific conditions in order for effect size estimates to be convincingly causal (e.g., Angrist & Pischke, 2009; Millimet & Bellemare, 2023; Steiner & Cook, 2013). For example, Jones and Winters (2022) seek to compare students with themselves in years in which they do or do not get exposed to a co-teaching model in a combination of ways; they leverage two-way fixed effects, estimate both a fixed effect and lagged dependent variable model, and control for a number of potentially confounding covariates (e.g., student demographic characteristics). In their analysis, they use all alternative settings to co-teaching as a counterfactual condition and include both combined and separate estimates comparing co-taught general education service delivery with general education or self-contained service delivery without co-teaching. As acknowledged in their study, however, there are several assumptions inherent in their model estimation and thus in their conclusions, the critiques and considerations of which are essential to acknowledge.
As another example, the two Cole et al. studies rely on PSM to allow for comparisons between students in more and less inclusive settings. While PSM is a quasi-experimental method, it relies heavily on observed covariates available within a data set. Cole et al. describe their matching strategy in detail and address predicted concerns about the validity of the comparison groups, sample restrictions, and the relationships between certain service delivery models and outcomes, which may be unaccounted for (e.g., concerns about matching validity or selection bias, which may influence the estimated relationship between inclusion and student outcomes). Nonetheless, the method of PSM is highly sensitive to the availability and inclusion of observed variables (e.g., Steiner & Cook 2013; Steiner et al., 2011). Furthermore, their sample restrictions bring to mind some questions about match fit. With PSM, “true” selection into the treatment (in this case, more inclusive settings) is likely made based on both observed variables and latent constructs (Rubin, 2006, 2008). This is particularly problematic when the omitted covariates are ones that are likely key factors determining selection into the treatment. It is undeniable that student placement decisions by IEP teams consider a much fuller set of variables than the limited information in administrative data sets, thus we cannot rule out that the things that distinguish who is and is not placed in more inclusive settings are also the things contributing to differences in the observed outcome variables.
Each of these studies undoubtedly serves as a contribution to the literature and demonstrates a strong attempt to answer the questions we seek to know about service delivery models and student outcomes. And although a methodological critique of the studies in this review is beyond the scope of our study, it is important to reflect on and consider the ways in which each study reveals both the promise and the challenge of using administrative data to address questions of policy and practice in special education, especially causal questions. Using studies by Cole et al. (2021, 2023) as an example, the promise is that the authors can look across a large number of students and identify the sub-group of students who look most similar to one another on the covariates captured in the state administrative data. The challenge is that administrative data sets are limited in the number, type, and specificity of variables available to researchers (Conaway et al., 2015). The infancy of this type of work in special education means that there is unexplored territory in terms of methodology. Along with more comprehensive descriptive research—such as studies by Barrett et al. (2020), Kleinert et al. (2015), Schifter (2016) and Theobald et al. (2019)—that could be done across different contexts, more advanced causal applications could help in identifying the impacts of different service delivery models on outcomes for SWDs. Although the focus of this study is on service delivery models and student outcomes, there is also great potential for leveraging administrative data to explore topics, such as teacher effectiveness, teacher labor markets, and special education placement, which can supplement a growing body of literature that explores topics pertinent to special education in such a way (e.g., Ballis & Heath, 2021; Gilmour et al., 2022; Theobald et al., 2022). These future studies can and should come from a range of theoretical frameworks and fields, including but not limited to economics, sociology, and psychology. While there is often a hyperfocus on econometric methods and economic thinking in policy analysis, including in the field of education (Berman, 2022), the broad ranging nature of administrative data makes it a perfect vehicle for necessary interdisciplinary collaboration.
Additionally of note, administrative data are particularly useful in its ability to combine data from multiple sectors (e.g., public health, employment, and voting records) that could extend our understanding of how different service delivery models and aspects of special education shape students’ experiences and outcomes not just within K-12 schools, but beyond their time in school as well. To go a step further, the field might benefit from replications and extensions of the studies in this review. For example, if replicating and expanding upon the findings of Cole et al. (2023), one might seek to merge Indiana department of education data with administrative data tracking postsecondary enrollment in the state (as Theobald et al., 2019 do in their study) or postsecondary employment in the state to determine the relationship between inclusion and graduation rates, college enrollment, and later employment. Best practices in secondary data analysis, including the principles of Open Science, should be considered in such studies (Lombardi et al., 2023). Thus, each study included in this review represents a useful contribution to the field by building a foundation of literature. There is not yet the depth and breadth of this type of research as there is for other areas of special education research, which allow for more advanced syntheses of findings, such as meta-analyses. As such, more such research is needed.
As we note in the introduction, special education has been defined by the wide variation in how services get enacted in schools. Administrative studies can complement existing work, which has tended to rely on smaller scale studies, by providing a bird’s eye view on service delivery models across an entire state or district’s student population. They provide the opportunity to monitor patterns in service delivery over time (as well as changes in relationships between different models and student outcomes over time) for a population-level sample of students in a natural context. The seven studies included in this review (Barrett et al., 2020; Cole et al., 2021, 2023; Jones & Winters, 2022; Kleinert et al., 2015; Schifter, 2016; Theobald et al., 2019) point to significant, positive effects of inclusion. Yet, we assert that the field would benefit from additional studies that attend to variation in service delivery for SWDs. For example, if we were to extend the questions asked of co-teaching by Jones and Winters (2022), we may consider, how does the effectiveness of co-teaching differ among different teaching partnership models (e.g., one general education teacher and one special education teacher, one general education teacher and one teacher certified to teach English learners (ELs), or one teacher and one paraeducator)? In addition, how might the effectiveness of those models differ in a self-contained setting versus a general education classroom setting? How might the effect sizes found by Jones and Winters (2022) differ in a southeastern state? How might they differ when inclusive of non-test measures? Given that student placements along the LRE continuum differ regionally and among different profiles of students (e.g., race/ethnicity, gender, socioeconomic status, disability category; Chakraborti-Ghosh, 2010), these and more questions must be answered to further examine how different service delivery models work in practice, for whom, under what conditions, and for which outcomes. This is not only a methodological or academic exercise, but also one that has practical consequences for students and school practitioners, with implications for educational equity and students’ rights to access inclusive and effective educational services.
Limitations
There are numerous limitations to this review, in addition to the limitations stated previously for administrative data analysis in general. One of the major limitations is that the inclusion criteria were relatively narrow in study design. We identified a number of studies as relevant and peripheral, but ultimately excluded, and many studies could have been included in a more expansive review of, for example, large-scale quantitative data analysis focused on the relationship between service delivery settings and student outcomes (e.g., Lombardi et al., 2013). Alternatively, there may have been numerous more studies included that use administrative data to examine outcomes for SWDs that have more to do with special education in general than with service delivery settings (e.g., Ballis & Heath, 2021; Schwartz et al., 2021). Further still, expanding our scope internationally would have increased the number of articles that fit our criteria (e.g., Contreras et al., 2020), and would have provided an interesting comparison across settings. The limited scope of our review was driven by our desire to focus explicitly on the ways that researchers have used full-scale, population-level data to describe and compare service delivery settings for SWDs, and to explore their relationship with student outcomes. Thus, while this narrow scope limits the number of articles included and the conclusions that we can draw, it does shed light on the newness—and, hopefully, the importance—of this type of work in the United States and provide some models for replication and extension.
Implications
How much attention should we pay to this study’s findings when the conclusion is relatively straightforward (inclusion tends to be associated with small, positive effects) and the number of studies included (seven) is quite small? First and foremost, the limited number of studies—and the limited scope of these studies—is itself a finding. The implication for researchers is that we simply need more of this kind of work. We need to broaden the number of scholars in special education conducting studies on administrative data sets and, to reflect full population of SWDs in the United States, we need to broaden the number of states from which these kinds of studies are conducted.
There is also an important implication surrounding how states collect and manage their administrative data sets. Collectively, our field’s ability to monitor patterns in service delivery models over time and to estimate the effectiveness of these models depends on how these data systems are set up. Which variables are included? Which variables are mandatory versus which are optional for districts to complete? How are response categories determined and operationalized? It is not a coincidence that the studies in our review all focused on the question of inclusion and that they largely made comparisons between three categories of time in general education (< 40%, 40%–79%, > 80%). This was because states commonly report time in general education by these three categories.
If we want to tell a richer, more detailed story of service delivery, we need data that more closely reflect students’ academic experiences. For example, data systems fall short in being able to tell us whether students receive “push-in” support or how much time they spend receiving intervention in a “pull-out” setting, yet these are common service delivery model options employed in schools across the United States. When it comes to linking school personnel to students, the data systems privilege general education practices, where a group of students are tied to a single “teacher of record.” This is in contrast to special education, where individual students receive a range of services often from a range of personnel (e.g., teachers, paraeducators, related service providers). Some state administrative systems make efforts to more closely map students to specific class “sections.” By this same logic, it would go a long way for states to provide more detailed information on all personnel tied to these same class sections.
Even with the data we currently have, there are still opportunities to expand on the findings reported in this review. For one, there is value in replicating the findings across state contexts to better reflect the diversity of U.S. schools. But also, given historical patterns where students of color have been systematically assigned to less inclusive settings, it will be worthwhile for researchers to more clearly focus on issues of race and equity in their analyses.
Directions for Future Research
Limited number of author teams and data sources
With such a small number of included articles, it is inevitable that the number of author teams and number of states represented is so limited. But, in each case, there are reasons that we might want to encourage the field to broaden. One challenge for the field of special education in advocating for research leveraging administrative data sets is that so few scholars (a) are trained in the methods needed to work with these data and (b) have fostered partnerships with departments of education that would lead to such data analysis. Similar concerns have been raised by scholars in encouraging the field to expand the number of rigorous studies focused on special education teacher preparation (Brownell et al., 2020); there are simply a much larger number of special education researchers focusing on classroom-based interventions. A byproduct of the limited number of author teams doing work in this area is that the knowledge produced by these researchers represents a small number of states. In this case, we see Indiana, Washington, and Massachusetts explicitly represented. To increase the external validity of findings from administrative data sets, we encourage a larger number of states be represented in future research on service delivery models.
Opportunities for replication
Many in special education have expressed the need for the field to pivot toward the increased use of replication (e.g., Coyne et al., 2016; Travers et al., 2016). Replication gives us an opportunity to better understand whether findings from one context generalize to others, which is worth exploring in future work. Large survey data, such as the NLTS2, have been used to develop insights and replications about key questions in special education (e.g., Lombardi et al., 2022), but such analyses do not fill the same role nor allow for the same population-level conclusions as do analyses using administrative data. Replications would provide an opportunity for cross-state analyses of shared research questions, as well as meta-analyses that compile evidence from numerous studies of a similar nature.
Conclusion
Special education is a unique field, especially for policy analysis, because it is both inherently individual (e.g., the individualized nature of services specified in each student’s IEP) and simultaneously universal (i.e., IDEA is a federal law that is applied across states and contexts). Administrative data give us an opportunity to examine how special education is being enacted in practice, at-scale, which complements extant literature on how service delivery can and should work for SWDs. Future analyses of administrative data should situate themselves in the context of the rich body of research on inclusion and service delivery for SWDs. Special education researchers are uniquely positioned to do this type of research given their comprehensive foundational knowledge about the nature of the field and—often gained from prior special education teaching experiences—the ways that special education is enacted in practice. Yet, the small number of studies in this review reflects that special education scholars are not necessarily taking up this opportunity. It is imperative that future research on this topic be conducted by those with content expertise in special education, and we see this as an opportunity for special education scholars to examine the state of special education from the bird’s-eye, longitudinal lens afforded by administrative data.
This lends itself to interdisciplinary collaboration—both within and across schools of education—which we see as another important avenue for future research. One of the benefits of administrative data is that they are collected across numerous institutions, providing data related not only to education, but also to public health, safety, employment, and numerous other realms. This provides researchers with an opportunity to collaborate across these intersecting fields in a way that can provide additional context to a greater totality of life experiences for SWDs, their peers, their families, and the school staff who support them. This may be especially useful for examining postsecondary outcomes and opportunities for SWDs (e.g., studying the use of postsecondary vocational programs; Rabren et al., 2003).
In sum, these seven studies—while limited in number and in scope—each provide useful guidance for future research on service delivery models. They offer strong examples for how to document patterns across students over time and how to resolve some of the selection challenges that inevitably emerge when trying to compare students’ experiences in different models. There is much more work to be done. We encourage researchers to take advantage of the unique affordances offered by administrative data and to find new ways to capitalize on the many different administrative systems housed by states. This review may not offer a clear answer to the question of whether some service delivery models tend to be more effective than others, but it does provide the beginnings of a roadmap for how that question could be better answered over time.
Footnotes
Acknowledgements
The authors thank Leanna Stiefel and Allison Gilmour for the opportunity to contribute to this special issue. The authors also thank generous colleagues, including Elizabeth Bettini and Josefina Senese, for their support throughout the writing process.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.va
