Abstract
Regression discontinuity (RD) designs are commonly used for program evaluation with continuous treatment assignment variables. But in practice, treatment assignment is frequently based on ordinal variables. In this study, we propose an RD design with an ordinal running variable to assess the effects of extended time accommodations (ETA) for English-language learners (ELLs). ETA eligibility is determined by ordinal ELL English-proficiency categories of National Assessment of Educational Progress data. We discuss the identification and estimation of the average treatment effect (ATE), intent-to-treat effect, and the local ATE at the cutoff. We also propose a series of sensitivity analyses to probe the effect estimates’ robustness to the choices of scaling functions and cutoff scores and remaining confounding.
Keywords
1. Introduction
In educational assessment, there have been ongoing efforts to include English-language learners (ELLs) and students with disabilities (SDs) in the National Assessment of Educational Progress (NAEP) assessment by providing appropriate testing accommodations (National Research Council, 2002). Among various testing accommodations, the extended time accommodation (ETA) is the most frequently offered accommodation in the NAEP assessment and other testing programs (Gregg & Nelson, 2012); the recent 2017 NAEP assessment included about 90% of ELLs and SD students, and about 10% of these students received ETA (National Center for Education Statistics [NCES], 2017a, 2017b). Despite the common usage of ETA, there is little guidance on how to evaluate ETA and no systematic research studies assessing ETA’s effectiveness (Jonson et al., 2019). Given that it is unethical and impractical to conduct a randomized experiment in this setting, the goal of this article is to propose a regression discontinuity (RD) design with an ordinal running variable for evaluating program effectiveness.
RD designs have been used for policy and program evaluation where subjects’ treatment status is determined by whether their treatment assignment variable (also called running or forcing variable) exceeds a predefined cutoff. If the running variable is continuous, as required by standard RD designs, the average treatment effect (ATE) at the cutoff is nonparametrically identified and can be estimated by comparing the average outcomes of subjects “just below” and “just above” the cutoff (Hahn et al., 2001; Imbens & Lemieux, 2008; Lee & Lemieux, 2010). However, in some settings, the running variable is discrete, reported in coarse intervals, or an ordinal variable with a few categories only. For instance, in NAEP assessments, student eligibility for ETA is determined by ELL English-proficiency scores, an ordinal variable with six categories: No Proficiency, ELL Beginning, ELL Intermediate, ELL Advanced, Formerly ELL, and Never ELL. Students with ELL Advanced (here, the cutoff) or lower proficiency level are offered ETA. Due to the running variable’s discrete and ordinal scale, the ATE at the cutoff is no longer nonparametrically identified because in the close vicinity of the cutoff score only ETA eligible students are observed (i.e., there is no overlap of eligible and ineligible students even in the limit at the cutoff). Thus, the identification of the ATE at the cutoff requires an appropriate scaling of the ordinal categories together with a correctly specified parametric outcome model to extrapolate the average control outcome of ineligible ETA students to the cutoff category of ELL Advanced students.
In this article, we extend the identification and estimation strategy proposed by Lee and Card (2008) for discrete running variables to ordinal variables. With ordinal running variables, RD designs face several challenges: First, the categories of the ordinal running variable need to be mapped onto a numeric scale by choosing an appropriate scaling function. For instance, using the ranks {1,…, 6} is a possible but not necessarily the only choice for the six ELL English-proficiency categories above. Second, a meaningful cutoff score between ELL Advanced and Formerly ELL must be determined with respect to the chosen numeric scale. The choice of the cutoff score determines the target population (at the cutoff) as well as the magnitude of the ATE. Third, the parametric functional form to the left and the right of the cutoff score must be correctly specified with respect to the chosen numeric scale, such that extrapolations to the cutoff score are accurate. Fourth, the statistical uncertainty due to the discreteness of the scaled running variable should be accounted for when estimating standard errors (Lee & Card, 2008). Fifth, given the increased number of assumptions with ordinal running variables and the potential violation of these assumptions, we need to conduct a set of sensitivity analyses to strengthen causal conclusions drawn from the analysis. These include probing the effect estimate’s sensitivity to (i) the choice of the scaling function, (ii) the choice of the scaled cutoff score, and (iii) remaining confounding due to model misspecification (see Section 3 for details). We remark that there are other approaches that can accommodate ordinal or discrete running variables in RD designs, such as using propensity scores as a surrogate continuous running variable under the local randomization framework (Li et al., 2021) and linking analysis of covariance and local randomization heuristics (Sales & Hansen, 2020).
Throughout this article, we use the ETA example based on the 2017 NAEP data to demonstrate our proposed approach. We discuss some points that investigators should consider when choosing the scaling function and the cutoff score. We then analyze the intent-to-treat (ITT) and the local ATE (LATE) that accounts for noncompliance with respect to the assigned ETA status.
The remainder of this article is organized as follows. Section 2 briefly reviews the sharp and fuzzy RD framework with a continuous running variable. Section 3 discusses the assumptions required for RD designs with an ordinal running variable. Section 4 describes the analyses and results of our empirical example concerning the effects of ETA using the 2017 NAEP data for mathematics; this section also includes the aforementioned sensitivity analyses to the assumptions laid out in Section 3. Conclusions are given in Section 5.
2. Setup
2.1. Notation
We use the Neyman–Rubin potential outcomes framework (Neyman, 1923; Rubin, 1974) and its extension to multilevel/clustered data by Hong and Raudenbush (2006) to define treatment effects in sharp and fuzzy RD designs. The NAEP sampling design involves a multilevel structure, where students, the study units, are nested within schools. Let
2.2. Review: Sharp RD Designs With a Continuous Running Variable
We first review the standard sharp RD design with a continuous running variable. Suppose our running variable, ELL English proficiency, is continuous where students scoring below or at the cutoff,
Because
(A1) Local Continuity of Potential Outcomes:
This assumption states that the mean potential treatment and control outcomes right below the cutoff are equal to the corresponding mean potential outcomes right above the cutoff. The assumption allows us to think of an RD design as a local randomized experiment where students near the cutoff are randomly assigned to treatment and control conditions (Lee & Lemieux, 2010). Under (A1), the ATE at the cutoff,
In our setting, the ATE at the cutoff represents the average effect of ETA for students scoring right at the eligibility cutoff. We can estimate
2.3. Review: Fuzzy RD Designs With a Continuous Running Variable
In practice, study administrators frequently do not adhere to the assignment rules or participants do not comply with the assigned treatment or control status. For instance, students eligible for ETA according to their English-proficiency scores may not receive ETA and ineligible students might actually receive ETA due to school- or administrator-specific rules or exemptions. In the presence of noncompliance, we have a fuzzy RD design that can identify the ITT and LATE at the cutoff score. Using the ETA eligibility status,
(A2) Local Monotonicity:
(A3) Local Exclusion Restriction:
The local monotonicity assumption rules out the presence of defiers at the cutoff, that is, students who would receive ETA if not eligible for ETA but would not receive ETA if eligible. Assumptions (A1)–(A3) allow us to identify the ATE for the latent subpopulation of compliers at the cutoff, that is, students who would receive ETA if they were eligible for ETA and who would not receive ETA if ineligible (i.e.,
In our study,
2.4. Review: A Graphical Perspective
The data-generating process underlying RD designs can be formalized by a causal diagram—a directed acyclic graph (DAG; Elwert, 2013; Morgan & Winship, 2014; Pearl, 1988, 2009; Steiner et al., 2017). The DAG for the ETA evaluation is shown in Figure 1(a). The graph highlights that the ELL eligibility status (A) is solely determined by a student’s ELL English-proficiency score (X), while the ETA receipt status (Z) depends on the eligibility status (A) and on observed and unobserved covariate sets (

Causal directed acyclic graph (DAG) and causal graphical identification for evaluating the effects of ETA. ETA Eligible (A) represents students’ ETA eligibility status. ELL EP (X) represents ELL English proficiency. ETA Received (Z) represents whether students received ETA or not, and Outcome represents students’ math-proficiency outcome.
Although the graph suggests that conditioning on ELL English proficiency (X) blocks the confounding backdoor paths between ELL eligibility (A) and students’ outcome (Y),
1
the causal effect of A on Y is nonetheless not identified because the positivity assumption is not met. That is, for each value of the running variable X, we only observe either eligible or ineligible students, but never both (i.e., there is complete lack of overlap). The causal effect of ETA receipt (Z) on the outcome (Y) would be identified conditional on
Figure 1(b) shows the causal graph for the RD design at the limiting cutoff score,
3. RD Design With an Ordinal Running Variable
So far, we have assumed that the running variable X is continuous. However, in practice, many RD designs rely on a discrete metric running variable, such that the causal effects at the limiting cutoff are no longer nonparametrically identified. The identification of causal effects then requires parametric functional form assumptions to bridge the gap between the neighboring discrete values at the cutoff (Lee & Card, 2008). With ordinal running variables, as for our ETA study, causal identification is even more challenging because the ordinal categories first need to be mapped onto an appropriate numerical scale. In this section, we discuss the identification and estimation of causal effects from RD designs with an ordinal running variable. Drawing valid and reliable causal conclusions from such RD designs requires three main steps. First, researchers need to decide on a reasonable scaling function for the ordinal running variable. Second, they need to correctly specify the outcome regression with respect to the chosen scaling function and estimate standard errors that reflect specification errors due to the running variable’s discreteness. Third, given uncertainties about the appropriate scaling and correct model specification, researchers should always conduct a set of sensitivity analyses to probe the conclusions’ robustness to (i) the choice of the scaling function, (ii) the choice of the scaled cutoff score, and (iii) remaining confounding due to model misspecification. The following subsections describe these steps in detail and state the causal identifying assumptions.
3.1. Scaling Function
The scaling function
There are a variety of scaling functions that map ordinal categories onto real numbers. First, ordinal categories can be arranged in ascending order, and their ranks are used as scale values (Crocker & Algina, 2006). For instance, in our ETA study, ELL English proficiency has six categories ranging from No Proficiency to Never ELL. Thus, No Proficiency translates into a scale value of
However, rank-based scale values might be a poor choice when equal distances between consecutive categories are not a suitable representation of the ordinal levels in the running variable. In this case, researchers should consider optimal scaling techniques that use observed variables directly related to the categories of the ordinal running variable (Bradley et al., 1962). Such a variable could be the underlying continuous variable that measures the same proficiency/skill or a close proxy thereof (e.g., a continuous composite English-proficiency score, domain scores like reading and writing, or English-proficiency scores from previous grades). Then, the optimal scale values can be determined using optimal scaling methods for categorical data like categorical regression or categorical principal components analysis. Here, optimality is often defined by maximization of variance, maximization of pairwise linear relationships, or maximization of homogeneity among variables, to name a few. Also, if multiple categories receive similar scale values, we can collapse these categories into one category (Bradley et al., 1962; Machines, 2019; Meulman, 1998; Meulman et al., 2019).
In the absence of related continuous variables that could be used for optimal scaling techniques, it is sometimes possible to infer scale values from external sources like published cut scores used to form the categories. Specifically, if there were clearly defined cut scores on the ELL English-proficiency variable, then the midpoints of the cut scores could be used as scale values. However, the midpoints of the lowest and highest category might not be meaningfully defined if the minimum or maximum of the ELL proficiency score is not known or is a very extreme value that is rarely observed. If cut scores for the ordinal running variable are not known, we can use other related classifications based on the same or similar underlying continuous variable. For example, the Wisconsin Department of Instruction defines English-language-proficiency classifications with seven ordinal levels (ELL Beginning Preproduction, ELL Beginning Production, ELL Intermediate, ELL Advanced Intermediate, ELL Advanced, Formerly ELL, and Never ELL) that are very similar to the observed ordinal categories of English proficiency in NAEP. After mapping the different ordinal categories, we could use their cut scores and corresponding midpoints to obtain scale values.
3.2. Sharp RD Designs With an Ordinal Running Variable
As mentioned before, with an ordinal running variable, nonparametric causal identification breaks down because in the close vicinity of the cutoff score (i.e., the ELL Advanced category) only students eligible for ETA are observed and there are no ineligible control students. Given the identification failure at the limiting cutoff, the limiting graph in Figure 1(b) no longer applies either. Thus, we are back to the graph in Figure 1(a), which indicates that the observed and unobserved sets of covariates
To this end, we extend the approach suggested by Lee and Card (2008) for discrete running variables to ordinal running variables. Instead of the local continuity assumption (A1), we now assume a correctly specified outcome regression, such that the expected control outcome can be correctly inferred by extrapolating the outcome regression for control units to the cutoff category.
(A4) Outcome Regression Function:
Here,
In estimating and conducting inference for
3.3. Fuzzy RD Designs With an Ordinal Running Variable
For fuzzy RD designs with an ordinal running variable, the ITT at the cutoff category is identified and estimated just like the ATE in the sharp RD design, where the treatment assignment status
(A5) Treatment Regression Function:
As before,
3.4. Sensitivity Analyses
The discussions of the assumptions for the sharp and fuzzy RD design revealed that the causal effects at the cutoff are identified only if the scaling function and outcome (and treatment) regressions are correctly specified. Given that the correct specification of the functions is uncertain in practice, researchers should always conduct sensitivity analyses to check the conclusions’ robustness to (i) the choice of scaling functions, (ii) the choice of cutoff values, and (iii) the presence of remaining confounding resulting from a misspecified outcome or treatment regression (that leads to biased extrapolations of the control or treatment outcomes to the cutoff value).
First, in choosing different scaling functions
Second, given a particular scaling function, one may evaluate the effect estimates at different cutoff value xc
between the scale value of the cutoff category
Finally, it is advisable to probe whether the conclusions drawn are sensitive to unblocked confounding due to model misspecification of hS (and gS for the LATE). The effect estimates at the cutoff are unbiased only if the extrapolations are based on a correctly specified functional form hS (and gS ) with respect to the chosen scaling function S. With a misspecified functional form, the extrapolation to the cutoff score may become invalid and fail to completely remove confounding bias between the outcomes of treatment subjects in the cutoff category and the outcomes of control subjects in the neighboring category. For example, suppose that we simply compare the mean outcomes of the neighboring categories Advanced ELL (treated subjects) and Formerly ELL (control subjects) in our empirical example. That is, we make no attempt to remove any confounding between the treatment and control groups. Thus, the resulting unadjusted effect estimate likely suffers from confounding bias due to group differences in ELL English-proficiency categories and in any other student characteristics like ability or the number of English-language books read per month. The RD design with an ordinal running variable tries to overcome differences between the neighboring treatment and control groups by adjusting for covariates through a correct specification of hS (and gS ). Since misspecified functional forms may remove a part but not all the confounding bias, researchers should conduct sensitivity analyses and evaluate the degree of effect variations obtained from the sensitivity analysis. We demonstrate all these sensitivity analyses with our empirical example in the next section.
4. Empirical Example: Testing Accommodations in NAEP
4.1. Data and Variables
NAEP is the largest nationally representative and continuing assessment of what students in the United States know and can do in various disciplines. The data have been collected by NCES within the Institute of Education Sciences. The 2017 NAEP assessments were conducted for Grades 4 and 8 in mathematics, reading, and writing. The NAEP sampling procedures ensure that the students and schools selected in NAEP are representative of the target population. The NAEP assessment strives to minimize participant burden by giving students a subset of items from the total item pool (Johnson, 1992; Oranje & Kolstad, 2019). For more details of the NAEP methods and procedures, see the NAEP page of the NCES website (https://nces.ed.gov/nationsreportcard/tdw/).
In our study, we used the NAEP Grade-4 2017 restricted-use data for mathematics. For the data analysis, we excluded (i) schools with only one student, (ii) SDs, and (iii) ELL students whose prior performance was below the grade level of performance of NAEP; here, ELL students’ prior performance was evaluated by their teachers or school staff members through an ELL questionnaire. After sample exclusion, our final analysis sample consisted of 116,910 students from 7,450 schools (78.2% of the original reporting sample). 2
In the 2017 NAEP data, we used the math proficiency as the outcome
In demonstrating our proposed approach, we only used the last of the 20 plausible values in math proficiency as the outcome
4.2. Choice of Scaling Function and Cutoff Score
To assess the effects of ETA with an RD design, we first chose the scaling function for the ordinal ELL English categories based on their ranks,

Observed means in reading by English-language learner English proficiency. Note. Numbers in parentheses represent sample sizes and are rounded to nearest tens. Source. U.S. Department of Education, National Center for Education Statistics, National Assessment of Educational Progress (NAEP) 2017.
4.3. Parametric Model Specifications
Given two-sided noncompliance, we estimate the ITT and the LATE at the cutoff. Following Lee and Card (2008), we model the categories of the ordinal running variable as random effects in a hierarchical linear model that also accounts for the clustered data structure of students nested within schools. The reading proficiency
In the model equation, the ITT at the cutoff is given by

Regression discontinuity design for evaluating the effects of extended time accommodations in mathematics. Note. NOP = No Proficiency; BEG = ELL Beginning; INT = ELL Intermediate; ADV = ELL Advanced; FOR = Formerly ELL; NEV = Never ELL. The solid black lines represent the estimated regression function, and the dotted gray line represents the extrapolated line from the regression function. Gray points indicate students’ math scores, and open red points indicate covariate-adjusted observed outcome means with cell means coding. Source. U.S. Department of Education, National Center for Education Statistics, National Assessment of Educational Progress (NAEP) 2017.
For estimating the LATE effect at the cutoff, we use an instrumental variable regression or two-stage least squares (TSLS) regression. That is, we treat the model for
In the first-stage regression (2),
4.4. Results
Table 1 summarizes the students’ ETA eligibility, as defined by ELL status, and whether they actually received the ETA. Overall, about 4.2% of the students (4,940 students) in our study sample were eligible for ETA; these are denoted as ELL students in the table. Among those who were eligible for ETA, about 32.5% of the ETA-eligible students (i.e., ELL) actually received ETA. Also, we saw that few students received ETA even though they were not eligible for ETA (i.e., non-ELL), likely due to test irregularities; see more details on the compliance rate by ELL English-proficiency categories in Supplemental Appendix B.
Compliance for Extended Time Accommodations (ETAs) by English-Language Learner (ELL) Status
Note. Numbers are rounded to nearest tens. Details may not sum to a total due to rounding.
Source. U.S. Department of Education, National Center for Education Statistics, National Assessment of Educational Progress (NAEP) 2017.
Figure 3 provides a visual representation of the RD design, where the x-axis represents the rank-based scale values of English-proficiency categories with the cutoff point defined at ELL Advanced and the y-axis represents the math proficiency scores. For each ELL category, covariate-adjusted observed outcome means (with cell means coding) are shown by red circles. Since the observed and predicted means (from the fitted model) are very similar, except for the No Proficiency category with relatively few observations, the linear model provides good fit to the data near the cutoff. We observe that the mean math-proficiency score of Never ELL was similar to that of Formerly ELL, and the mean math-proficiency score increased when ELL English proficiency increased from No Proficiency to ELL Advanced. We can also visually see that the ITT effect at the cutoff is small.
Table 2 summarizes ITT and LATE estimates. We interpret 95% confidence intervals as “compatibility intervals” (Amrhein et al., 2019), that is, the set of possible true effects that are compatible with our data for the given model. As seen from Table 2, the ITT estimate is 2.66 points and small. Also, under the above assumptions, our data are compatible with ITT effects as small as −4.36 points and as large as 6.12 points. The interval suggests that the ITT could be (close to) zero or positive and we do not have sufficient power to reject the null hypothesis of no effect. In contrast, the LATE estimate is larger than the ITT estimate, and all values in the interval are positive; the LATE estimate is 12.99 points, and the interval indicates compatibility with effects between 7.25 and 19.53 points.
Intent-to-Treatment (ITT) Effect and Local Average Treatment Effect (LATE) Estimates of Extended Time Accommodations at the Cutoff
Source. U.S. Department of Education, National Center for Education Statistics, National Assessment of Educational Progress (NAEP) 2017.
Finally, we remark that given that we have only two categories to the left of the cutoff category, the extrapolation of the control outcomes from the Formerly ELL category to the Advanced ELL category strongly depends on the mean math proficiency of the Never ELL category. If students with Never ELL would on average score slightly higher or lower, the effect estimates at the cutoff could change their sign. Thus, the student composition of the Never ELL category is potentially influential on the ETA evaluation. In NAEP, the Never ELL category includes students who (1) are native English speakers, (2) have never been designated as ELL, or (3) were designated as ELL students more than 2 years ago. The inclusion of native English speakers in the Never ELL category is problematic because they do not belong to our target population of interest (i.e., nonnative speakers). Unfortunately, the NAEP data do not provide individual-level information about whether a student belongs to one of the three subcategories. To address potential bias concerns, we did a subgroup analysis only with Hispanic students because a smaller portion of them are native English speakers (de Brey et al., 2019). We found that the ITT effect may be slightly smaller than the ITT with the original sample, while the LATE does not differ much (see results in Supplemental Appendix D).
4.5. Sensitivity Analyses
4.5.1. Methods
We assess the sensitivity of our conclusions against (i) the choice of the scaling function, (ii) the choice of the cutoff value, and (iii) remaining confounding. To probe the results’ sensitivity to different choices of the scaling function, we first decreased or increased the gap between the two neighboring cutoff categories (i.e., ELL Advanced and Formerly ELL) on our original scale. It assesses the effect estimates’ sensitivity to uncertainties in the appropriate scaling between the two neighboring categories at the cutoff. A larger gap increases reliance on extrapolation than a smaller gap, and thus, the former might be more susceptible to model misspecification. We consider two cases, one with a reduced gap and the other with an increased gap. For the reduced gap, we used scale values (−3, −2, −1, 0, 0.5, and 1.5), and for the increased gap, we used (−3, −2, −1, 0, 1.5, and 2.5). All of the scale values were centered at the cutoff of ELL Advanced.
Next, we used a different set of scale values for all categories (and not only the neighboring cutoff categories) based on the cut scores of the proficiency categories from the ACCESS for ELLs exam in World-Class Instructional Design and Assessment (WIDA). 5 WIDA’s (2013, 2019) English-proficiency levels consist of Entering, Emerging, Developing, Expanding, Bridging, and Reaching for ELL students and former ELL students. From the description of WIDA’s proficiency levels, we considered Entering, Emerging, Developing, Bridging, and Reaching as being aligned with No proficiency, ELL Beginning, ELL Intermediate, ELL Advanced, and Formerly ELL, respectively, and we used the cut scores from reading proficiency levels of fourth graders (WIDA, 2013) to determine the scale values. Specifically, based on the cut scores of the WIDA’s proficiency levels, we computed the mean scores (i.e., midpoints) of the first four proficiency levels and used the relative differences between consecutive proficiency levels as scale values. For noneligible students, we assigned the same scale value to the last two categories because Formerly ELL and Never ELL’s reading performance was similar from Figure 2. Ultimately, we used the scale values of (−5.6, −2.2, −1, 0, 1, and 1) for our sensitivity analysis, with 0 indicating ELL Advanced.
We also conducted a sensitivity analysis with respect to the choice of the cutoff point. Specifically, we varied the cutoff value by redefining the cutoff value as the mean between the scale values of ELL Advanced and Formerly ELL; we call the new cutoff point “Between.”
Lastly, for the sensitivity analysis against remaining confounding, we used the general bias formula from VanderWeele and Arah (2011), where the potential bias d arising from unmeasured confounders can be represented by the product
4.5.2. Results
Figure 4 visualizes the sensitivity analysis for ITT estimates by changing the choice of the scaling functions and the choice of the cutoff values. Figure 5 displays the corresponding point estimates and confidence intervals.

Sensitivity analysis against scaling. Note. NOP = No Proficiency; BEG = ELL Beginning; INT = ELL Intermediate; ADV = ELL Advanced; FOR = Formerly ELL; NEV = Never ELL. The solid black lines represent the estimated regression function, and the dotted gray lines represent the extrapolated line from the regression function. Gray points indicate students’ math scores. Source. U.S. Department of Education, National Center for Education Statistics, National Assessment of Educational Progress (NAEP) 2017.

Estimates and confidence intervals of the intent-to-treatment effect and local average treatment effect from sensitivity analysis against scaling. Source. U.S. Department of Education, National Center for Education Statistics, National Assessment of Educational Progress (NAEP) 2017.
From Figures 4 and 5, we observed that all the ITTs are compatible with no effect although the confidence intervals of the ITT vary depending on the combinations of scaling, gap, and cutoff points. In contrast, the LATEs remain positive and the confidence intervals of the LATE are very similar across different choices of the scaling function. 6 Overall, our original conclusions about the ITT and LATE are robust to different specifications of the scaling function and the cutoff value.
Next, we tested the sensitivity of our results to remaining confounding and evaluated the degree of effect variations introduced by negative and positive unmeasured confounding bias. Specifically, we estimated a new adjusted treatment effects
For the ITT estimate, we use a negative bias of
5. Conclusions
In this article, we proposed to use an RD design with an ordinal discrete running variable. We used a scale function S to convert the ordinal levels of the running variable to numeric scale values and modified Lee and Card’s (2008) framework to accommodate the ordinal running variable. We assessed the sensitivity of our results with respect to the choice of the scaling function and the cutoff value. We also assessed the sensitivity of our results to remaining confounding arising from an incorrect scaling function or an imperfect functional form. We demonstrated the proposed approach by investigating the effects of ETA on students’ math performance based on the 2017 NAEP data. Overall, we found that our ITT results of being eligible for ETA at the cutoff are compatible with slightly negative but also positive effects, whereas our LATE results of receiving ETA indicate positive effects. Our sensitivity analyses indicate that our original ITT and LATE are robust to various design-based factors for sensitivity analyses (i.e., the choice of the scaling function, the choice of the cutoff value, and remaining confounding).
Based on our findings, we provide some suggestions for future research concerning evaluation of testing accommodations based on an RD framework. First, Kolesár and Rothe (2018) showed that Lee and Card’s (2008) confidence intervals with standard errors that are clustered by a discrete running variable underestimate coverage properties. Future research would use alternative confidence intervals which guarantee accurate coverage properties, proposed by Kolesár and Rothe (2018). Second, if the NAEP assessment provides continuous ELL English proficiency that was used to determine ETA eligibility, it will enable researchers to estimate the ITT and LATE at the cutoff with less concerns about biases arising from the choice of the scaling function or the functional form of the outcome model. Third, if an ordinal running variable is present and a relevant underlying variable, such as pretest English scores in our setting, is included in the observed data, researchers could use optimal scaling techniques to choose appropriate scale values. Lastly, we did not consider whether students made use of ETA in our study. That is, even if the student received ETA, students may have not needed the extra allotted time. Based on prior works on accommodations, about 40% of the students who received ETA made use of the extra time in the NAEP assessment (Kim & Circi, 2018, 2019). Students’ actual use of ETA can be determined by process data, which are data provided by examinees’ responses to the testing devices while taking the test (Bergner & von Davier, 2018). If such data are available, we may have to use sequential compliance models and it would be interesting to incorporate them into RD designs in order to assess the effect of making use of ETA.
Supplemental Material
Supplemental Material, sj-pdf-1-jeb-10.3102_10769986221090275 - Regression Discontinuity Designs With an Ordinal Running Variable: Evaluating the Effects of Extended Time Accommodations for English-Language Learners
Supplemental Material, sj-pdf-1-jeb-10.3102_10769986221090275 for Regression Discontinuity Designs With an Ordinal Running Variable: Evaluating the Effects of Extended Time Accommodations for English-Language Learners by Youmi Suk, Peter M. Steiner, Jee-Seon Kim and Hyunseung Kang in Journal of Educational and Behavioral Statistics
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by a grant from the American Educational Research Association Division D.
Notes
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
