Abstract
This direct replication study compared the use of dichotomized likelihood ratios and interval likelihood ratios, derived using a prior sample of students, for predicting math risk in middle school. Data from the prior year state test and the Measures of Academic Progress were analyzed to evaluate differences in the efficiency and diagnostic accuracy of gated screening decisions. Post-test probabilities were interpreted using a threshold decision-making model to classify student risk during screening. Using interval likelihood ratios led to fewer students requiring additional testing after the first gate. But, when interval likelihood ratios were used, three tests were required to classify 6th- and 7th-grade students as at-risk or not at-risk. Only two tests were needed to classify students as at-risk or not at-risk when dichotomized likelihood ratios were used. Acceptable sensitivity and specificity estimates were obtained, regardless of the type of likelihood ratios used to estimate post-test probabilities. When predicting academic risk, interval likelihood ratios may be best reserved for situations where at least three successive tests are available to be used in a gated screening model.
Researchers have recently argued for the adoption of evidence-based assessment (EBA) principles to increase the accuracy and efficiency of universal screening in schools (Canivez, 2019; Pendergast et al., 2018; VanDerHeyden & Burns, 2018). A critical aspect of an EBA approach to academic screening is to use base rates and post-test probabilities to determine student risk. VanDerHeyden (2011) argued for the adoption of post-test probabilities to inform universal screening decisions made in schools nearly a decade ago. Although these metrics have started to appear in recent studies of universal academic screening (e.g., Nelson et al., 2017; VanDerHeyden et al., 2017; Van Norman, Klingbeil et al., 2017), many of these studies are limited for two important reasons.
First, most of the recent work has applied cut-scores to one sample of students and used the obtained sensitivity and specificity values to estimate post-test probabilities. However, educators must apply cut-scores and make decisions regarding students whose proficiency status on the criterion measure is unknown. Variation in the base rates of nonproficiency in future cohorts could affect the diagnostic accuracy obtained in future years (Jenkins et al., 2007; Meehl & Rosen, 1955). Cross-validation and replication in academic screening research could provide needed evidence for educators who must evaluate the viability of competing screening practices.
Second, researchers have largely estimated post-test probabilities using likelihood ratios (LRs) derived from sensitivity and specificity estimates, with student performance on the screening and criterion measures classified as a dichotomous outcome. The dichotomization of continuous scores may result in a loss of important diagnostic information (Brown & Reeves, 2003; Deeks & Altman, 2004). Instead, researchers in medicine and psychology commonly classify test performance as an ordinal, multicategory variable. Risk is then estimated for individuals within each ordinal category or interval. Thus, additional research on the application of interval LRs in the context of education may offer useful guidance for screening procedures in schools. In the current study, we sought to replicate our initial evaluation of using interval LRs to estimate post-test probabilities to determine math risk (Klingbeil et al., 2019).
Estimating LRs and Post-Test Probabilities
When applied to academic screening, an LR represents the ratio between (a) the probability of a screening test result for students who will not be proficient and (b) the probability of a screening test result for students who will be proficient on a criterion measure. An LR of 1 indicates the screening test provides no predictive information. LRs greater than 1 indicate the screening test result is more likely for a student who will not be proficient, and LRs less than 1 indicate the screening result is less likely for a student who will not be proficient. When comparing screening tests, a wider range between LRs is indicative of scores that provide more diagnostically useful information (Brown & Reeves, 2003).
LRs can be combined with the pretest (or prior) probability of nonproficiency to determine the post-test (or posterior) probability of nonproficiency using Bayes’ theorem. Applying Bayes’ theorem to calculate post-test probabilities requires several steps. First, LRs must be estimated for each possible outcome on the screening measure. Schools could estimate LRs by retrospectively analyzing their own screening data or by adopting published LRs reported in empirical studies (Youngstrom, 2013). Second, the pretest probability of the condition (e.g., nonproficiency) must be estimated for the setting in which the screening measure will be used. For example, when using a screening measure to identify students who are at-risk for failing a statewide test, historical base rates of nonproficiency in that setting could be used as the pretest probability of failure (Pendergast et al., 2018; Van Norman, Klingbeil et al., 2017). Third, the pretest probability of nonproficiency is transformed into odds. Fourth, the pretest odds are multiplied by the LR associated with the obtained screening test result to calculate the post-test odds of nonproficiency. Fifth, the post-test odds are transformed into post-test probabilities to facilitate interpretation.
Interpreting Post-Test Probabilities
The post-test probability of failure represents the probability that a student who obtained a specific screening test result will not demonstrate proficiency on the criterion measure. Post-test probabilities are interpreted using a threshold decision-making model (Pauker & Kassirer, 1980) within an EBA approach (Youngstrom, 2013). Threshold decision-making divides the range of post-test probabilities into three zones: (a) provide intervention without further assessment, (b) conduct more assessment, and (c) withhold further assessment and intervention (Youngstrom, 2013). VanDerHeyden (2013) recommended that students with post-test probabilities ≥.50 should receive intervention and students with post-test probabilities ≤.10 should not participate in further assessment or intervention. Students with post-test probabilities between .11 and .49 may require further assessment to determine whether intervention is needed. For these students, the obtained post-test probability becomes the revised pretest probability and the process can be repeated using results from another test (Gallagher, 1998).
Methods to Estimate LRs
LRs are typically estimated when test performance is classified as a dichotomous outcome (i.e., at-risk or not at-risk) or as ordinal performance categories (Gallagher, 1998). When screening test results are categorized into at-risk or not at-risk classifications, the application of Bayes’ theorem results in positive and negative post-test probabilities. Positive post-test probabilities indicate the likelihood that a student who was identified as at-risk on the screening measure will not be proficient on the criterion. Negative post-test probabilities indicate the likelihood that a student who was identified as not at-risk on the screening measure will not be proficient on the criterion.
Classifying screening test performance as a dichotomous outcome simplifies decision-making (VanDerHeyden & Burns, 2018), but it may result in the loss of useful diagnostic information (Brown & Reeves, 2003; Deeks & Altman, 2004). That is, when applying Bayes’ theorem, all students who were identified as at-risk on the screening test are treated as having the same level of risk, despite differences in the magnitude of risk conveyed by the screening test scores (Sonis, 1999). In addition, errors in interpreting post-test probabilities are most pronounced for students near the threshold used to dichotomize screening results (Brown & Reeves, 2003).
Many popular academic screening tools provide three- or four-level performance classifications to aid in the interpretation of student performance. A separate LR can be estimated for each interval of test scores (Gallagher, 1998). Interval LRs are interpreted the same as LRs derived from sensitivity and specificity without the differentiation between positive or negative LRs being required. Bayes’ theorem is applied in the manner described above to combine information from the screening test and pretest probability to determine the post-test probability of failure.
Klingbeil et al. (2019) evaluated differences in screening efficiency and accuracy when students’ post-test probabilities of math risk were estimated using dichotomized or interval LRs. The end-of-year statewide achievement test in math was the criterion measure. Screening measures included the prior-year statewide achievement test and the Measures of Academic Progress (MAP; Northwest Evaluation Association [NWEA], 2014). Both measures classified student performance into Below Basic, Basic, Proficient, and Advanced categories, which were used to estimate interval LRs. In comparison, dichotomous LRs were estimated by classifying students who scored in the Below Basic and Basic categories as at-risk and students who scored in the Proficient and Advanced categories as not at-risk. Within each grade, data from the entire sample were used to estimate the interval and dichotomized LRs for each screening measure. Therefore, the sample size and base rates of nonproficiency were constant across both methods of estimating LRs. The resulting LRs are shown in Table 1. More details regarding the estimation of both types of LRs are provided in Supplemental Materials.
Likelihood Ratios Derived Using Data From Prior School Year.
Note. Dichotomized likelihood ratios were estimated using the cut-scores provided by the Wisconsin Department of Public Instruction (Prior Year Forward Exam) or the Northwest Evaluation Association (MAP). ILR = interval likelihood ratio; +LR = positive likelihood ratio; −LR = negative likelihood ratio; MAP = Measures of Academic Progress.
Two of the gated screening models investigated by Klingbeil et al. (2019) are relevant to this study. The first model included the prior year state test, fall MAP, and winter MAP. The second model included the prior year state test followed by the winter MAP. Students with post-test probabilities ≥.50 were considered at-risk and students with post-test probabilities ≤.10 were considered not at-risk (VanDerHeyden, 2013). These students did not continue further in the emulated gated screening framework. Students with post-test probabilities between .11 and .49 went on to the next screening gate. The post-test probabilities from the prior screening gate were used as the revised pretest probability, and the results of the next screening assessment were applied. The final at-risk or not at-risk decision for each student, regardless of the gate in which it was obtained, was used to estimate the diagnostic accuracy of the gated models. Across grades and gated models, sensitivities (range = .79–.89) were generally lower than the specificities (range = .82–.90). The sensitivity and specificity values were within .03 when comparing the results from the interval and dichotomized LRs.
A potential benefit of using interval LRs outlined by Klingbeil et al. (2019) was the reduction of students who required additional assessment after the first gate in the screening process. When dichotomized LRs were used, between 60% and 74% of students required further assessment after the first gate. All students who performed in the not at-risk range had post-test probabilities warranting further assessment. In comparison, when interval LRs were used to estimate post-test probabilities, between 46% and 53% of students required further assessment. Students who performed in the Advanced range had post-test probabilities below .10 and did not require further assessment.
Purpose and Research Questions
Interpreting screening test performance using interval LRs may help schools realize one of the critical benefits of gated screening—the reduction in the number of students requiring further assessment. Yet, the findings of Klingbeil et al. (2019) are limited because the pretest probabilities and the LRs were estimated and then applied to the same sample of students which may have biased the diagnostic accuracy estimates. Estimating pretest probabilities and LRs using extant data and applying them to make decisions about students in subsequent academic years is a closer approximation of how screening occurs in schools. The purpose of this study was to further evaluate the use of interval LRs to estimate post-test probability of math risk among middle school students. Two research questions guided this study:
Method
Setting and Participants
We conducted a retrospective analysis of data that were collected in a large suburban school district in Wisconsin. The setting and inclusion criteria were the same as those in Klingbeil et al. (2019). Data were collected at both district middle schools during the 2017 to 2018 school year. As required by the district, all measures used in this study were administered to all enrolled students, unless the student had an Individualized Education Program specifying participation in an alternative assessment program (n = 13). A total of 1,593 students in Grades 6 (n = 512), 7 (n = 518), and 8 (n = 563) participated.
Approximately 68.7% of students were identified as White, 17.2% as Asian, 6.7% as Hispanic/Latinx, 4.4% as two or more races, and 2.8% as Black. Nearly 15.5% of students were identified as gifted and talented, 10.3% were qualified for free or reduced-price lunch, 8.2% were qualified for special education services, and 1.3% were identified as English learners. A series of chi-square analyses indicated the current sample did not significantly differ from the sample in Klingbeil et al. (2019) in terms of grade, race/ethnicity, sex, and the other demographic characteristics listed above (see Table S1 in Supplemental Materials).
Measures
Forward exam
The Forward Exam is a computerized summative assessment administered statewide. The Wisconsin Department of Public Instruction (Wisconsin DPI; 2018) published a technical report that provides evidence of internal consistency, construct validity, and divergent validity. Student performance on the Forward Exam is reported as a continuous scaled score that can be classified as one of the four performance levels which reflect student understanding and ability to apply grade-level knowledge and skills associated with college content-readiness (Wisconsin DPI, 2018). The four classifications are Below Basic, Basic, Proficient, and Advanced.
The 2018 Forward Exam was the criterion measure. We classified student performance on the criterion measure as not proficient (i.e., Basic or Below Basic) or proficient (i.e., Proficient or Advanced) using the state-provided cut-scores. Cut-scores and associated statewide percentile ranks were 626 (56th percentile), 647 (57th percentile), and 667 (64th percentile) in Grades 6 through 8, respectively (Wisconsin DPI, 2018).
The 2017 Forward Exam was used as a screening measure. The 2017 Forward Exam was administered the previous spring when participants were in the preceding Grade. For the interval LR analyses, student performance was classified as one of four ordinal performance levels provided by the state. For the dichotomized LR analyses, we used the state provided cut-scores that differentiated between Basic and Proficient performance to classify student performance as at-risk or not at-risk (Wisconsin DPI, 2018).
MAP
The district administered the common-core aligned MAP math assessment (NWEA, 2014) in the fall, winter, and spring. Results from the fall 2017 and winter 2018 administrations were used as screening measures in this study. NWEA (2014) reported adequate reliability and provided evidence regarding content, concurrent, predictive, and construct validity. MAP scores are reported in vertically scaled Rasch Units (i.e., RIT scores) that range from 100 to 350. RIT scores are suitable for measuring achievement across grades with higher scores representing higher achievement.
NWEA (2017) published results of a linking study between MAP and the Forward Exam that indicated the MAP scores corresponding with each of the Forward Exam performance levels (see Table 2). We used the NWEA recommended cut-scores that differentiated between Basic and Proficient performance to classify students as at-risk or not at-risk to apply the dichotomized LRs.
Threshold Definitions and Dichotomous Cut-Scores for Each Screening Measure.
Note. Bold values are the dichotomous cut-scores that differentiated between below proficient and proficient performance. MAP = Measures of Academic Progress.
Grade for the 2017 Forward Exam reflects the grade the students were in when they took the test in the previous year.
Procedure
Data collection procedures in this study were identical to those in Klingbeil et al. (2019). School staff administered the MAP to all students without an alternative assessment plan following district protocol. The district estimated that MAP screening took approximately 30 to 90 min depending on the student. District staff administered the Forward Exam following Wisconsin guidelines in spring 2017 and spring 2018. There was approximately 12 months between the 2017 and 2018 administrations of the Forward Exam. In comparison, there was approximately 6 months between administration of the fall MAP and the 2018 Forward Exam and 2 months between the winter MAP and the 2018 Forward Exam. District staff transferred deidentified demographic and achievement data to the first author for analysis. The institutional review board at the University of Texas at Austin determined this study did not meet the federal definition of human subject research and did not require monitoring.
Data analysis
Missing data
There were 195 students (12.2%) missing data on at least one achievement variable. Missingness was most pronounced on the 2017 Forward Exam (10.0%), 2018 Forward Exam (2.8%), fall MAP (1.9%), and winter MAP (1.7%). Students may have been missing data due to joining or leaving the district during the school year, being absent while testing, or other unknown factors. Little’s missing completely at random (MCAR) test was significant, χ2 (39) = 136.055, p < .001, suggesting the missingness was not completely at random (Enders, 2010). Further analyses regarding relationships between the demographic variables and missingness are reported in Supplemental Materials.
We used multiple imputation to simulate the missing data in SPSS (v. 25; IBM Corp., 2017). The multiple imputation procedures were identical to those used by Klingbeil et al. (2019). Predictor variables included gender, race/ethnicity, English language learner status, special education status, and free or reduced-price lunch status. Students’ scores on the MAP (fall, winter, and spring) and Forward Exam (2017 and 2018) in math were imputed and used as predictor variables. We created 20 imputed data sets and reported the pooled estimates for all analyses in the study.
Analytic plan
We conducted all other analyses in R (R Core Team, 2020). All analyses were stratified by grade. We compared the differences between using interval LRs and dichotomized LRs using the following steps. First, we estimated the initial pretest probability and pretest odds using the district-wide percentage of students, enrolled in Grades 6 through 8 during the previous school year, who did not pass the Forward Exam during the prior school year. The rates of nonproficiency were 32.6% (odds = .483), 33.8% (odds = .509), and 46.0% (odds = .853) in Grades 6, 7, and 8, respectively.
Second, we classified student performance on the 2017 Forward Exam and MAP scores as at-risk or not at-risk using vendor recommended cut-scores. We also classified student performance on each screening measure as one of the four ordinal performance categories. We used the same cut-scores as Klingbeil et al. (2019). Third, we multiplied the interval or dichotomized LR value associated with each student’s screening test result by the pretest odds to estimate the post-test odds of failure. We used the interval and dichotomized LRs reported by Klingbeil et al. (2019) and shown in Table 1. Fourth, the post-test odds were reexpressed as post-test probabilities.
We used the post-test probabilities, estimated using dichotomized LRs or interval LRs, to emulate a gated screening framework. The 2017 Forward Exam, administered while students were in the preceding grade, was the first screening measure in each model because it was a state-mandated assessment. In Model 1, we used the fall MAP and winter MAP scores as the second and third step, respectively. In Model 2, we used the winter MAP as the second step (omitting the fall MAP data). Model 2 was consistent with previous suggestions that screening decisions in the fall could be made based on the prior year state test with additional screening occurring the middle of the year (e.g., Gersten et al., 2009). Next, we used the recommended benchmarks of VanDerHeyden (2013) to interpret the post-test probabilities after each step of the gated screening model.
Interval LRs
When interval LRs were used, students were classified as not at-risk when their post-test probability of failing the 2018 Forward Exam was ≤.10. If the post-test probability was ≥.50, the student was classified as at-risk. In either case, the student was not moved onto the next step in the screening process. Students with post-test probabilities between .11 and .49 continued to the next step of the emulated gated framework. The post-test probabilities from the previous step were used as the revised pretest probability and converted into odds. We multiplied the revised pretest odds by the interval LRs which corresponded to the student’s performance on the next measure in the gated model. After the final step in the model, we classified the number of students with final post-test probabilities between .11 and .49 as undifferentiated. For Model 1, if there were no students with post-test probabilities in the further testing range after the fall MAP, the results from the winter MAP were not used.
Dichotomized LRs
Applying dichotomized LRs requires estimation of positive and negative post-test probabilities. When students were classified as at-risk on the screening measure, we multiplied the pretest odds by the dichotomized positive LR to estimate the positive post-test probability. When the positive post-test probability was ≥.50, the student was classified as at-risk and did not move on to the next step of the gated model. Students with positive post-test probabilities ≤.49 continued to the next step of the emulated gated framework. For students who were classified as not at-risk on the screening measure, we multiplied the pretest odds by the dichotomized negative LR to estimate the negative post-test probability. Students with negative post-test probabilities ≤.10 were classified as not at-risk and did not move on to the next step of the gated model. Students with negative post-test probabilities >.10 continued to the next step of the emulated gated framework.
Diagnostic accuracy
We estimated the sensitivity and specificity of the predictions made based on the emulated gated model. Students who could not be classified as at-risk or not at-risk after the final step in the gated model were considered undifferentiated. For all other students, their at-risk or not at-risk classification was based on their post-test probabilities, regardless of the step in which that designation occurred.
Results
Descriptive statistics and correlations are shown in supplemental materials. Approximately 34.4% of sixth grade students, 36.1% of seventh grade students, and 39.1% of eighth grade students did not achieve proficiency on the 2018 Forward Exam. As stated above, the rates of nonproficiency for the previous cohort of students were 32.6% in Grade 6, 33.8% in Grade 7, and 46.0% in Grade 8. Across the two cohorts, the proportion of nonproficient students was not significantly different in Grade 6, χ2(1) = 0.36, p = .55, or Grade 7, χ2(1) = 0.65, p = .42. The proportion of Grade 8 students who did not achieve proficiency was significantly smaller than that in the previous cohort, χ2(1) = 5.31, p = .02.
Differences in Efficiency
Our first research question compared the number of students requiring further assessment between the interval and dichotomized LR approach to estimating post-test risk. The number of sixth grade students in each category and the associated pretest and post-test probabilities after each step is depicted in Figures 1–4. Similar figures for Grades 7 and 8 are provided in the supplemental materials. The number of students in each post-test probability range is shown in Table 3. Across all three grades, between 64.5% and 72.1% of students required further assessment after Step 1 when dichotomized LRs were used to estimate post-test probabilities. Between 39.6% and 46.2% of students required further assessment after Step 1 when interval LRs were used.

Sixth grade students flow through gated screening Model l when interval likelihood ratios were used to estimate post-test probabilities.

Sixth grade students flow through gated screening Model 1 when dichotomized likelihood ratios were used to estimate post-test probabilities.

Sixth grade students flow through gated screening Model 2 when interval likelihood ratios were used to estimate post-test probabilities.

Sixth grade students flow through gated screening Model 2 when dichotomized likelihood ratios were used to estimate post-test probabilities.
Number of Students in Post-Test Probability Threshold Categories.
Note. Outcomes are pooled estimates from 20 multiply imputed data sets. We rounded the pooled number of students in each category to the nearest whole number, which resulted in small differences in the number of students in some models. Prior to Step 1, the pretest probability was .326 in Grade 6, .338 in Grade 7, and .460 in Grade 8. Students with post-test probabilities ≤.10 were classified as not at-risk and students with post-test probabilities above ≥.50 were classified as at-risk. Students with post-test probabilities in the .11 to .49 range continued to the next step in the gated framework. Tests in bold were planned but not required because all students were identified as at-risk at the prior step. 17-FWD = performance classification from 2017 Forward Exam; F-MAP = Measures of Academic Progress performance classification from fall 2017; W-MAP = Measures of Academic Progress performance classification from winter 2018.
The use of dichotomized or interval LRs to estimate post-test probabilities also affected the number of students who required further assessment after Step 2 in Grades 6 and 7. When dichotomized LRs were used to estimate post-test probabilities, only two steps were needed to classify all students as at-risk or not at-risk based on their post-test probability. Using interval LRs resulted in 134 (26.2%) students in Grade 6 and 193 (37.3%) students in Grade 7 who could not be definitively classified as at-risk or not at-risk until after the results of the winter MAP were applied (Model 1). Students who could not be differentiated before the winter MAP had post-test probabilities of .14 in Grade 6 and .12 in Grade 7 after the fall MAP. In Model 2, there was a large percentage of the sixth (27.0%) and seventh grade (37.8%) students who could not be definitively classified as at-risk or not at-risk after the winter MAP. The post-test probability after the winter MAP was .13 and .11 in Grades 6 and 7, respectively. Only two steps were required to classify all Grade 8 students as at-risk or not at-risk, regardless of the type of LR used (see Table 3).
Differences in Diagnostic Accuracy
Our second research question compared differences in diagnostic accuracy when interval LRs or dichotomized LRs were used to estimate post-test probabilities. Sensitivity and specificity values exceeded .80, regardless of the type of LR or gated model used (see Table 4). For Model 1, using interval LRs to estimate post-test probability resulted in slightly higher (+.03) sensitivity values than when dichotomized LRs were used in Grades 6 and 7. Specificity values were nearly identical across the two approaches for estimating LRs. When interval LRs were used in Model 2, approximately 27% of Grade 6 students (see Figure 2) and 38% of Grade 7 students remained undifferentiated after the winter MAP (see Figure S3). We did not estimate the sensitivity and specificity for these models due to the bias that would result from excluding undifferentiated students. Using interval LRs within Model 2 had little practical value in Grades 6 and 7.
Diagnostic Accuracy of Screening Models After the Final Step.
Note. Outcomes are pooled estimates from 20 multiply imputed data sets. We rounded the pooled number of students in each category to the nearest whole number which, resulted in small differences in the number of students in some models. Tests in bold were planned but not required because all students were identified as at-risk or not at-risk based on their post-test probabilities at a prior step. Classification accuracy was calculated as TP + TN / (U + TP + TN + FP + FN). 17-FWD = performance classification from 2017 Forward Exam; F-MAP = Measures of Academic Progress performance classification from fall 2017; W-MAP = Measures of Academic Progress performance classification from winter 2018; U = students who earned post-test probability scores that necessitated further testing after all assessment had been given; TP = true positive; TN = true negative; FP = false positive; FN = false negative; Sn = sensitivity; Sp = specificity; CA = classification accuracy.
Sensitivity/specificity estimates are not presented due to the bias resulting from the number of undifferentiated students after the last step in the model (shown in the U column).
All students were classified as at-risk or not at-risk when dichotomized LRs were used, and the resulting sensitivity and specificities were acceptable in both grades. In Grade 8, decisions based on the 2017 Forward Exam and fall MAP (Model 1) resulted in higher sensitivity and slightly lower specificity than decisions made based on the 2017 Forward Exam and winter MAP (Model 2), regardless of the type of LRs used.
Discussion
In studies of academic screening, researchers have rarely applied pretest probabilities estimated using different sample of students or treated screening test performance as an ordinal variable that may preserve useful diagnostic information (Brown & Reeves, 2003). Klingbeil et al. (2019) conducted an initial evaluation of whether using interval LRs improved upon the diagnostic accuracy and efficiency of a gated screening framework but the LRs were applied to the same sample from which they were derived. The purpose of this direct replication was to evaluate the screening efficiency and accuracy when applying dichotomized and interval LRs, derived from a previous sample, to a different cohort of students.
Screening Efficiency and Diagnostic Accuracy
The added precision of using interval LRs to interpret the first screening test was evident in all three grades. As in Klingbeil et al. (2019), the difference resulted from the post-test probability for students who scored in the Advanced level falling below the threshold for withholding testing. However, relative to the dichotomized LR approach, the use of interval LRs increased the number of students requiring further assessment after the second step of the screening model in Grades 6 and 7. When applying dichotomized LRs to the MAP results, there were no students with post-test probabilities warranting further testing. Approximately 26% of Grade 6 students and 37% of Grade 7 students remained in the further assessment range after the second step when interval LRs were used.
The students in Grades 6 and 7 who remained in the further testing threshold after the second gate scored in the Proficient category on the 2017 Forward Exam and the fall MAP (Model 1) or winter MAP (Model 2). These students had resulting post-test probabilities ranging from .11 to .14, depending on the grade and model. When dichotomous LRs were used, students scoring in the Proficient and Advanced categories were grouped together into a not at-risk category. Combining the students in the Proficient and Advanced categories likely lowered the obtained negative post-test probabilities below the further testing threshold.
When post-test probabilities were estimated using interval LRs, the number of students requiring a third test (Model 1) or remaining undifferentiated (Model 2) diverged from Klingbeil et al. (2019). In the previous study, a third screening test was only required in Grade 6. Similarly, only 1.7% of students remained undifferentiated after the second test in Model 2. Because Klingbeil et al. (2019) estimated and then applied the LRs to the same cohort of students, a conservative interpretation is that three tests may be required before all students can be classified as at-risk or not at-risk when interval LRs are used. Notably, only two tests were required in Grade 8 in this study and Klingbeil et al. (2019). This difference may be related to the relatively higher pretest probability in Grade 8.
Based on the wider range between the interval LRs compared with the dichotomized LRs (see Table 1), it is reasonable to expect better diagnostic accuracy from the interval LR approach (Brown & Reeves, 2003). Yet, we found similar sensitivity and specificity values when predicting students’ end-of-year proficiency status in math, regardless of the LRs used to estimate post-test probability. The negligible differences in sensitivity and specificity corroborated the findings of Klingbeil et al. (2019).
Limitations
These results should be interpreted in the context of their limitations. The generalizability of these findings is limited to middle schools with similar demographics (e.g., predominately White, high socioeconomic status [SES]) and proficiency rates. Similarly, the results are limited to situations where the MAP and the Forward Exam are used to predict math risk. The base rates of nonproficiency in a sample, along with the choice of criterion measure, will affect diagnostic accuracy estimates of any screening tool (Meehl & Rosen, 1955; VanDerHeyden, 2013). Determining whether these findings generalize to other statewide achievement tests, screening measures, grades, or academic skills will require future conceptual replications of this study.
Another limitation is that we conducted a retrospective analysis of an emulated gated screening process. These data were collected in schools that did not use a gated screening approach to determine student risk. Thus, we could not control for any math interventions provided to students, which may have affected the diagnostic accuracy results. Prospective studies in schools that use interval LRs to estimate post-test probabilities, within a gated screening framework, would provide more compelling evidence supporting the adoption of this type of approach in practice.
Future Directions for Research
The results observed in the present study add to the body of existing research examining the value of deriving screening methods on one sample and subsequently applying those methods to a new cohort of students (e.g., Nelson et al., 2017). In that regard, we observed differences in efficiency and diagnostic accuracy across time, but those differences were small enough to suggest that schools may be able to apply screening methods derived from one cohort of students to future cohorts. Barring a large shift in proficiency rates, it seems likely that similar results would be observed across time so long as the criterion and screening measures remain the same.
It may also be worth examining the degree to which increasingly specific categories of risk lead to benefits in efficiency or diagnostic accuracy. In the present study, we used interval LRs categorized by three cut-scores that corresponded to state-created or NWEA-provided performance categories. Yet, these categories were not designed to predict risk. Youngstrom (2013) suggested that screening test results be divided into thirds or quintiles to estimate interval LRs. Future research might examine the benefits of creating empirically derived cut-scores designed to partition student performance into multiple categories which maximize screening precision. Another potential direction is to evaluate the potential benefits of interpreting extant data (e.g., prior year state test scores) using interval LRs in Step 1 and then applying a dichotomous LR to newly collected screening data in Step 2 of a gated screening model.
Future research is also needed to assess the value of interpretation guidelines for post-test probabilities. There are two primary points of interest when interpreting post-test probabilities—a test threshold and a treatment threshold (Pauker & Kassirer, 1980). Students with post-test probabilities below the test threshold do not participate in additional screening, students above the treatment threshold are provided treatment, and students with post-test probabilities between the two thresholds require additional testing. VanDerHeyden (2013) identified a test threshold based on the acceptable level of risk for withholding testing (.10) and a treatment threshold based on the unacceptable level of risk for withholding intervention (.50). These guidelines have since been adopted in other studies (e.g., Van Norman, Klingbeil et al. 2017), but the nature by which those values were originally selected is relevant because they have substantial downstream effects on efficiency and diagnostic accuracy. For example, if the test threshold was set at .15 in this study, only two tests would have been required to classify all students, regardless of the type of LRs used to estimate post-test probability. To that end, it is important to recognize that the thresholds proposed by VanDerHeyden (2013) were not empirically derived and require further empirical investigation moving forward. The appropriateness of changing the test threshold could be evaluated based on its effect on the diagnostic accuracy and the efficiency of the screening process.
Taken together, the results of this study and Klingbeil et al. (2019) suggest that the primary benefit of using interval LRs to estimate post-test probability could be to reduce the number of students requiring further assessment after the first screening test. In situations where all students are required to complete the MAP, regardless of their risk status, it may mitigate the potential time savings afforded by using interval LRs to interpret the prior year achievement test. Whether the use of interval LRs at Step 1 could afford more certainty in the decisions made for some students is unknown. Additional research on school-based applications of EBA principles for predicting academic risk, including the methods for estimating LRs and determining test and treatment thresholds, is needed.
Supplemental Material
Supplemental_Materials_RR_3_AEI_6_9_20 – Supplemental material for Using Interval Likelihood Ratios in Gated Screening: A Direct Replication Study
Supplemental material, Supplemental_Materials_RR_3_AEI_6_9_20 for Using Interval Likelihood Ratios in Gated Screening: A Direct Replication Study by David A. Klingbeil, Ethan R. Van Norman and Peter M. Nelson in Assessment for Effective Intervention
Footnotes
Author’s Note
David A. Klingbeil is now affiliated with University of Wisconsin-Madison, WI, USA.
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
Supplemental Material
Supplemental material for this article is available on the Assessment for Effective Interventions website with the online version of this article.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
