Abstract
One perceived advantage of computer-based testing is that accessibility tools can be embedded within the testing format, allowing students with disabilities to use them when necessary to remove unique barriers within testing. However, an important assumption is that students activate and use the tools when needed. Initial data from large-scale computer-based testing suggest many students with disabilities are not using them; information is needed to understand why. Both computer skills and motivation are likely necessary for students to use accessibility tools; therefore, we explored whether prior computer use, math motivation, and test motivation predicted accessibility tool use on a national math test. We further explored the relationship between accessibility tool use and test performance. Accessibility tool use was relatively infrequent. Test motivation was weakly associated with text-to-speech use. Use of eliminate choice and scratchwork tools were weakly associated with performance. When combined with related empirical work, findings suggest a potential need to improve student test motivation and corresponding use of accessibility tools to improve validity of low-stakes test scores. However, given the weak relationships identified between tool use and performance, evidence-based math interventions are anticipated to be more helpful for improving math performance than mere promotion of accessibility tool use.
With the increase of technology use in Grades K–12 education, so came the increased use of technology in testing within K–12 education (Bridgeman, 2009; Gelbart, 2018; Thurlow et al., 2010). Computer-based testing can facilitate use of innovative item types, as well as offer more immediate feedback (Bridgeman, 2009). Beyond the general benefits of computer-based tests, researchers suggest computer-based tests may be beneficial to students with disabilities. This benefit is anticipated to be, in part, due to the fact that computer-based tests can allow for greater individualization, including student-selected access to a variety of accommodations and accessibility tools (Bridgeman, 2009; Flowers et al., 2011).
A variety of different accessibility tools have become more readily available during computer-based testing over the past few decades to permit students to engage with the test as they deem helpful. These tools are often designed to mimic or expand upon supports traditionally available during paper-based testing. For example, one accessibility tool that may be particularly helpful for students with reading difficulties is the text-to-speech (TTS) tool, in which a computer reads text aloud to students (Stodden et al., 2012). TTS is intended to function somewhat similarly to a “human read aloud” option, although often has an additional feature of allowing students to directly control when and which text they activate for computer-based reading support. Although considered somewhat controversial and often limited in availability on tests specifically designed to measure reading decoding and fluency, TTS is frequently available for use on tests measuring other skills (Smarter Balanced, 2022). Computer-based accessibility tools are also often available to support visibility of and focused attention on test items (e.g., magnifying, highlighting, masking/eliminate answer choice tools; Smarter Balanced, 2022), which previously occurred via large print test materials and the ability to mark-up printed test booklets. On math tests, both calculators and digital scratchwork tools are often available to students, the latter of which allow students to draw, write notes, and engage in computation processes in a manner similar to how they might have used scratch paper to solve math problems during a paper-based test or activity (Kwak & Gweon, 2019).
As computer-based tests are increasingly used to evaluate student achievement, including that of students with disabilities, researchers have engaged in work to compare computer-based and paper-based testing formats to explore whether computer-based testing, including the unique accessibility tool features afforded, results in either comparable or improved performance for students with disabilities. Among secondary students with learning disabilities, Calhoon et al. (2000) found no differences in mathematics performance between accommodated test administrations offered via computer and those offered via paper-and-pencil testing. In their study, students accessed reading support multiple times using the computer-based test; they had continuous access to a human reader on the paper-based test and could ask for sections to be reread as appropriate by the human reader. Although the reading support itself was beneficial, the medium (e.g., human vs. computer) did not produce different results. In contrast, Flowers et al. (2011) found that across academic areas, including mathematics, students scored lower on a computer-based test than on a paper-and-pencil test when both included associated reading supports. Additional researchers have similarly suggested computer-based tests may negatively impact student performance in mathematics, as compared with paper-and-pencil tests. A meta-analysis suggested computer-based testing disadvantages students with disabilities (Pan, 2016).
Given the many aforementioned advantages of computer-based tests and increasing use of computers for education in general, it is critical to explore reasons why many of the studies mentioned above found computer-based tests to disadvantage students with disabilities. Given the potential advantages that individualized access to accessibility tools has the potential to offer, further exploration is needed to better identify and understand conditions under which computer-based testing that includes accessibility tools is beneficial versus detrimental to the performance of students with disabilities (Bridgeman, 2009; Flowers et al., 2011). One critical issue not always considered when comparing computer-based testing with paper-based testing is the extent to which students who needed accessibility tools to engage in the target skills intended to be measured actually used them. In contrast to accommodations on paper-and-pencil tests that are often automatically provided (via specific test accommodation test booklet or human assistant), use of computer-based accessibility tools and accommodations are often at the discretion of the student. If students who need accessibility tools to access the test fail to activate them on items for which they are needed, they may correspondingly fail to access the test item content and response processes necessary to display their knowledge. Their scores may systematically under-represent their achievement with respect to what is intended to be measured. A corresponding failure to use necessary accessibility tools has the potential to, at least in part, explain the lower performance of students with disabilities on computer-based tests. Relatedly, it has been argued that when flexibility and choice is built into the testing program, there is greater need to obtain quality evidence of score comparability (Camara & Davis, 2022), such should arguably be an important consideration for tests in which students have choice with regard to accessibility tool use.
Researchers have started to investigate accessibility tool use during testing, and findings point to important considerations for future research. Lee et al. (2021) examined accessibility tool use on a sixth-grade statewide computer-based assessment, including both English language arts and math tests. For several accessibility tools available to all students (e.g., highlighter, masking, line reader), the proportion of students without disabilities using the tools was higher than the proportion of students with disabilities using them. The specific proportions of students using them ranged from 12% to 52% for students without disabilities and 9% to 47% for students with disabilities (ranges reflect variation for different tools and content areas investigated). Less use was evident on the math test compared with the English language arts test. With respect to TTS, which was only made available to those identified with a need for the support, not all students who were identified with a need for TTS used it, and use dropped off considerably across the course of each test. Such findings beg the question of whether all students who need accessibility support know how to and are motivated to use the associated tools, as well as whether the tools are well-matched to their preferences and needs. If there are additional access-related skills (e.g., metacognitive skills necessary to recognize when an accessibility tool may be needed to address a unique barrier a student experiences, knowledge of how to activate the tool) and dispositions (e.g., motivation to use the tool consistently throughout the test) necessary among students to ultimately activate and use accessibility tools, this has the potential to unfairly disadvantage those who need them on a test. Because students with disabilities are a group of students specifically anticipated to need certain accessibility tools and accommodations to overcome disability-related barriers, a better understanding of their accessibility tool use and the factors that predict use will be helpful to better understand and address related concerns. An exploration of factors that predict student use of accessibility tools is anticipated to help test developers, researchers, policymakers, and Individualized Education Program teams better understand the conditions that may be necessary to promote optimal accessibility of tests, specifically when students are responsible for creating their own accessible test environments via self-directed accessibility tool activation.
Possible Predictors of Accessibility Tool Use
Knowledge and skills for using accessibility tools in a digital environment, as well as motivation to use them, are anticipated to be critical for activating and using accessibility tools. The predictors we selected to explore empirically included one corresponding to computer knowledge and skills (i.e., prior computer use) and two associated with motivation to engage in math and testing (i.e., general math motivation and motivation to perform well on the corresponding test). Each of these predictors is described more fully below.
As computers have become more widely used for testing, a general concern is that those students with less exposure and experience with computers may be at a unique disadvantage during testing (Thurlow et al., 2010), given that they will not be as familiar with specific skills involved in navigating the digital environment. Limited prior access to computers may correspondingly limit a student’s likelihood of using the associated computer-based accessibility tools. Although not studying mathematics or accessibility tool use for students with disabilities in particular, Tate et al. (2016) found students with prior computer experience related to writing had higher writing scores on the computer-based National Assessment of Educational Progress (NAEP) writing assessment. Indeed, familiarity with writing in the digital environment may have helped students more meaningfully engage in writing within the digital environment to subsequently obtain higher scores. To our knowledge, prior computer use has not yet been explored as a possible factor predicting accessibility tool use among students with disabilities.
Another possible predictor of accessibility tool use during computer-based testing is student motivation. As mentioned earlier, activation of some computer-based accessibility tools, such as TTS, often requires the student to actively choose to use the tool, rather than having it automatically provided, such as may be the default when using a human-based reader. As such, students must take an active role in the process of both activating and using them, which likely requires additional student awareness and effort among non-proficient readers to demonstrate underlying knowledge and skills on a test compared with students who are proficient readers. Student motivation to engage in academic activities can vary considerably (Cleary & Chen, 2009); those with low academic-related motivation may be particularly less likely to activate the associated accessibility tools designed to facilitate their meaningful engagement in test content. Researchers have identified a relationship between motivation in general and student performance on tests, such as NAEP (LaFave et al., 2022). Of concern with computer-based tests is a possibility that motivation (as opposed to underlying academic skills) could account for a particularly large amount of construct-irrelevant variance in test scores, particularly for students with disabilities who may uniquely need motivation to activate and use accessibility tools throughout a test. In considering the possible predictive factor of motivation, test-specific motivation is an area of investigation increasingly recognized as something related but distinct from more general academic motivation that can influence test engagement and ultimately the validity of test scores (Soland et al., 2019). A lack of motivation to perform well on a test may be particularly evident on a test that has limited consequences specific to students (i.e., is low stakes for students), and therefore it may be particularly important to explore this factor in the context of such tests. To our knowledge, neither general motivation nor test-specific motivation has been explored as possible predictors of accessibility tool use among students with disabilities, despite the notion that such a relationship might contribute to a more comprehensive understanding of test score differences among students with disabilities on computer-based tests.
Current Study
Accessibility tools are increasingly embedded in large-scale testing programs to promote valid testing among students with diverse needs, but questions remain about the extent to which these tools help or hinder student performance. Further research is needed. Although students are often provided brief tutorials on how to use these tools immediately prior to testing, it is questionable whether students make optimal use of the tools during the test to show what they know and can do, particularly if the test score bears limited consequences for students. The increasing availability of test process data can allow for related exploration of the use of accessibility tools given that it offers rich information on students’ actual use of the tools in real time. Therefore, using test process data from the 2017 eighth-grade NAEP math test, a test with low stakes for students, we explored student use of various accessibility tools, including the extent to which several factors predict accessibility tool initiation and use among students with high-incidence disabilities. Predictors examined include prior computer use, student motivation and interest in math, and student test-based motivation and effort. We specifically explored student use of three accessibility tools separately, all of which are examples of accessibility tools a student must initiate use of on the respective test: TTS, the eliminate choice tool, and the scratchwork tool. Our research questions were as follows:
Method
Process data made available to the first author from the U.S. Department of Education’s Institute for Educational Sciences (IES) were analyzed. These included response process data from the 2017 administration of one block of 15 math items administered digitally from the NAEP eighth-grade math test. More details about the test are available at https://www.nationsreportcard.gov/process_data/. Per IES guidelines, we report all sample sizes rounded to the nearest 10.
Participants
The NAEP is administered to a U.S. sample of students that is nationally representative; however, students are assigned to complete specific blocks of items such that each block is completed by only a subsample of the full national sample. Currently, one block of the 2017 eighth-grade math test administered digitally contains the process data necessary to examine accessibility tool use, and it was available to the first author via a restricted use license. It is important to note that NAEP block administration is not organized to ensure national representation. As such, block-level results cannot be interpreted as accurately representing findings for a nationally representative sample of students; relatedly, sampling weights commonly applied with full NAEP data sets are not considered appropriate with data such as this representing a single block of item responses. For analytic purposes, we selected data corresponding to those students coded as having one or more of the following high-incidence disabilities: Learning Disability (LD), Emotional Disturbance (ED), Other Health Impaired (OHI), and/or Speech/Language Impaired (SLI). Of the 28,160 eighth-grade students who participated in the item block made available for analysis, 2,520 (8.9%) were identified as having one (or more) of the four selected high-incidence disabilities; they subsequently formed the sample for the current analyses. Among these 2,520 students, 1,650 (65.3%) were male and 880 were female (34.7%); 1,510 (59.6%) were eligible for free- or reduced-price lunch, 950 (37.8%) were not eligible for free- or reduced-price lunch, and such information was not available for the remaining 70 (2.6%) of students. Over 50% (1,330; 52.9%) of students were White, 490 (19.3%) were African American, 490 (19.6%) were Hispanic, 50 (1.5%) were Asian, 60 (2.5%) were American Indian/Alaska Native, and 20 (0.7%) were Native Hawaiian or Pacific Islander. Disability frequencies within the data set were as follows: 1,420 (56.1%) LD only, 650 (25.7%) OHI only, 160 (6.3%) ED only, 160 (6.5%) SLI only, and 140 (5.5%) two or more of the four given disabilities.
Introduction to NAEP Process Data
The NAEP is administered for the purpose of providing aggregated information on the status of student achievement across the nation; there are no direct consequences for individual students corresponding to their test scores (i.e., low stakes for students). Student knowledge and skills in the following areas are measured on the NAEP 2017 eighth-grade math test: number properties and operations; measurement; geometry; data analysis, statistics, and probability; and algebra (Nation’s Report Card, n.d.). The 15-item math block selected for analysis included the following item types: multiple choice–single select, multiple choice–multiple select, matching, zone–multiple select, and grid items (National Center for Educational Statistics [NCES], 2020). All items included at least some printed words/text, except for the first item which was merely a computation item. A variety of accessibility tools were available to all students, and a brief tutorial was provided to students in advance of item block presentation. This tutorial provided an orientation to the digital testing environment and the available accessibility tools. Information on student use of accessibility tools and item responses was recorded for each item for each student and included in the available data set. In addition, after math item block administration, students completed a questionnaire that included a series of questions about their background and testing experiences, with responses to these items also made available to the first author and used for the analyses.
Target Variables
Demographic Variables, Including Disability Type
Demographic information, including disability type, was provided by school staff with knowledge of the individual students completing the test. Disability categories (i.e., LD, ED, OHI, SLI) were dummy coded to facilitate analysis. In doing so, students who were reported as having more than one of the four targeted disabilities were coded as “0” for all disability variables.
Computer Use
Seven items that students completed as part of the questionnaire were selected to comprise a measure of prior math-related computer use. More specifically, each item asked students to report the frequency with which they used computers or other digital devices for math-related tasks. Each item required students to select from a list of choices the response that aligned closest with the frequency of their associated computer use (e.g., never, once, two or three times, four or five times, more than five times per year). Each item response was then coded as a number (1–5) representing the given student’s frequency of use (higher representing more use), and responses to the seven items were summed to obtain an individual student’s total computer use score. The coefficient alpha associated with this seven-item scale was .79 for the entire sample and .70 for those in the high-incidence disability sample used for the analyses. Sample scores ranged from 7 to 35.
Math Motivation
Seven items that students completed as part of the questionnaire were selected to comprise a measure of math motivation. More specifically, each item included a statement about a desire to learn or understand math and required students to indicate how much the given statement described a person like them (e.g., “not at all like me,” “a little bit like me,” “somewhat like me,” “quite a bit like me,” “exactly like me”). Each item response was then coded as a number (1–5) representing their desire to learn and understand math (higher indicating more desire), and responses to the seven items were summed to obtain a total math motivation score. The coefficient alpha associated with this seven-item scale was .94 for the entire sample and .91 for those in the high-incidence disability sample. Sample scores ranged from 7 to 35.
Test Motivation
Two items students completed as part of the questionnaire were selected to comprise a measure of test motivation: “How important was it to you to do well on this test?” [response options = not very important (1), somewhat important (2), important (3), or very important (4)] and “How much effort did you apply to succeed on this test?” [response options = no effort at all (1), very little effort (2), some effort (3), quite a bit of effort (4), and a lot of effort (5)]. Item responses were summed to obtain a total test motivation score. The coefficient alpha associated with this two-item scale was .72 for the entire sample and .72 for those in the high-incidence disability sample used for the analyses. Sample scores ranged from 2 to 9.
Use of Text-to-Speech (TTS)
All students were able to select and deselect the TTS function continuously throughout the item block. Information was available within the process features data set at the item level on how many times students activated the TTS tool. We first recoded this information to determine for each student whether or not they used TTS for a given item. We then created a raw score total representing the number of items on which each student used TTS. For analysis, each student was then coded as (a) not using TTS for any items, (b) using TTS for one to seven items (i.e., fewer than half of the items in the block), or (c) using TTS for eight or more items. We anticipated some students would not need to use the tool given strong reading skills, others may need to use the tool on some items given slightly below proficiency with reading, and others may need to use the tool for most to all items given reading skills far below proficiency (see Lee et al., 2021 for more information on use of TTS).
Use of Eliminate Choice
All students were able to select and deselect item responses using the “eliminate choice” function continuously while completing the multiple-choice items, of which there were three in the given item block. Information was available within the process data set at the item level on how many times students used the eliminate choice tool. We first recoded this information to determine for each student whether they used elimination for a given multiple choice item. We then created a raw score total representing the number of items on which each student used this tool. Students were coded as (a) not using eliminate choice for any items (i.e., 0), (b) using eliminate choice for one item (i.e., 1), or (c) using eliminate choice two or more items (i.e., 2).
Use of Scratchwork
All students were able to use a “scratchwork” option at any point during the test. Information was available within the process features data set at the item level about whether or not students used the scratchwork option. We created a raw score total based on the number of items each student used this tool. Many items could easily be completed without scratchwork; indeed, for few items did many students use this tool. Therefore, for analysis, students were coded as (a) not using scratchwork, (b) using scratchwork for one item, or (c) using scratchwork for more than one item.
Math Test Performance
A total raw score of 25 was possible for the given item block. Seven items were dichotomously scored, seven items had a maximum score of 2, and one item had a maximum score of 4. The total block score (with students weighted equally) had good reliability in the full sample, α = .80, and corresponded well to the weighted reliabilities reported for corresponding 15-item blocks by NCES. The reliability for the sample of students with high-incidence disabilities that was analyzed was α = .70.
Data Analyses
Complete data were available for nearly all of the predictor and response variables, with the exception of the measures derived from the questionnaire data (i.e., computer use [16% missing], math motivation [20% missing], and test motivation [10% missing]). Little’s Missing Completely at Random test was significant, χ2(27) = 45.608, p = .014, suggesting the data did not meet the criteria for being Missing Completely at Random (Little, 1988). As a result, multiple imputation methods were applied for analyses involving these measures (Baraldi & Enders, 2010). Ten imputed data sets were created using the multiple imputation function in SPSS; all other variables included in the study were used as part of the imputation process (i.e., gender, race/ethnicity, free- or reduced-price lunch status, disability type, total points, computer use, math motivation, test motivation, use of the three accessibility tools). When variables with missing data were included in the given analysis, results are presented either as pooled statistics or via reporting of the median and range for the 10 imputed samples (Manly & Wells, 2015).
Three separate ordinal logistic regression analyses were completed that included each of the predictor variables (i.e., dummy coded disability categories, computer use, math motivation, and test motivation) predicting use of the three accessibility tools (i.e., TTS, choice elimination, and scratchwork) separately. Similarly, three one-way analyses of variance (ANOVAs) (one for each accessibility tool) were completed to examine the relationship between accessibility tool use and total raw math test score. The Holm-Sidak method was used to account for multiple tests and associated adjusted p values calculated and reported for the corresponding 31 significance tests completed.
Results
Descriptive Analyses
Of the 2,520 students studied, 1,610 (63.6%) never used TTS, 650 (25.8%) used TTS on one to seven items, and 270 (10.5%) used TTS on eight or more items. Almost 70% (1,750; 69.2%) never used the eliminate choice tool, 520 (20.8%) used the eliminate choice tool on one item, and 250 (10.0%) used the eliminate choice tool on two or three items. Less than half of students in the sample, more specifically 1,180 (46.7%), never used the scratchwork tool, 550 (21.9%) used the scratchwork tool for one item, and 790 (31.3%) used the scratchwork tool on two or more items.
Among the 10 imputed samples, the median of the mean values for computer use was 22.15 (range = 22.08–22.25), which represents an average item-level score of approx. “3” (i.e., mid-level frequencies of computer use) on the corresponding 1 to 5 scale. The median of the standard deviations for computer use was 6.02 (range = 5.95–6.04). The median of the mean values for math motivation was 24.94 (range = 24.74–25.08), which represents an average item-level score of approx. “3” (i.e., “somewhat like me”) on the corresponding 1 to 5 math motivation scale. The median of the standard deviations for math motivation was 7.24 (range = 7.16–7.38). The median of the mean values for test motivation was 6.62 (range = 6.59–6.66); given the different corresponding values for the two items used, it is difficult to offer context for interpreting this median score apart from that it represents “middle scores” on the item scales (i.e., somewhat important to important, and some effort to quite a bit of effort). The median of the standard deviations for test motivation was 1.57 (range = 1.57–1.59). The average total raw score was 5.63 (out of a total possible score of 25), with a standard deviation of 3.45.
Predictors of Accessibility Tool Use
Text-to-Speech (TTS) Use
Ordinal logistic regression was used to examine the extent to which disability type, computer use, math motivation, and test motivation predicted TTS use during the item block. No collinearity problems were evident given tolerance values greater than 0.1. Moreover, the assumption of proportional odds was met, as assessed by a full likelihood ratio test comparing the fit of the proportional odds model to a model with varying location parameters: χ2(7) = 10.32 (imputed samples range = 10.03 to 11.45), adj. p = .91. The model statistically significantly predicted the dependent variable over and above the intercept-only model, χ2(7) = 53.05 (imputed samples range = 49.53 to 57.00), adj. p < .001. Table 1 provides the model results, including the medians and ranges for non-pooled statistics. An increase in test motivation (expressed in rating scale points) was associated with an increase in the odds of using text to speech, with an odds ratio of 1.12 (95% confidence interval: [1.06, 1.18]), Wald χ2(1) = 16.03, adj. p < .01. No other predictor variables in the model significantly predicted TTS use.
Ordinal Logistic Regression Model for Disability Type, Computer Use, Math Motivation, and Test Motivation Predicting Text-to-Speech Use.
Source. U.S. Department of Education, National Center for Educational Statistics (NCES), 2017 NAEP eighth-grade Mathematics Process Data, Student Features File Partial Form (restricted use data set made available October, 2020).
Note. LD = learning disability; OHI = other health impairment; ED = emotional disturbance; SLI = speech or language impairment.
Pooled results. bMedian (range) for 10 imputed groups.
Eliminate Choice Use
Ordinal logistic regression was used to examine the extent to which disability type, computer use, math motivation, and test motivation predicted use of the eliminate choice tool during the item block. No collinearity problems were evident given tolerance values greater than 0.1. Moreover, the assumption of proportional odds was met, as assessed by a full likelihood ratio test comparing the fit of the proportional odds model to a model with varying location parameters: χ2(7) = 5.98 (imputed samples range = 4.79–7.35), adj. p = .98. The model did not predict the dependent variable over and above the intercept-only model, χ2(7) = 13.89 (imputed samples range = 11.11–16.45), adj. p = .61.
Scratchwork Use
Ordinal logistic regression was used to examine the extent to which disability type, computer use, math motivation, and test motivation predicted scratchwork use during the item block. No collinearity problems were evident given tolerance values greater than 0.1. Moreover, the assumption of proportional odds was met, as assessed by a full likelihood ratio test comparing the fit of the proportional odds model to a model with varying location parameters: χ2(7) = 6.39 (imputed samples range = 5.20–9.01), adj. p = .99. The model statistically significantly predicted the dependent variable over and above the intercept-only model, χ2(7) = 27.45 (imputed samples range = 23.73–31.24), adj. p =.03. Table 2 provides the model results, including the medians and ranges for non-pooled statistics. No specific predictors were identified as significant following use of the Holm-Sidak adjustment.
Ordinal Logistic Regression Model for Disability Type, Computer Use, Math Motivation, and Test Motivation Predicting Scratchwork Use.
Source. U.S. Department of Education, National Center for Educational Statistics (NCES), 2017 NAEP eighth-grade Mathematics Process Data, Student Features File Partial Form (restricted use data set made available October, 2020).
Note. LD = learning disability; OHI = other health impairment; ED = emotional disturbance; SLI = speech or language impairment.
Pooled results. bMedian (range) for 10 imputed groups.
Relationship Between Accessibility Tool Use and Math Test Performance
Text-to-Speech (TTS) Use
A one-way ANOVA was conducted to determine whether math test performance was different for the groups representing three different levels of TTS use. There were several outliers, as assessed by boxplot; however, one-way ANOVA can be considered robust to related violations (Maxwell & Delany, 2004). Data were presented as mean ± standard deviation. The total points score increased from high TTS use (n = 270, 5.15 ± 2.67), to medium TTS use (n = 650, 5.56 ± 3.29), to no TTS use (n = 1610, 5.74 ± 3.61), in that order. However, these group differences were not significant, Welch’s F(2, 753.70) = 5.00, adj. p = .14.
Eliminate Choice Use
A one-way ANOVA was conducted to determine whether math test performance was different for the groups representing three different levels of eliminate choice tool use. Data were presented as mean ± standard deviation. The total points score increased from no eliminate choice use (n = 1750, 5.40 ± 3.22), to medium eliminate choice use (n = 520, 6.08 ± 3.76), to high eliminate choice use (n = 250, 6.27 ± 4.06), in that order. The total points score was statistically significantly different for different levels of eliminate choice use, Welch’s F(2, 554.59) = 10.060, adj. p < .001, partial eta-squared = .003. The assumption of homogeneity of variance was violated, as assessed by Levene’s test of homogeneity of variances (adj. p < .001), and so the Games-Howell test was used to test for specific mean score differences. The increase in math performance between the no use to medium use of .68 (95% CI, [.25, 1.10]) was statistically significant (adj. p < .01). No other significant differences between groups were identified.
Scratchwork Use
A one-way ANOVA was conducted to determine whether math test performance was different for the groups representing three different levels of scratchwork tool use. Data were presented as mean ± standard deviation. The total points score increased from no scratchwork use (n = 1,180, 5.44 + or – 3.41), to medium scratchwork use (n = 550, 5.51 ± 3.53), to high scratchwork use (n = 790, 6.01 ± 3.41), in that order. The total points score was statistically significantly different for different levels of scratchwork use, Welch’s F(2, 2520) = 6.89, adj. p = .02, partial eta-squared = .005. There was homogeneity of variances, as assessed by Levene’s test for equality of variances (adj. p = .52), and so Tukey post hoc tests were applied. The increase in math performance between the no use to high use of .57 (95% CI [.20, .94]) was statistically significant (adj. p =.02). No other significant group differences were identified.
Discussion
An innovative feature of computer-based tests is the opportunity for students to use accessibility tools, which may result in more valid test scores for a variety of students. However, use of some tools has been shown to be particularly low for students with disabilities who may particularly need them (Lee et al., 2021). This begs questions about why so few students make use of them. We correspondingly explored the extent to which three factors (i.e., computer use, math motivation, and test motivation) predicted use of three different accessibility tools among students with high-incidence disabilities. In addition, we examined the relationship between accessibility tool use and math test performance. Overall prediction models, including disability types and the three corresponding factors, significantly predicted TTS and scratchwork use. Test motivation was identified as a significant individual predictor of TTS use, although the corresponding effect size was weak. Use of the eliminate choice option and scratchwork were both positively associated with performance, but again with only weak effect sizes.
One important finding to initially note is that the majority of students with high-incidence disabilities never made use of TTS or eliminate choice. Although it is possible these students did not have a need for and would not have benefited from using the tools, it is important to recognize that overall use seems to be relatively infrequent. Of the students with LD, which represented the most common disability type of those included in the analysis and for whom it is anticipated included many students with difficulties with reading (Moll et al., 2014), approximately 60% never used TTS, despite all but one of the test items requiring reading skills to access. Although additional information is needed to know whether students truly needed to use TTS given that we do not have information on their actual reading skill level, it is a bit concerning that so few students made use of this feature given the likelihood that many may struggle with reading skills and may need to use it to truly access the test content. In contrast to TTS, the eliminate choice accessibility tool may be viewed as representing a test-taking strategy that is not necessarily required to understand and answer items. However, limited use of choice elimination among the students in our study aligns with other findings in the literature suggesting many students with disabilities need to be taught and encouraged to use strategies for test-taking (Ray, 2018); merely being offered access to associated accessibility tools may not be enough to ensure adequate and appropriate use of them.
Models including disability type, computer use, and the two types of motivation were predictive of use of both TTS and scratchwork. However, among the individual predictor variables, only test motivation (i.e., perceived importance of successful test performance) was found to be significant and was merely a significant predictor for TTS use. Moreover, the odds ratio value of 1.12 suggests a very weak effect when considering an odds ratio value of 1.50 as the threshold for a small effect (Chen et al., 2010). Although additional research may be helpful to identify the extent to which efforts to increase students’ perceptions of the importance of successful test performance increase their use of needed accessibility tools, it may be particularly worthwhile to explore other potential reasons for low accessibility tool use. For example, qualitative methods to identify students’ reasons for not using TTS may be worthwhile to better understand and correspondingly address lack of use.
Contrary to our expectation, prior computer use was not associated with greater use of accessibility tools. This finding is promising with respect to ensuring equity in appropriate access to test items among students with differing prior access to computers. Had a significant correlation been identified, this would have pointed to the possibility that students with greater prior access to computers have an increased opportunity to show what they know and can do on the test. Instead, results seem to suggest students with and without extensive prior computer use may be on more-or-less equal footing with use of accessibility features on this computer-based test. The advanced tutorial provided, which includes student practice with the accessibility tools, may have been sufficient for students, including those with little computer experience, to effectively use the tools.
Use of eliminate choice and scratchwork were both positively associated with math test performance; however, the associated effect sizes were quite weak, with associated partial eta-squared values of .003 and .005, respectively. Such values are below the .01 threshold deemed to indicate evidence of a small effect size by Cohen (1988). It is important to note that other efforts (in place of increased use of accessibility tools) may be much more likely to lead to large effects on math achievement test scores. For example, strategic use of several digital tools for teaching math skills (e.g., dynamic math tools, intelligent tutoring systems, drill and practice activities) have been found to have medium and large effect sizes on math and science test scores of secondary students (Hillmayr et al., 2020).
Although our results identified only weak relationships among predictor and outcome variables, the general pattern of the results does point to some potential value in ensuring students have a desire to succeed on the test and that they make use of scratchwork and the eliminate choice tools to engage and potentially perform optimally on the test—these may be areas for future exploration. Other research has highlighted lack of student test engagement as something that can result in depressed and invalid test scores (Wise et al., 2021). Test disengagement has the potential to particularly impact the testing of students with disabilities on computer-based tests, many of whom may need to engage in additional activities (e.g., activation of TTS) in order to access and engage in the test in a way that ensures the test scores are valid. This may be particularly the case on tests for which individual students experience no consequences, such as the test that was the focus of this investigation. If students correspondingly do not view test performance as important and would rather be engaging in other activities with their time, they may avoid accessibility tool activation and correspondingly rush through the test to be able to engage in more valued activities (e.g., Los et al., 2022).
Limitations
This work was correlational in nature; although it can help in identifying potentially helpful areas for future investigation, the results alone do not have implications for what affects accessibility tool use and test performance. It is also important to highlight the measures used to indicate computer use, and the two different types of motivation were not well-established measures and instead represented combinations of related items used for the purpose of our exploratory analyses. Although reliabilities derived from the sample for these tools were adequate, research using more established measures is warranted. Finally, it is important to emphasize that although cases were selected from a nationally representative rather than a convenience sample, they do not necessarily represent students with high-incidence disabilities nationally.
Implications for Future Research and Practice
Similar to other work (e.g., Lee et al., 2021), our findings of very limited accessibility tool use among students with disabilities beg the question of whether mere embedding of accessibility tools on a computer-based test ensures test accessibility, specifically on a test with low states for students. Additional efforts may be needed to ensure accessibility tool use among those who need them. Relatedly, test motivation appears to be something that may be helpful for researchers to explore when considering how to potentially increase use of accessibility tools among students with high-incidence disabilities during testing with low stakes for students. In the current study, students’ test motivation varied considerably; practical methods for identifying students who lack test motivation may be an important first step, followed by the investigation of various approaches to fostering their motivation. In recent years, researchers have begun to investigate issues related to test disengagement (e.g., Wise et al., 2021), which may be a particularly important area to address in light of an increasing use of low-stakes test scores to measure achievement. A meta-analysis of related literature points to the helpfulness of increasing test relevance and providing external incentives to improve test engagement (Rios, 2021).
Determining the contexts in which use of various accessibility tools should be encouraged in order to promote more valid test scores should also be a focus of future research. Certainly, not all students will benefit from use of every accessibility tool; for example, not all students need to use TTS (Silvestri et al., 2021), and it may in fact represent a distraction and hindrance for some. However, given that both use of eliminate choice and scratchwork were associated with higher test performance, it may be worthwhile to more closely experimentally examine whether use of these tools indeed facilitates stronger performance on tests. A major advantage of process data such as those explored in the current study is that researchers can more carefully examine fidelity in accessibility tool use, which is something that has been neglected in earlier work on test accessibility tools and accommodations. As a result, much of the existing research on test accommodations must be considered critically, given that lack of identified effectiveness may have resulted from a student’s in-the-moment choice not to use the accommodation or accessibility tool, rather than failure of the accommodation or accessibility tool itself to foster greater access. More research is needed to help understand how to promote use of accessibility tools that are indeed beneficial. Moreover, limited use on tests begs questions of whether students are using accessibility when needed during instruction; if not, poor performance on tests may be a result of failure to use needed accessibility tools during both instruction and testing.
An implication for practice that stems both from our work and the scholarship of others (e.g., Ray, 2018) is the importance of equipping students with high-incidence disabilities with metacognitive strategies for engaging effectively in test items and similar instructional tasks. More specifically, it may be particularly important to teach and encourage students to actively think about their own accessibility tool needs during a given task, and to correspondingly intentionally decide whether or not to use various accessibility tools that are expected to support them in accurate completion of the required tasks. Based on our findings, many students with disabilities currently appear to ignore these supports during testing, when information suggests they may help students perform better. Although it is important to recognize that individual student needs and preferences may vary, and that students should ultimately be encouraged to use those that are helpful, equipping students with knowledge of accessibility tools that can help them perform better, and motivating them to want to perform better may be important strategies to use. This may help ensure stronger student engagement in the targeted academic tasks.
At the same time, it is important to recognize that efforts to promote accessibility tool use have yet to attain the same level of evidence for their overall influence on student learning and achievement when compared with other intervention-based methods. Although findings from this study point to some options to consider for potentially improving accessibility during computer-based testing and corresponding test score validity among students with disabilities in the future, existing evidence-based interventions have evidence of larger effects on student learning and achievement and should be the focus of educational efforts among those students who may, with intensive instruction, be able to develop the associated skills.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
