Abstract
For large-scale assessments, data are often collected with missing responses. Despite the wide use of item response theory (IRT) in many testing programs, however, the existing literature offers little insight into the effectiveness of various approaches to handling missing responses in the context of scale linking. Scale linking is commonly used in large-scale assessments to maintain scale comparability over multiple forms of a test. Under a common-item nonequivalent group design (CINEG), missing data that occur to common items potentially influence the linking coefficients and, consequently, may affect scale comparability, test validity, and reliability. The objective of this study was to evaluate the effect of six missing data handling approaches, including listwise deletion (LWD), treating missing data as incorrect responses (IN), corrected item mean imputation (CM), imputing with a response function (RF), multiple imputation (MI), and full information likelihood information (FIML), on IRT scale linking accuracy when missing data occur to common items. Under a set of simulation conditions, the relative performance of the six missing data treatment methods under two missing mechanisms was explored. Results showed that RF, MI, and FIML produced less errors for conducting scale linking whereas LWD was associated with the most errors regardless of various testing conditions.
In large-scale assessments, multiple forms of a test are often administered to multiple groups of test takers whose ability distributions are not equivalent. As a result, parameters that are calibrated separately for each group will be on different scales due to the indeterminacy property (S. Kim & Kolen, 2007; Kolen & Brennan, 2014) of the item response theory (IRT). This property makes it difficult to directly compare person and item parameters across different groups. More importantly, the differences between IRT scales need to be adjusted to conduct various psychometric work, such as item analysis, differential item functioning analysis, form construction, and so on. To deal with this issue, scale linking is usually used to position estimates from different groups to be on a common scale (Kolen & Brennan, 2014; Lee & Lee, 2018). Specifically, linking is indispensable when tests are administered under the common-item nonequivalent groups design (CINEG), where groups differ in ability and, thus, forms are built to share a set of common items in an effort to achieve scale comparability.
Research on IRT scale linking has focused mainly on understanding the efficiency of using various linking approaches (e.g., moment methods and characteristic curve methods; Baker & Al-Karni, 1991; S. Kim & Kolen, 2006; S. Kim & Lee, 2006; von Davier & von Davier, 2007), examining the performance of different calibration methods (Hanson & Béguin, 2002; S. Kim & Kolen, 2006; Lee & Ban, 2009), or extending scale linking to tests with complex structures or with multiple constructs (S. H. Kim & Cohen, 1998; S. Kim & Kolen, 2006; S. Kim & Lee, 2006; Li et al., 2004; Li & Lissitz, 2000; Oshima et al., 2000). These studies often assumed that responses collected from examinees are complete before dealing with linking issues. Unfortunately, achieving a complete dataset is almost infeasible in reality. For instance, in Programme for International Student Assessment (PISA) 2012, missing data were found to range from 0.3% to 13.3% at the item level for mathematics and from 0.8% to 10.5% for reading items (OECD, 2012).
Under the CINEG design, linking coefficients are often estimated separately based on the responses to common items. If missing data that occur with the common items are not handled properly, estimated item and ability parameters are possible to be affected, influencing the accuracy of scale linking estimates. In that case, the estimated abilities of examinees will also be biased after being transformed to the common scale due to the inaccurate linking coefficients. Thus, it is crucial to understand the appropriate method to deal with missing data in the scale-linking context such that sound decisions can be made to ensure test score validity and improve test fairness.
Several missing data handling approaches have been explored within the IRT context (Cetin-Berber et al., 2019; Finch, 2008; Pohl et al., 2014; Shin, 2016), such as their impacts on the estimation accuracy of item and ability parameters. However, the performance of the methods on scale linking is still unknown. This study was conducted to fill this gap in the literature and to provide implications for practitioners and researchers on how to deal with missing responses for correctly placing group abilities on a common scale. Specifically, missing data treatment approaches were examined under two missing data mechanisms for a set of simulation conditions, including examinees’ ability distributions, ratio of common items, missing rates, percentages of common items involving missing data, and test lengths. Six missing data handling approaches investigated in this study included (a) listwise deletion (LWD), (b) treating the missing data as incorrect responses (IN), (c) substituting the missing data with corrected mean (CM; Bernaards & Sijtsma, 2000), (d) applying imputation using response function (RF; Sijtsma & van der Ark, 2003), (e) using multiple imputation approach (MI; Rubin, 1987), and (f) utilizing full information maximum likelihood method (FIML; Enders, 2001a, 2001b; Finkbeiner, 1979).
Scale-Linking Approaches
This study focuses on the impact of missing data treatment methods on two IRT characteristic curve methods, including Haebara (1980) and Stocking-Lord methods (1983), because they have been found to yield more accurate results compared with the moment methods (Hanson & Béguin, 2002; S. Kim & Kolen, 2006; S. Kim & Lee, 2006; LeBeau, 2017). Both the Haebara and the Stocking-Lord methods search for optimal scale transformation constants (slopes and intercepts) to mitigate the difference between characteristic curves across common items. However, they differ in that the Haebara method finds the coefficients by minimizing the differences between item characteristic curves, whereas the Stocking-Lord approach achieves the goal by reducing the differences between test characteristic curves (see Kolen & Brennan, 2014 for more details).
Missing Mechanisms
According to Rubin’s theorems (1976), missing data are grouped into three mechanisms, including missing completely at random (MCAR), missing at random (MAR), and missing not at random (MNAR), depending on how the probability of missingness is related to the missing data. Data that are MCAR imply that the cause of the missingness is completely random. Under the MAR mechanism, the probability of having a data point missing is not dependent on the missing point itself. Instead, it is linked to some additional measured variables (e.g., total scores over non-missing items can be considered a measured variable). For MNAR data, the likelihood of a missing response cannot be explained by any measurable variables but is caused by unmeasured variable(s) (e.g., missingness depends on the individual’s ability). Missing data handling approaches have shown to perform differently based on the type of missing mechanisms (e.g., Cetin-Berber et al., 2019; Finch, 2010; Robitzsch & Rupp, 2009). Mislevy and Wu (1996) suggested considering different missing data mechanisms when estimating parameters for obtaining more accurate results. Sachse et al. (2019) analyzed the PISA data from 35 countries and identified that missing mechanisms and the presence of missing data were considerably different across multiple time points, countries, and domains. Through a simulation study, these two factors were proved to have significant impacts on trend estimates for large-scale assessments.
Large testing organizations tend to deal with missing responses based on missing types, such as non-administered items, omitted items, and non-reached items. For instance, the National Assessment of Educational Progress (NAEP) assessment considers missing responses that appear before the last observed response as omissions and treats them as fractionally correct, whereas the missing responses at the end of a block of items are considered as non-reached items and treated as not present (National Center for Educational Statistics [NCES], 2008). In real-world practice, it is very unlikely that common items are not administered or not present to examinees unless there is a technical or administrative issue, which is rare to occur. Although missing responses on common items could possibly fall under all three missing data mechanisms, it is more reasonable to assume that responses are omitted or non-reached for common items. Previous research found that both omitted items and non-reached items can be affected by the examinee’s proficiency levels (Köhler et al., 2015a; Rose et al., 2010; Sachse et al., 2019). However, they each can be associated with distinctive factors. The presence of omitted items could be attributed to item format (Köhler et al., 2015a) and item difficulty (Rose et al., 2010), whereas nonreached items are possibly influenced by one’s motivation level and test-taking strategy (Mislevy & Wu, 1988). These findings indicate that missingness for omitted items and non-reached items are likely to be MAR and MNAR rather than MCAR. The current study, as the first to examine missingness in the scale linking context, intended to identify the most appropriate methods to deal with omitted items under both MAR and MNAR missing mechanisms.
Missing Data Handling Methods
Listwise Deletion and Treating Missing Data as Incorrect
In this study, six missing data handling approaches were investigated either because they have been found to yield accurate estimation of parameters or they are easy to implement and have been commonly used (Cheema, 2014; Finch, 2008; Hawthorne et al., 2005). In terms of practicality, two of the most straightforward and easiest methods to apply are probably LWD and IN. However, treating missing responses with LWD by excluding the entire record of an examinee if any single value is missing was found to be associated with biased estimates (Enders, 2001b; Robitzsch & Rupp, 2009; Sinharay et al., 2001), except when it is under the MCAR mechanism (van Ginkel et al., 2010, 2020). The use of LWD inherently reduces overall sample size and, in turn, lowers statistical power. Considering field testing where a small number of examinees are involved, applying LWD with missing data seems to be more problematic because little data may remain if examinees who miss answering any items will be removed.
IN assumes that the missingness is MNAR. Specifically, missingness is believed to be resulted from test takers’ limited knowledge or skills to perform the tasks, and thus missing responses are treated as incorrect. However, in practice, missingness can also be related to examinees’ gender, motivation, self-efficacy, and enjoyment of working on the subject (Di Chiacchio et al., 2016), which are not MNAR. Research has shown that IN is associated with biased parameter estimates or theta estimates (De Ayala et al., 2001; Mislevy, 2017; Robitzsch & Rupp, 2009; Rose et al., 2010), but this approach is still commonly used in practice for handling missing data due to its relatively simple application.
In addition to the traditional methods introduced earlier, various imputation methods that are more complex and computationally demanding have been proposed to treat missingness in educational assessments. Imputation is a process where, “the missing values are filled in, and the resultant completed data are analyzed by standard methods” (Little & Rubin, 2020, p. 24). Both single imputation methods, including CM and RF, and multiple imputation (MI) methods are described subsequently. In addition, FIML is also discussed and examined in this study.
Corrected Item Mean Substitution
The CM method incorporates a weight to represent the examinee’s performance relative to the average performance of all examinees on the non-missing items (Bernaards & Sijtsma, 2000). For instance, a higher value is imputed to the missing item response when the examinee’s performance on non-missing items is above the average performance. The imputed value
where
Response Function
The RF approach is a method using nonparametric regression to impute values for missing data based on a latent trait parameter of an examinee (Sijtsma & van der Ark, 2003). This method assumes that an examinee’s score is related to a latent parameter denoted by the rest score,
where
Both CM and RF are single imputation strategies, which generate a complete dataset by filling in a value for each missing point. Finch (2008) found that the amount of bias in difficulty estimates associated with CM was comparable with other commonly used approaches, although a larger bias was found in item discrimination estimates with this method. The use of RF led to less bias than IN and CM but was associated with higher standard errors than CM in many conditions. These two methods were included in this study because they were designed to be applied to categorical data (Finch, 2008). In addition, both methods require relatively less intensive computation. However, using one model to restore missing data does not reflect sampling variability, which might underestimate the standard error in the subsequent statistical analysis (Little & Rubin, 2020).
Multiple Imputation Algorithm
To correct the major flaw of single imputation approaches as discussed earlier, Rubin (1987) introduced the MI algorithm. In MI, each missing data point is imputed M times. The imputed values are estimated based on the means and variances of the observed data. A standard statistical analysis is then carried out on each imputed dataset. The M sets of data are merged into one dataset (i.e., parameter estimates) by averaging values over M sets of results. For more details about MI, see Little and Rubin (2020) and Finch (2008).
It was found that MI led to accurate results in parameter estimation (Finch, 2008; Kalkan et al., 2018; Mislevy, 2017; Schafer & Graham, 2002; Sijtsma & van der Ark, 2003). The number of imputations for implementing MI needs to be determined by investigators and different numbers of imputations potentially lead to different results (Graham et al., 2007). Some researchers argue that a large number of imputations are needed to ensure the accuracy of parameter estimation (Enders, 2010; White et al., 2011). However, running too many imputations is often time-consuming and not realistic from a practical aspect. Parameter estimation involving MI is a computationally demanding process, but the introduction of several computer software, such as AMOS (Arbuckle, 2014), has made MI more accessible.
Full Information Maximum Likelihood
Schafer and Graham (2002) pointed out that MI and maximum likelihood are the two state-of-the-art methods to treat missing data in that both work well with MAR data. FIML uses all available information to estimate parameters without imputing incomplete data (Graham, 2009; Schafer and Graham, 2002). This algorithm is also one of the direct maximum likelihood (ML) approaches, as the linear parameter estimates are generated from the raw data, and no other procedure is required (Enders, 2001a). Enders (2001a) offered a more comprehensive description of FIML. The parameters to be estimated in FIML are obtained by maximizing the log-likelihood function as follows:
where
Treating missing responses with FIML was found to have more accurate results than traditional missing treatment approaches, such as IN, LWD, pairwise deletion, and mean imputation and produce similar or more accurate results to MI (Cetin-Berber et al., 2019; Enders, 2001b; Xiao & Bulut, 2020). Another advantage of FIML is that it does not require additional steps to impute missing values. FIML has become a part of built-in functions in many computer programs, such as AMOS and flexMIRT version 3.62 (Cai, 2020). Hence, practitioners or researchers may be able to easily incorporate this method into their practice.
Proper techniques for handling missing responses have been demonstrated in the literature in the context of IRT applications, mostly on the effect on the accuracy of item and ability parameter estimation (e.g., De Ayala et al., 2001; Debeer et al., 2017; Finch, 2008; Rose et al., 2015). Other scholars focused on dealing with missing responses in the context of competence tests (Köhler et al., 2015b, 2017), computerized adaptive testing (Cetin-Berber et al., 2019), and differential item functioning (Finch, 2011; Robitzsch & Rupp, 2009); however, thus far, researchers have paid relatively little attention to applications for test scale comparability. Shin (2016) investigated the effects of eight missing data approaches on vertical scaling using real data. IN was found to lead to higher discrimination and difficulty parameters, but it was associated with lower pseudo-guessing parameters. Two of the multiple imputation approaches, treating missing responses as not present (NP) and combining IN and NP (INNP), yielded similar results. In sum, different missing data treatment approaches resulted in different parameter estimates as well as vertical scaling results. However, the findings from this study have limited interpretation due to a lack of a proper evaluation criterion for real data analysis.
Research Questions
Although many studies compare missing data methods in the IRT context, the effect of these methods on scale linking is still unclear. From a practical perspective, how to handle missing responses when transforming scales into a common scale for large-scale assessments has received little attention. The primary purpose of this simulation study was to understand the relative performance of six methods to treat missing data on scale linking, with an intention to help practitioners choose an appropriate treatment.
Specifically, missing responses within two missing mechanisms, including MAR and MNAR, were examined. Also, how to handle omitted responses to common items was the focus of the current study. The following research questions guided the study:
Method
Simulation Conditions
The six missing data handling methods were compared across two test lengths, including 30-item and 60-item test forms, as used in previous research (Cetin-Berber et al., 2019). As transformation coefficients are calculated based on responses to common items over examinees under the CINEG design, this study took common items-related factors into consideration. The ratio of common items was varied with two levels, including 20% and 40%. Kolen and Brennan (2014) mentioned that the rule of thumb is to have a minimum of 20% of common items for tests involving more than 40 items. In addition, the percentage of common items that has missing responses varied at two levels: 20% and 40%. Taken together, for instance, if the percentage of common items with missing data is fixed at 20%, then missing responses occurred on five common items for a 60-item test form with 24 (40%) common items. Furthermore, three levels of missing rates were examined in the study, including 8%, 15%, and 30%, the rates frequently observed in practice according to previous research (Cetin-Berber et al., 2019; Finch, 2008).
In addition, the ability distribution of examinees taking the new test form was examined at three levels, including N(0, 1), N(0.25, 1.12), and N(0.5, 1.22). The three different proficiency levels were selected after consultation with prior research with an intention to consider three distinct scenarios in which the new and old groups are similar, somewhat different, or extremely different (Kang & Petersen, 2012). The examinees in the old group were drawn from a standard normal distribution N(0, 1). The sample size was fixed at 3,000. Table 1 presents all conditions considered in this study. Results under all conditions were investigated using the six missing data methods for both the Haebara and the Stocking-Lord linking approaches.
Simulation Conditions.
Data Generation
To illustrate the simulation procedure, two test forms with 60 items and 20% common items were used as an example. First, 60 items were randomly drawn from an item pool of 800 sets of item parameters that were calibrated from real data, and were considered as the old form—Form Y. Among the items in Form Y, 20% of common items were deliberately selected such that the distributions of the common items represented the full test in terms of statistical characteristics, for example, item difficulty level. Next, the new form, Form X, was built to include common items already selected and another 40 items unique to Form X. The unique items on Form X were selected from the same item pool. In addition, efforts were spent to make sure that the distributions of the unique items on both forms were similar. In this way, the two forms were created to have similar item characteristics to reflect operational test construction. Finally, the abovementioned process was repeated for two test lengths (30 and 60 items) and two levels of common items proportions (20% and 40%). Responses to items on each test form were simulated based on the IRT 3PL model using the R program (R Core Team, 2021). The means and standard deviations of item parameters used in simulations are provided in Table 2.
Means and Standard Deviations of Item Parameters Used in Simulations
For the MAR scenario, the missing data generation process followed the procedure used in the previous literature (De Ayala et al., 2001; Enders, 2004; Finch, 2008). In this study, the number-correct score over the non-missing items was viewed as the observed variable inversely related to the probability of a data point that was missing. Specifically, to conduct the simulation, the number-correct score over the non-missing items was summed and then divided into three categories, each of which was assigned a probability of a missing response. For a 60-item test form where five common items were missing, the maximum value of the correct score over the non-missing items was 55. Thus, the possible correct scores for all examinees were classified into three categories, 0 to 18, 19 to 36, and 37 to 55. A missing probability was assigned to each category in which the higher score was inversely related to the missing probability. One condition here was that the averaged probability for each score category should be approximately equal to the desired missing percentage. For example, the probability for each category could be close to 0.4, 0.3, and 0.2 for 0 to 18, 19 to 36, and 37 to 55, respectively, such that the desired missing percentage, on average, was 30%. A range of ± 0.5% of the missing probability was allowed for each category to ensure the correct representation of the ratio of missingness while giving some extra room for the data to be generated (Finch, 2008). The probability of a missing value was then compared with a random value drawn from a uniform distribution. Based on the results of the comparison, the value was reserved or deleted in the datasets.
Under MNAR, responses to test items were assigned a corresponding missing probability such that examinees answering the items incorrectly were more likely to be assigned a higher probability of missingness (Finch, 2008). When generating data, the average probability of missing was set approximately equal to the target missing percentage. For example, to achieve an overall missing percentage of 15%, the missing probability of a correct response was assigned a value of 0.1, whereas the probability of an incorrect response was 0.27. As with MAR, the probability of a missing response was compared with a random value generated from a uniform distribution
The deletion of data points in the study was carried out only within common items. In fact, missing data within both common and unique items of each form may influence parameter estimates because they are calibrated based on responses to all items in both forms. However, when using separate calibration, transformation coefficients are calculated based on the responses to common items, which makes it more reasonable to assume that missing data within common items lead to a more significant impact on linking coefficients, compared to when missing occurs to unique items. This assumption led this study to consider only a situation where missing happened within common items. The parameters of the common items involving missing items are summarized in Table 3.
Item Parameters of the Common Items With Missing Responses
Six Missing Data Handling Approaches
The first four methods (LWD, IN, CM, and RF) were applied with the incomplete datasets using R. The R package MICE (van Buuren et al., 2015) was employed to implement MI. The package offers an option to conduct MI with dichotomous items by using an iterative approach to estimate logistic regressions and impute missing values using regression estimates for the dependent variable (Vidotto et al., 2015). After treating the incomplete datasets using each method, the computer program flexMIRT was used to estimate item and person parameters. For FIML, missing data were handled by using flexMIRT directly. Once item and ability parameters were obtained, POLYST (S. Kim & Kolen, 2003) was used to conduct scale linking with the Haebara (1980) and the Stocking and Lord (1983) methods.
Evaluation Criteria
The extent to which the estimated slopes and intercepts are deviant from the true linking relationship was used to evaluate the relative performance of the six missing data methods. If no error is involved in the linking process, the exact values of the transformation constants A and B are equal to
Three indices were used including: absolute bias
where
Results
Overall Performance of Missing Data Handling Methods
Aggregated results are presented across all study conditions for each type of missing data mechanism in Table 4. The results from the Haebara approach are presented only, as similar patterns were observed between the Haebara and the Stocking-Lord linking approaches.
Linking Error From Different Missing Data Handling Methods Under MAR and MNAR Using the Haebara Linking Approach
Note. The largest and the smallest values across the six methods were bolded and underlined, respectively. MAR = missing at random; MNAR = missing not at random; LWD = listwise deletion; IN = incorrect responses; CM = corrected mean; RF = response function; MI = multiple imputation; FIML = full information maximum likelihood; RMSE= root mean squared error.
Missing at Random
Table 4 shows minimal differences in linking errors among the missing data handling methods under the MAR missing mechanisms, except for LWD. Treating missing data with LWD produced a significantly large amount of errors than the other methods in recovering both slopes and intercepts. IN yielded the second largest bias and root mean squared error (RMSE) in slopes but led to the smallest errors in intercepts under MAR for the Haebara approach. CM produced small errors in slopes but yielded the second largest errors in intercepts. The other three approaches, including MI, FIML, and RF, tended to be not only associated with small errors in slopes but also in intercepts. In general, these three methods provided the most accurate linking results when the missingness was MAR.
Missing Not at Random
Although differences in the performance of the six methods were found between MAR and MNAR, the overall patterns were similar. The LWD approach yielded the largest amount of errors in both slopes and intercepts under MNAR, as can be seen in Table 4. Overall, differences among the approaches except LWD were minimal. In terms of the performance of the other five approaches, both CM and IN seemed to be associated with slightly larger bias and RMSE in slopes than the other three approaches. For intercepts, CM yielded larger errors than the other four approaches across all three evaluation criteria, and IN generally led to small errors. As with the results for MAR, RF, MI, and FIML were likely to produce more accurate linking coefficients in both slopes and intercepts as compared with other approaches.
The Impact of Test and Examinee Conditions
Linking accuracy was calculated and compared across different test conditions for each simulation factor to better connect the performance of missing data treatment methods with real-world practices. The results are presented in Tables 5 to 9 and discussed in the following paragraphs.
Linking Error From Different Missing Data Handling Methods by Test Lengths Using the Haebara Approach
Note. MAR = missing at random; MNAR = missing not at random; LWD = listwise deletion; IN = incorrect responses; CM = corrected mean; RF = response function; MI = multiple imputation; FIML = full information maximum likelihood; RMSE = root mean squared error.
Linking Error From Different Missing Data Handling Methods by Proportions of Common Items Using the Haebara Approach
Note. MAR = missing at random; MNAR = missing not at random; LWD = listwise deletion; IN = incorrect responses; CM = corrected mean; RF = response function; MI = multiple imputation; FIML = full information maximum likelihood; RMSE= root mean squared error.
Linking Error From Different Missing Data Handling Methods by Missing Rates Using the Haebara Approach
Note. MAR = missing at random; MNAR = missing not at random; LWD = listwise deletion; IN = incorrect responses; CM = corrected mean; RF = response function; MI = multiple imputation; FIML = full information maximum likelihood; RMSE= root mean squared error.
Linking Error From Different Missing Data Handling Methods by Percentages of Common Items Involving Missing Data Using the Haebara Approach
Note. MAR = missing at random; MNAR = missing not at random; LWD = listwise deletion; IN = incorrect responses; CM = corrected mean; RF = response function; MI = multiple imputation; FIML = full information maximum likelihood; RMSE= root mean squared error.
Linking Error From Different Missing Data Handling Methods by Ability Distributions Using the Haebara Approach
Note. MAR = missing at random; MNAR = missing not at random; LWD = listwise deletion; IN = incorrect responses; CM = corrected mean; RF = response function; MI = multiple imputation; FIML = full information maximum likelihood; RMSE= root mean squared error.
Test Length
In MAR and MNAR, most of the missing data handling methods, except for LWD, followed a similar pattern where they tended to produce slightly greater accuracy for a form with more items, as seen in Table 5. CM generally followed this trend, but it led to a similar amount of bias and RMSE in intercepts for a 30-item test and a 60-item test. Notably, treating missing responses with LWD showed a completely different pattern by leading to more errors, particularly in intercepts, as the number of items on a test form increased. By further examining the datasets, it was found that using LWD, responses from more examinees were deleted for a longer test, which caused greater errors in linking.
Ratio of Common Items
In general, most of the missing data treatment methods yielded slightly more accurate linking coefficients when there was a larger ratio of common items on a test form, as shown in Table 6. The exception was with LWD and CM. LWD yielded substantially larger errors in slopes and intercepts as the ratio of common items increased. This finding is partly explained by the fact that more responses involved missing values when there were more common items and using LWD led to deletion of those cases, which then affected the accuracy of scale linking negatively. CM yielded smaller errors for a higher ratio of common items under most conditions, but it produced larger values of bias and RMSE in intercepts regardless of the ratio of common items under MNAR. In some real-world practices, achieving a proportion of common items of 20% or 40% may not be possible be challenging. Thus, a condition where 10% of the items are common items was also investigated in the study for the 60-item test forms. No substantial changes were identified in the results under this condition in terms of the performance of the six missing data handling methods.
Missing Rates
In general, MI, FIML, and RF were shown to be robust to missing rates: a drift was minimal in linking accuracy across the three levels of missing rates in both MAR and MNAR, which can be found in Table 7. In contrast, the influence of missing rates on the other methods was more apparent. LWD introduced substantial errors for a test form with higher missing rates. Note that when LWD was used, the rate of change in errors was greater for a form with a missing percentage above 15% than a form with a lower missing percentage. Also, when CM and IN were employed, larger bias and RMSE were found in slopes and intercepts for a missing percentage of 30% as compared with those for 8% and 15% under most conditions.
The Proportion of Common Items With Missing Data
The performance of RF, MI, and FIML was very similar regardless of how many common items involved missing data, based on Table 8. Similarly, the errors produced by IN and CM remained constant under most test conditions with a few exceptions. For instance, there was only a slight increase in bias and RMSE in slopes for IN as missing data were observed within more common items. Similarly, relatively more errors were also found in intercept using CM, especially under the MNAR mechanism. LWD yielded considerably more errors for a form with a larger proportion of common items containing missing values.
Ability Distribution
According to Table 9, the largest error was consistently found when the new group differed the most from the old group, which was N(0.5, 1.22). Slightly more or similar linking errors were found when the new group followed N(0.25, 1.12) as compared with N(0, 1) when using IN, CM, RF, MI, or FIML. The rate of change in errors was greater when the new group had an ability distribution of N(0.25, 1.12) or higher as compared with the rate for a lower proficiency, which was more evident for IN and CM. LWD also yielded a large amount of errors when the new group was of a higher proficiency.
The Influence of Linking Approaches
In addition to the five simulation factors discussed, the efficiency of the six methods to treat missing data was also compared for both the Haebara and the Stocking-Lord methods. In general, a fairly consistent results were found between the two linking methods. The Stocking-Lord approach seemed to produce slightly more errors in slopes than the Haebara approach, with a few exceptions. For intercepts, results from the two linking methods were almost identical.
Discussion
Large-scale testing programs often administer multiple forms of a test to eliminate the chances of cheating and improve test security. One of the challenges under IRT is how to maintain scales consistently across different forms of a test. Scale linking methods are commonly used for large-scale assessments to achieve group comparability when common items are included in multiple test administrations. Under a CINEG, scale linking is a prerequisite to conducting many psychometric works when using a separate calibration method. For example, without having accurate linking coefficients, it is impossible for researchers to obtain precise equating coefficients, which will then undermine the validity of test scores. So far, many aspects of scale linking have been examined, but the understanding of the proper method to address missing responses in the process of scale linking is still rudimentary. Specifically, there is no clear guidance in selecting an appropriate approach to handling missing data with the consideration of real-world test conditions and missingness assumptions.
This study presented the results of a simulation study to understand the relative performance of using six different missing data handling methods on scale linking under two missing data mechanisms. Furthermore, a set of simulation conditions was varied, with the intent to create a comprehensive picture of the behaviors of missing data treatment methods in the context of scale linking.
In general, RF, MI, and FIML consistently demonstrated their superior performance over other methods based on the overall results, although some discrepancies were found between two mechanisms. In contrast, LWD always produced the largest errors regardless of the levels of factors used for simulating the datasets. CM and IN were identified to be associated with slightly larger or similar errors as compared with RF, MI, and FIML. However, it is crucial for researchers and practitioners to pay close attention to the specifics of test conditions when applying CM and IN in practice. For example, CM tended to generate substantially more errors when 15% of the responses to the common items were missing than that of a smaller missing rate. IN resulted in larger errors as the proficiency of the examinees taking the new form followed a distribution of N(0.5, 1.22) as compared with a less proficient group.
One of the major findings of this study was that MI and FIML seemed to consistently yield the smallest errors under the two missing mechanisms, which is in line with previous work (Enders, 2001b; Finch, 2008; Olinsky et al., 2003; Peyre et al., 2011). Prior research found that MI and ML tended to have a similar amount of errors (Collins et al., 2001), and the current study confirmed the trend in the context of scale linking. Using MI may be computationally demanding and time-consuming in practice, especially when imputation is needed for large datasets from multiple administrations. Although the MICE package has made multiple imputations more accessible, researchers still need to carry out the linking procedure multiple times to obtain the final averaged linking estimates over multiple imputations. Imputation is not required for FIML, but it also needs certain software to be available.
As introduced previously, the most inaccurate results were led by LWD in both MAR and MNAR, which seemed to be consistent with what was found in previous research (Enders, 2004; Robitzsch & Rupp, 2009). Once a large proportion of responses are deleted by using LWD, the remaining sample may not well represent the whole population such that biased results can be generated. In the context of scale linking, scales estimated from two forms are supposed to be placed on the same scale based on the responses to common items. If a subgroup of examinees who underperform did not respond to those items, item parameter estimates are likely to be influenced by not having a whole range of examinees, which will in turn affect scale linking accuracy.
A number of studies pointed out that using IN to handle missing data in the IRT context is not ideal (Cetin-Berber et al., 2019; Finch, 2008; Köhler et al., 2017; Pohl et al., 2014; Zhang & Walker, 2008). However, the linking results associated with IN presented in this study under MAR and MNAR were shown to be mixed. Specifically, this method led to slightly larger bias and RMSE than RF, MI, and FIML in slopes but had comparable or smaller results in intercepts. Notably, this method is relatively sensitive to a few simulation factors examined in this study, such as missing rate, the proportion of common items involving missing responses, and the difference in the ability distributions of the two groups for linking. Therefore, given that IN is one of the commonly used methods to treat missing responses in large-scale assessments, for example in PISA 2018 (OECD, 2020), more caution is needed to implement IN by taking different test conditions into consideration, especially in the context of scale linking.
The findings of the two single imputation methods seemed to be more complex. CM tended to produce slightly larger errors compared with other methods, except for LWD, under a few conditions. This might be related to the fact that the CM method involves the computation of an item mean using scores over non-missing responses divided by the number of examinees. With the increase in missing responses, information becomes limited in computing the item mean. By contrast, RF constantly resulted in small errors that were comparable to those of MI and FIML, which is aligned with the findings from previous studies (Finch, 2008, 2011). The calculation of RF does not require an item mean. Rather, it depends on rest scores for imputation, which accounts for examinees’ information across the entire form. Based on the design of the study, missingness only occurred within common items. Borrowing information from unique items with no missing data seemed to contribute to maintaining the accuracy of the imputation.
In terms of the simulation factors, RF, FIML, and MI led to the most accurate linking coefficients regardless of the study conditions. LWD was the most sensitive method to the choice of simulation factors under a variety of the conditions. The performance of IN and CM was slightly or moderately influenced by the simulation conditions. In general, the performance of the six missing data handling methods was consistent between the Haebara and the Stocking-Lord linking approaches. However, CM had slightly better performance when the Stocking-Lord method was used.
Conclusion and Future Research
Based on the results observed in this study, RF, MI, and FIML revealed to introduce a relatively small amount of errors for conducting scale linking. It is an important finding that RF, as a less complex method, also demonstrated robust performance under most of the test conditions investigated in the study. Other than using the two well-known missing data handling methods, MI and FIML, researchers may also consider imputing missing responses with RF for scale transformation. Another notable finding of this study is that researchers may avoid using LWD as it consistently revealed a large amount of error across various study conditions.
The list of missing data treatment approaches examined in the study is by no means exhaustive. In addition, the current study only focused on how to handle missing responses for omitted items specifically, without considering the treatment of non-reached items which might be more closely related to examinee’s motivational factors. Research has demonstrated that model-based approaches can provide more accurate parameter estimates (Debeer et al., 2017; Rose et al., 2010) due to their ability in treating omitted and no-reached items differently. In future studies, the accuracy of more complex model-based approaches on scale transformation can be investigated for missing responses caused by different reasons. In addition, the sample size was fixed at 3,000 in this study. Researchers might consider varying this factor and exploring its interaction with multiple missing data handling approaches.
The current study serves as a starting point in developing a better understanding of how to treat missing responses in maintaining IRT scales. Efforts to make linking estimates accurate when missing responses are present are an essential step in ensuring the comparability of item and person parameters and the validity of test scores. The goal of the study is to assist psychometricians and researchers in selecting the most appropriate approach for dealing with missing data in scale transformation. Future work may consider extending the current study design to the context of test equating for improving test form comparability and test fairness with the presence of missing responses.
Footnotes
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
