Abstract
Two Monte Carlo simulation studies investigated the effectiveness of the mean adjusted χ2/df statistic proposed by Drasgow and colleagues and, because of problems with the method, a new approach for assessing the goodness of fit of an item response theory model was developed. It has been previously recommended that mean adjusted χ2/df values greater than 3 using a cross-validation data set indicate substantial misfit. The authors used simulations to examine this critical value across different test lengths (15, 30, 45) and sample sizes (500, 1,000, 1,500, 5,000). The one-, two- and three-parameter logistic models were fitted to data simulated from different logistic models, including unidimensional and multidimensional models. In general, a fixed cutoff value was insufficient to ascertain item response theory model–data fit. Consequently, the authors propose the use of the parametric bootstrap to investigate misfit and evaluated its performance. This new approach produced appropriate Type I error rates and had substantial power to detect misfit across simulated conditions. In a third study, the authors applied the parametric bootstrap approach to LSAT data to determine which dichomotous item response theory model produced the best fit. Future applications of the mean adjusted χ2/df statistic are discussed.
Item response theory (IRT; Lord, 1980) is widely used for the measurement of cognitive ability. A critical issue in the application of IRT is ensuring that the model fits the data. One statistic commonly used to evaluate IRT model–data fit is the adjusted (to a sample size of 3,000) χ2/df ratio with values greater than 3 in a cross-validation data set proposed as indicating model misfit (Drasgow, Levine, Tsien, Williams, & Mead, 1995). This statistic has been widely used for examining IRT model–data fit for the purposes of differential item functioning (DIF; Kulas & Finkelstein, 2007; Rivers, Meade, & Fuller, 2009; Robie, Zickar, & Schmit, 2001; Zickar & Robie, 1999), scale development and evaluation (Rauch, Schweizer, & Moosbrugger, 2008; Scherbaum, Cohen-Charash, & Kern, 2006; Schmit, Kihm, & Robie, 2000), and IRT model comparisons to determine which model to use (Chernyshenko, Stark, Chan, Drasgow, & Williams, 2001; Stark, Chernyshenko, Drasgow, & Williams, 2006; Tay, Drasgow, Rounds, & Williams, 2009).
Although the adjusted χ2/df ratio has been widely used, there has not been a systematic examination of its effectiveness for evaluating IRT model–data fit for dichotomously scored items. In contrast, IRT indices for a variety of other applications have been evaluated: for example, Raju, van der Linden, and Fleer’s (1995) DIF analysis (e.g., Bolt, 2002; Flowers, Oshima, & Raju, 1999; Meade, Lautenschlager, & Johnson, 2007), person–fit measures such as Drasgow, Levine, and Williams’s (1985) z3 (e.g., Drasgow, Levine, & McLaughlin, 1987), and Tatsuoka’s (1984) extended caution indices (e.g., St-Onge, Valois, Abdous, & Germain, 2009), among others. Recently, an examination of the adjusted χ2/df ratio statistic for polytomously scored items from scales of different lengths (10 and 20 items) with various sample sizes (500, 1,000, and 2,000) showed that a cutoff value of 3 produced large Type I error rates for cross-validation data (LaHuis, Clark, & O’Brien, 2011). In this study, we simulated dichotomous items to ascertain whether the adjusted χ2/df ratio can be used for assessing model-fit in this context. We also examine a wide range of data conditions. As an alternative to the cutoff value of 3, which may not be generalizable across conditions, we propose and test the use of a parametric bootstrap approach to the adjusted χ2/df ratio for assessing absolute model–data fit.
The Adjusted χ2/df Ratio: Singles, Doubles, and Triples
Drasgow et al. (1995) proposed the use of an IRT model–data fit based on the ordinary χ2 value,
where Oik is the observed frequency for response option k to item i, and Eik is the expected frequency, which is computed from the item response function for option k integrated over a standard normal density f(θ) multiplied by the total sample N,
This χ2-statistic was termed the singles χ2. However, because the singles χ2 is a marginal statistic, it is insensitive to certain types of misfit (Van den Wollenberg, 1982). Hence, another statistic, termed the doubles χ2, was developed for a pair of items i and i′. Observed frequencies are compared with the expected frequency of endorsement given by
By extension, the triples χ2 value is computed using a multiway contingency table for triples of items. We note that the doubles χ2 value is analogous to the bivariate residual statistic implemented in Latent GOLD (Vermunt & Magidson, 2000, 2002, 2005). Both examine the correspondence between the model-based expected values and the observed values.
One problem with the χ2 is its sensitivity to sample size because the expected value of a noncentral χ2 is equal to its df plus N times its noncentrality parameter δ,
To resolve this problem, Drasgow and his colleagues recommended that the χ2 indices be adjusted to a fixed sample size of 3,000, with the expectation that these indices would be comparable across different sample sizes. Thus, an estimate of the noncentrality parameter is
An observed χ2 can be adjusted to fixed sample size of, say, n = 3,000, by
After the adjustment, the values can be divided by the degrees of freedom to produce the adjusted χ2/df value,
As a heuristic, Drasgow and his colleagues recommended that a cutoff value of 3 be used as an index of misfit. In their calculations, they took the df as the number of cells in a contingency table minus one.
There are still several issues which are not resolved with the use of an adjusted χ2/df value. Foremost, there has not been a systematic examination of the proposed cutoff value across a variety of data conditions. Indeed, for this statistic and recommended critical value to be useful, it has to be generalizable across different sample sizes and test lengths. In addition, although cross-validation samples have been recommended as a way to ensure that models are not overfit to the data (Drasgow et al., 1995), it has been found that validation samples have little power to detect model misspecification when the sample sizes are small (<2,000; Tay, Ali, Drasgow, & Williams, 2011). Joe and Maydeu-Olivares (2006) have shown that the χ2-statistic does not follow a χ2 distribution for cross-validation samples. However, are there conditions where the adjusted χ2/df statistic can be used if the sample size is large enough?
Furthermore, a critical question is whether the singles-, doubles-, or triples-adjusted χ2/df fares better in detecting misfit and whether the same critical value applies across the three variations of the same statistic. We expected that because the singles χ2/df is a marginal statistic, it would be less sensitive in detecting model misspecifications as compared with the doubles- or triples-adjusted χ2/df values.
Study 1
In this study, we used a simulation to examine (a) whether the rule of thumb of 3 holds across a range of sample sizes and test lengths, (b) the effects of using a calibration or cross-validation sample, and (c) the usefulness of the mean singles-, doubles-, and triples-adjusted χ2/df. We focused on only dichotomously scored items as these are most commonly used in cognitive ability testing. Specifically, we simulated data based on the one-parameter (1PLM), two-parameter (2PLM), or three-parameter logistic (3PLM) IRT model and subsequently fit all three models to the data. Fitting the correct IRT model to the data (null condition) or a more general model to the data (e.g., 3PLM to 2PLM data) allows us to ascertain the number of instances the cutoff value of 3 is breached, producing the Type I error rate. Alternatively, by fitting a more restricted IRT model to the data (e.g., 2PLM fit to 3PLM data), we were able to ascertain power based on a cutoff value of 3. In addition to examining the power to detecting whether a more restricted unidimensional IRT model fits the data (e.g., 2PLM fit to 3PLM), we also examined whether unidimensional IRT models fit multidimensional IRT data. We evaluated the effectiveness of the adjusted χ2/df index for detecting model misspecification by adjusted power, which equals power minus the actual Type I error rate.
In this study, we used sample sizes of 500, 1,000, 1,500, and 5,000 simulees crossed with test lengths of 15, 30, and 45. Across each of these 4 (sample size) × 3 (test length) = 12 conditions for three-dimensionality, 200 replications were undertaken in which data were simulated, IRT models were fit, and the adjusted χ2/df ratio was computed. The dimensionality conditions included unidimensionality simulations as well as simulations of low and moderate multidimensionality where two latent traits were correlated at .80 and .60, respectively.
Simulation
First, data were simulated based on the 1PLM, 2PLM, or 3PLM. The 3PLM is given by
where the probability of correctly answering an item (ui = 1) is based on the simulated θ j sampled from a standard normal distribution, ai is the item discrimination parameter sampled from a log-normal (0, .5) distribution and divided by 1.702, bi is the item difficulty parameter randomly sampled from a uniform (−2, 2) distribution, and ci was sampled from a logit-normal (−1.1, 0.5) distribution resulting in a mean value of ci = 0.25. The 2PLM data were simulated by fixing ci = 0 and the 1PLM was simulated with ci = 0 and a single ai value for all the items.
Multidimensionality was simulated following procedures by Zhang and Stout (1999). The latent trait distribution was bivariate normal with zero means, unit variance, and a correlation of .80, or .60, reflecting low and moderate multidimensionalities, respectively. The odd-numbered items were set to measure the first dimension and the even-numbered items were set to measure the second dimension, resulting in a simple structure with two factors.
Next, all three IRT models were fit to the data using BILOG-MG (Zimowski, Muraki, Mislevy, & Bock, 1996) using the default settings, and the adjusted χ2/df ratios for each fitted IRT model was computed using FORSCORE (Williams & Levine, 1993) for the calibration samples as well as a cross-validation sample with the same number of simulated items and people. The measure used in this study was the mean adjusted χ2/df ratio for all singles, doubles, and triples for the test.
Type I Error and Adjusted Power
Both Type I error and adjusted power were determined by the proportions of replications that had mean adjusted χ2/df ratios greater than 3. These proportions were considered Type I errors when the correct model or a more general IRT model was fit to the data. In contrast, adjusted power was determined when a less restricted model was fit to data from a more general IRT model.
Results
For our results, we only present Type I error rates when the correct model was fit to the data because similar values were obtained when a more general IRT model was fit to the data. In view of space considerations, we present detailed results for moderate multidimensionality (latent traits correlated .60) in the tables. The trend was very similar when data were less multidimensional (latent traits correlated .80); in general, there was less power to detect low multidimensionality as compared with moderate multidimensionality.
1PLM data
Table 1 shows the results of fitting the 1PLM to unidimensional and multidimensional 1PLM data. For the calibration samples, the Type I error rates were uniformly close to zero and, therefore, we do not present them in Table 1. For the validation sample, high Type I error rates are evident for relatively smaller sample sizes of 500 but are close to the desired nominal rate when the sample size was much larger (i.e., 5,000). Indeed, we expected that item parameter estimation error would be larger for smaller sample sizes, leading to higher Type I error rates as compared with the other conditions.
Proportion of Mean Adjusted χ2/df Ratios Greater Than 3 for 1PLM Simulated Data
Note. 1PLM = One-parameter logistic model. Type I error rates for calibration samples were zero across all simulated conditions and are omitted.
When the data were multidimensional, the doubles- and triples-adjusted χ2/df ratios had higher adjusted power as sample size and test length increased. However, singles-adjusted χ2/df ratios were not useful in detecting multidimensionality as evidenced from low adjusted power. As expected, adjusted power increased as multidimensionality increased. Using the triples-adjusted χ2/df ratios, the average power across the conditions to detect moderate multidimensionality was .57 (latent traits correlated at .60) whereas the average power to detect low multidimensionality was .11 (latent traits correlated at .80).
2PLM data
As with the simulated 1PLM data, Type I error rates were zero for the calibration samples. Again, for validation samples, the Type I error rates were much higher for smaller sample sizes but were close to zero for sample sizes of 5,000 as shown in Table 2.
Proportion of Mean Adjusted χ2/df Ratios Greater Than 3 for 2PLM Simulated Data
Note. 1PLM = one-parameter logistic model; 2PLM = two-parameter logistic model. Type I error rates for calibration samples were zero across all simulated conditions and are omitted.
Using a cutoff value of 3, the adjusted power to detect misfit—when the 1PLM was fit to 2PLM data—for doubles- and triples-adjusted χ2/df ratios was moderate to excellent across all conditions, including calibration and validation data sets. Adjusted power increased as a function of a larger sample size and test length. However, unlike 1PLM data, the adjusted power to detect multidimensionality was poor. For calibration samples, adjusted power was uniformly close to zero for both low and moderate multidimensionality. Importantly, the mean adjusted χ2/df ratios for the multidimensional data were consistently elevated as compared with when 2PLM was correctly fit to unidimensional data, but values were lower than the cutoff value of 3. For example, the mean values of the triples-adjusted χ2/df ratio for 15 items ranged from −.71 to 2.02 when data were moderately multidimensional whereas they ranged from −2.24 to 0.68 when the data were unidimensional.
For validation samples of small to moderate sample sizes, Type I error rates were too high for power to be useful, whereas Type I error rates were low for sample sizes of 5,000. For this largest sample size, the adjusted power to detect moderate multidimensionality was low, ranging from .06 to .23 for doubles- and triples-adjusted χ2/df ratios. In addition, for sample sizes of 1,500 to 5,000, the mean adjusted χ2/df ratios for doubles and triples decreased from above the critical value of 3 to below 3. This resulted in low power to detect multidimensionality even when Type I error rates were low. These results clearly show that a simple cutoff value of 3 is not generalizable.
3PLM data
The Type I error rates were very low and uniformly close to zero for calibration data as with the simulated 2PLM and 1PLM data (Table 3). For validation data, Type I error rates also had the same trend as found before—Type I error rates were high for small sample sizes but were low for the largest sample size (i.e., 5,000).
Proportion of Mean Adjusted χ2/df Ratios Greater Than 3 for 3PLM Simulated Data
Note. 1PLM = One-parameter logistic model; 2PLM = two-parameter logistic model; 3PLM = three-parameter logistic model. For calibration sample, Type I error rates were zero across all simulated conditions; adjusted power was zero for multidimensional 3PLM (latent traits correlated at .60) and for 2PLM. These are omitted from the table.
When a 1PLM was fit to 3PLM data, the adjusted power was different between validation and calibration samples. Therefore, although there was little difference in adjusted power between validation and calibration samples when the 1PLM was fitted to 2PLM data, validation samples had higher adjusted power than calibration samples in this instance. When fitting the 2PLM to 3PLM data using a cutoff value of 3, there was little or no adjusted power to detect misfit for both calibration and validation data. Hence, these two models were not distinguishable using the rule of thumb. Similarly, there was little power to detect moderate multidimensionality for both calibration and validation data. Although the mean adjusted χ2/df was higher when a misspecified model was fit to the data, many values were lower than the critical value of 3 resulting in low power.
Discussion
This initial study contributes to the current literature on IRT model–data fit using the adjusted χ2/df ratios in several important ways. Past studies have frequently used the cutoff value of 3 across all three different adjusted χ2/df ratios—singles, doubles, and triples—as a justification that the IRT model fits the data. To our knowledge, this study was the first to use Monte Carlo simulations to rigorously determine the effectiveness of this cutoff value for the different adjusted χ2/df ratios. We found that adjusted χ2/df ratios were generally effective for detecting model misfit when a 1PLM was fit to 2PLM or 3PLM data in doubles and triples but not in singles. Furthermore, with large sample sizes of 5,000 or more, analysis of validation data was somewhat effective. This confirms the recommendations by Drasgow et al. (1995) and their rule of thumb. Thus, in spite of past research showing that the χ2 value does not follow a χ2 distribution in validation samples (Joe & Maydeu-Olivares, 2006), this approach to assessing goodness of fit appears to be reasonable in large samples. However, when moderate sample sizes are used (i.e., 500 to 1,500), the use of validation samples is not recommended because of high Type I error rates.
Nevertheless, our simulations suggest that the 3PLM and 2PLM are not easily distinguishable on the basis of the adjusted χ2/df ratio for a cutoff value of 3; similarly, there was little power to detect moderate multidimensionality for the 3PLM and 2PLM. This occurred because even though the mean adjusted χ2/df ratios were elevated for misspecified models, they did not usually exceed the cutoff value of 3. To ascertain the relative fit of several models, we recommend another strategy: Fit different IRT models to the data and compare the relative magnitudes of the adjusted χ2/df ratios for the calibration data set. Indeed, this strategy has been undertaken in the past when attempting to compare the fits of different IRT models with empirical data (e.g., Chernyshenko et al., 2001; Stark et al., 2006; Tay et al., 2009). Furthermore, simulation studies have shown that this method is helpful to distinguish dominance and ideal point data (Tay et al., 2011). This is particularly important given that the adjusted χ2/df ratio appears to be affected by sample size and test length despite the sample size adjustment. However, this method does not give an indication of absolute fit. In Study 2, we propose a parametric bootstrap approach for assessing absolute IRT model-fit.
Finally, we also determined that not all types of adjusted χ2/df ratios are equally effective for detecting misfit. Specifically, the singles-adjusted χ2/df ratio had very low power and we do not recommend its use. Instead, we recommend that the doubles- and triples-adjusted χ2/df ratios be used in the future. For the types of model misspecification studied here, triples had consistently larger power than doubles.
Study 2
Method
Instead of using a fixed cutoff value of 3 for examining model-fit, which Study 1 found to be ineffective across various conditions, we propose the use of a parametric bootstrap approach instead (Langeheine, Pannekok, & Van de Pol, 1996; Vermunt, 2001; von Davier, 1997). The general procedure is as follows:
Following procedures from Study 1, we simulated data from one of the dichotomous IRT models. This data set is referred to as the observed data.
We estimated item parameters for a model (i.e., a 1 PLM, 2 PLM, or 3PLM) and calculated the mean adjusted χ2/df ratio; this is referred to as the observed mean adjusted χ2/df value.
The estimated item parameters from the observed data were then used to simulate data using the IRT model from Step (b). This model was then fit to the simulated data. The mean adjusted χ2/df ratio for the IRT model was calculated for this new simulated data set.
Step (c) was repeated multiple times (1, . . ., m) so that a sampling distribution for the mean adjusted χ2/df ratio for the IRT model from Step (b) was obtained; the observed adjusted χ2/df value was then compared with the sampling distribution to obtain an empirical p value; a relatively larger observed adjusted χ2/df value would have a relatively smaller p value.
Steps (a) to (d) were then replicated (1, . . ., r) to determine power and Type I error rates. Power was determined by the proportion of replications that had an empirical p value less than .05 when a more restricted model was fitted to the observed data (e.g., by fitting a 1PLM to 3PLM observed data). In contrast, the Type I error rate is the proportion of replications that had an empirical p value less than .05 when the correct model, or a less restricted model, was fitted to the observed data (e.g., a 3PLM was fit to 1PLM observed data).
We used the same sample sizes and test lengths as the first simulation study. Across each of these 12 conditions, m = 50 replications were taken and r = 50 bootstrap samples drawn per replication. As before, for each of the conditions, the simulated observed data conformed to the 1PLM, 2PLM, or 3PLM. We decided to use a smaller number of replications and bootstrap samples because the results from m = 100 replications and r = 200 bootstrap samples were very similar in the trial runs as using a large number of replications and bootstrap samples is very time consuming. For example, for the 1PLM 15-item condition, the Type I error rates (m = 50 and r = 50) for singles, doubles, and triples were .02, .06, and .05, respectively; whereas a larger number of replications and bootstrap samples (m = 100 and r = 200) yielded Type I error rates of .02, .03, and .03, respectively. Furthermore, each replication was very involved, requiring 100 bootstrap samples of each IRT model (i.e., 1PLM, 2PLM, and 3PLM) and calculation of the singles-, doubles-, and triples-adjusted χ2/df ratio; as a result, each replication required 2 hours, and a total of 2 hours × 100 replications × 12 conditions × 3 different dichotomous IRT models = 14,400 computing hours would be needed for 200 bootstrap samples of 100 replications per condition.
Results
1PLM data
Table 4 presents the Type I error rates and the adjusted power to detect misfit using the bootstrap approach. Foremost, the Type I error rates were controlled better than with the greater-than-3 rule of thumb. They ranged from .00 to .14, for singles-, doubles-, and triples-adjusted χ2/df ratio when fitting the 1PLM to the 1PLM data. Although not presented in the table, we found that the Type I error rates (M = .79) were much higher when fitting the 3PLM to the 1PLM data. This was a function of the default c-prior in BILOG; assuming four choices in an item, the c-prior is given a mean of .25. This resulted in item parameter estimates that were consistently biased away from the simulated 1PLM item parameters. When a low 3PLM mean prior (i.e., c = .01) was specified for the 1PLM observed data, the bootstrap approach yielded excellent Type I error rates (M = .02) for the 3PLM, equivalent to the 1PLM. These results show that the 3PLM can fit 1PLM data with an appropriate prior for c; however, the results also suggest that we can determine when the 1PLM is sufficient because it fits better than the 3PLM with BILOG’s default c-prior.
Type I Error Rates and Power to Detect Misfit Using a Replication Approach
Note. 1PLM = One-parameter logistic model; 2PLM = two-parameter logistic model; 3PLM = three-parameter logistic model.
Table 4 shows that there was high adjusted power (M = .92) to detect moderate multidimensionality using doubles- and triples-adjusted χ2/df ratio. When their low multidimensionality was simulated, the average power (M = .84) for both doubles- and triples-adjusted χ2/df was lower, as would be expected.
2PLM data
For Type I error rates, the same trends were observed as with the 1PLM data. Although not presented in the table, we found that fitting a 2PLM resulted in excellent Type I error rates (M = .07). However, fitting the 3PLM with the default prior for c produced high Type I errors (M = .72). When a low BILOG c-prior was used, the Type I error rates were substantially lower (M = .02).
Importantly, the adjusted power to detect misfit was consistently higher than .84 for doubles and triples across all the simulated conditions when a 1PLM was fit to the 2PLM data. In contrast, the singles-adjusted χ2/df ratio had little power. This again confirms the notion that the singles-adjusted χ2/df ratio is much less useful than the doubles and triples. When fitting the 2PLM to moderately multidimensional 2PLM data, the adjusted power for doubles- and triples-adjusted χ2/df ratio was also high (M = .91). For less multidimensional 2PLM data, the adjusted power was lower (M = .83) but increased with sample size and test length. This is quite different from the results shown in Table 2 and illustrates the advantage of the parametric bootstrap approach.
3PLM data
When the 3PLM was fit to simulated 3PLM data, the Type I error rates were excellent to slightly inflated, ranging from .00 to .20, respectively. Although the Type I error rates at times were higher than the nominal .05 level, this is a vast improvement over using a critical value of 3. Across all conditions, the mean Type I error rates for singles-, doubles-, and triples-adjusted χ2/df ratios were .05, .09, and .09, respectively. When a more restricted IRT model—the 2PLM—was fit to 3PLM data, the power to detect misfit was excellent for longer tests (i.e., more than or equal to 30 items). Excluding the 15-item test length condition, the average adjusted powers of doubles- and triples-adjusted χ2/df ratios were .84 and .86, respectively. Adjusted power was lower, however, for the simulated 15-item tests. This shows that a longer test length is needed to differentiate the 2PLM and 3PLM. The power to detect misfit was higher when a 1PLM was fit to 3PLM data for the doubles- and triples-adjusted χ2/df ratios, with an average of .90 for both doubles and triples. In contrast, the singles-adjusted χ2/df ratio was much less effective.
Empirical Example
We applied the parametric bootstrap approach to the LSAT 6 data from Bock and Lieberman (1980), which consists of the responses of 1,000 participants to five dichotomous items. Joe and Maydeu-Olivares (2006) obtained goodness-of-fit statistics for the 1PLM, 2PLM, and 3PLM fit to a split-half sample of the LSAT 6 data; their p values were .52, .33, and .13, respectively. They suggested that the 1PLM fit provided a satisfactory fit because it is more parsimonious. Using the full sample, the parametric bootstrap triples-adjusted χ2/df ratio p values were .80, .49, and .46 for the 1PLM, 2PLM, and 3PLM, respectively; the parametric bootstrap doubles-adjusted χ2/df ratio p values were .83, .47, and .50 for the 1PLM, 2PLM, and 3PLM, respectively. These statistics show that the 1PLM fit the LSAT 6 data quite well. Importantly, the parametric bootstrap procedure produced results that were similar to that of Joe and Maydeu-Olivares (2006).
General Discussion
Proposing a simple rule of thumb, such as a value of 3, for ascertaining misfit using the mean adjusted χ2/df is helpful but may be misleading in some circumstances. This is because the cutoff value of 3 does not apply across conditions with varying sample sizes and test lengths. This rule of thumb is especially ineffective for smaller sample sizes. In addition, there are important differences in whether this rule is applied to the calibration or validation sample. To overcome the problems with the original adjusted χ2/df ratio, we proposed and tested a bootstrap approach and found that it was useful for ascertaining model–data fit. Specifically, it provided reasonably good control of Type I errors with very high power to identify misspecified models. Currently, we are developing software that implements this approach so that it can be easily used by researchers and practitioners—specifically, users can input item parameters and the number of bootstrap samples. The shell program then implements the routine by calling the relevant IRT software to create an appropriate sampling distribution.
In addition to finding that the performance of the mean adjusted χ2/df can be improved by the bootstrap method, this study provides several other interesting results. First, the singles statistic provides little power and should not be used. Apparently, BILOG (and presumably other IRT estimation software) can fit the marginal numbers of right and numbers of wrong regardless of whether an IRT model fits. In contrast, fitting 2 × 2 and 2 × 2 × 2 contingency tables presents a more challenging task. We found that more restrictive models were unable to fit data from more general models in that the chi-squares that resulted from the two-way and the three-way tables were relatively large. Interestingly, the bootstrap method provided fairly good control of Type I error rates when the correct model was fit and when a more general model was fit with appropriate specifications for the priors. It is arguable whether a Type I error occurs when the 3PLM with an inappropriate prior for c is rejected when fit to 1PLM data: Clearly, there is misspecification.
Drasgow et al. (1995) developed the adjusted χ2/df measure in the context of fitting IRT models to large data sets. In retrospect, it is clear that this measure confounds goodness of fit of an IRT model with estimation error when it is applied in a cross-validation sample. Tables 1 and 2 show that this confound is very serious when sample sizes are 1,500 and less but not a problem for samples of 5,000. In essence, the bootstrap method ascertains the degree to which estimation error affects the χ2/df ratio and thereby provides a basis for judging what constitutes an unexpectedly large ratio.
How does the parametric bootstrap procedure compare with other newer IRT fit statistics? A comparison of the parametric bootstrap adjusted χ2/df ratio shows that it may have higher adjusted power than that of the S-χ2 procedure (Orlando & Thissen, 2000). We calculated the adjusted power from Table 1 of Orlando and Thissen’s (2000) article. For a sample size of 1,000 simulees and 40 items, the adjusted power was .52 when fitting the 1PLM to 3PLM data; .47 when fitting the 1PLM to 2PLM data; and .07 when fitting the 2PLM to 3PLM data. In contrast, for a sample size of 1,000 simulees and 30 items, the parametric bootstrap adjusted χ2/df ratio has adjusted powers of .90, .94, and .94, respectively. Adjusted power was even higher for a test length of 45 items.
Future Research
In this article, we focused on the adjusted χ2/df ratio because it has been frequently used even though its properties were poorly understood. Our simulations clearly identified a shortcoming and consequently we proposed and examined an improvement to the adjusted χ2/df ratio. In future research it will be important to carefully examine how the proposed parametric bootstrap approach compares with other IRT fit indices.
We also found that a more general model such as the 3PLM may not fit simulated 1PLM data well. When the c-prior was adjusted lower, the model fit the data substantially better as indicated from the low Type I error rates. Practically, this suggests that overparameterization may be detected with the parametric bootstrap approach.
Conclusion
In sum, evaluating the fit of an IRT model is important for practical applications (e.g., DIF analyses) and theoretical considerations (e.g., determining whether a dominance or ideal point process underlies responses). In this article, we have proposed an improvement on the adjusted χ2/df ratio measure and shown it to function better than the original. Although the parametric bootstrap method requires extensive computations (about 2 hours on a modern personal computer to compare three different IRT models), its qualitatively superior performance appears to justify its use. We expect that in the future, increased computation power would reduce the time needed to obtain a sampling distribution and enable a more extensive bootstrapping process.
Footnotes
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
The author(s) received no financial support for the research, authorship, and/or publication of this article.
