Abstract
In applications of item response theory (IRT), an estimate of the reliability of the ability estimates or sum scores is often reported. However, analytical expressions for the standard errors of the estimators of the reliability coefficients are not available in the literature and therefore the variability associated with the estimated reliability is typically not reported. In this study, the asymptotic variances of the IRT marginal and test reliability coefficient estimators are derived for dichotomous and polytomous IRT models assuming an underlying asymptotically normally distributed item parameter estimator. The results are used to construct confidence intervals for the reliability coefficients. Simulations are presented which show that the confidence intervals for the test reliability coefficient have good coverage properties in finite samples under a variety of settings with the generalized partial credit model and the three-parameter logistic model. Meanwhile, it is shown that the estimator of the marginal reliability coefficient has finite sample bias resulting in confidence intervals that do not attain the nominal level for small sample sizes but that the bias tends to zero as the sample size increases.
In classical test theory, the reliability of a test plays a central role. The reliability is a measure of the consistency of the scores from a test. Reliability, together with the concept of item and test information, also plays a role in item response theory (IRT). Two different measures of reliability in IRT are marginal reliability (Cheng, Yuan, & Liu, 2012) and test reliability (Kim & Feldt, 2010; Lord, 1977, 1980). The marginal reliability denotes the ratio of the true score variance to the total variance, expressed with respect to the estimated latent abilities. Test reliability, on the other hand, denotes the ratio of true score variance to total variance expressed with respect to the sum scores. Note that, in the terminology used in this article, both the marginal reliability and the test reliability refer to reliability with regard to a population as a whole and hence are in some sense both “marginal” measures of reliability. Having a measure of the reliability of the scores from a test in the context of IRT is useful since it provides an indication of the overall consistency of the test scores in measuring the underlying trait or the observed score generated from the IRT model. Since the object of a test in IRT is to measure the latent concept for everyone taking the test, an overall measure of the reliability of the test scores across the spectrum of the latent distribution is important to consider.
When using IRT in empirical research, a reliability coefficient is often reported. The reported reliability coefficients are estimates of the true reliability coefficient and these estimates have a certain amount of variance associated with them due to the item parameter estimation. The large sample variance for several reliability coefficient estimators in classical test theory have been derived (Yuan & Bentler, 2002). However, analytical estimates of the variance of the IRT reliability coefficients are not available in the literature and an estimated standard error or confidence interval is therefore usually not reported. In this article, the large sample variance of the estimators of the IRT marginal reliability coefficient and the IRT test reliability coefficient are derived using standard asymptotic theory. The results may be used by empirical researchers to estimate confidence intervals for the reliability coefficients.
The article is structured as follows. First, IRT is introduced and common models for either dichotomous or polytomous data are briefly described. Then, IRT marginal and test reliability coefficients are defined and the asymptotic variance of estimators of these are derived. It is then investigated how well the large sample results work with finite samples through a simulation study. Finally, concluding remarks are given.
Item Response Theory
Consider a test consisting of
where
where
where
where, for the 3-PL model,
where (Hambleton & Swaminathan, 1985)
For a polytomous IRT model, define the expected item information as (Magis, 2015)
where, for the GPCM (Muraki, 1992),
and
resulting in the expected information for the GPCM having the expression
The test information is the sum of the information for each item, i.e.,
Let
Item Response Theory Marginal Reliability
One measure of reliability for scores on the latent variable metric which has been proposed in the literature is the marginal reliability (Green, Bock, Humphreys, Linn, & Reckase, 1984), sometimes referred to as parallel forms reliability (Kim, 2012). With the assumption of an ability distribution with density
The marginal reliability coefficient can be interpreted as the reliability with regard to the maximum likelihood ability estimates. The density function
Item Response Theory Test Reliability
The IRT test reliability coefficient
where, following Kim and Feldt (2010),
where the conditional error variance
Consider a test that has possible observed scores from 0 to
where
Asymptotic Variance of Item Response Theory Reliability Coefficient Estimators
Define
Item Response Theory Marginal Reliability
The integral in Equation (13) is approximated by a sum using Gauss-Hermite quadrature and we thus obtain the estimator of
where
where
In Equation (20), the expression of
and for polytomous IRT with the GPCM is
For the 3-PL model,
where
Item Response Theory Test Reliability
The integral implicit in the numerator of Equation (14) is approximated by a sum using Gauss-Hermite quadrature and so are the integrals required to calculate
Again,
where
Note that
and
where the derivatives
Confidence Interval Estimation
For large sample sizes, approximate confidence intervals for the reliability coefficients can be constructed using the derived variances and a normal approximation. Hence, approximate 95% confidence intervals for
Simulation Study
Simulations were conducted to study the finite sample properties of the derived asymptotic variances and confidence intervals. The values of the latent variable were drawn from the
Generalized Partial Credit Model Item Parameters for the Nine-Item Scale Used in the Simulation Study.
Three-Parameter Logistic Item Parameters Used in the Simulation Study.
Results
In Table 3 the results for the simulation with the GPCM are given. For the marginal reliability coefficient, there exists a small but statistically significant bias for sample sizes lower than 2,000. The asymptotic standard errors are accurate as an estimate of the sampling variability for all sample sizes. The empirical coverage rates of the 95% confidence intervals are slightly lower than the nominal level for sample sizes 250 and 500 but with sample sizes 1,000 and higher the empirical coverage rate is not statistically significantly different from 95%. For the test reliability coefficient, the bias is not statistically significantly different from zero for any sample size. The asymptotic standard error is accurate for all sample sizes considered. The empirical coverage rates of the 95% confidence intervals for the test reliability coefficient are not statistically significantly different to the nominal level for any sample size.
Mean Bias (×100), Asymptotic and Monte Carlo Standard Errors (×100), and Coverage of 95% Confidence Intervals (%) for the GPCM Reliability Coefficient Estimators, With Estimated Standard Errors in Parentheses.
Note. GPCM = generalized partial credit model; ASE = asymptotic standard error; MCSE = Monte Carlo standard error; CI = confidence interval.
The results for the 3-PL model simulation are shown in Table 4. The estimator of the marginal reliability coefficient has a small but statistically significant bias under all settings considered. The asymptotic standard errors are accurate for all sample sizes but the empirical coverage rates of the confidence intervals are statistically significantly different from the nominal level of 95% with all sample sizes except the highest sample size of 8,000. For the test reliability coefficient, the bias is smaller and not statistically significantly different from zero with sample sizes 2,000 and higher. The asymptotic standard errors are accurate for all sample sizes and the empirical coverage rate is not statistically significantly different from the nominal level for any sample size.
Mean Bias (×100), Asymptotic and Monte Carlo Standard Errors (×100), and Coverage of 95% Confidence Intervals (%) for the 3-PL Reliability Coefficient Estimators, With Estimated Standard Errors in Parentheses.
Note. 3-PL = three-parameter logistic; ASE = asymptotic standard error; MCSE = Monte Carlo standard error; CI = confidence interval.
Concluding Remarks
With the results presented in this article, empirical researchers and practitioners have access to methods that evaluate the variability of reliability coefficient estimators in IRT and with which confidence intervals for the reliability coefficients can be estimated. The simulation study indicates that the estimated confidence intervals for the test reliability coefficient have good coverage properties with sample sizes as small as 250 with the GPCM and 1,000 with the 3-PL model. The estimator of the marginal reliability coefficient is slightly biased with small samples when using the GPCM and with all sample sizes considered in this article when using the 3-PL model. With the GPCM, the estimated confidence intervals were however still largely accurate and had correct coverage with sample size 1,000 while with the 3-PL model the confidence intervals had correct coverage only with the largest sample size considered. If sufficient computational resources are available, a nonparametric bootstrap approach (Davison & Hinkley, 1997) can be used to estimate the finite sample bias of the reliability coefficient estimators. This bias estimate can then be used together with the results in this article to generate bias-adjusted confidence intervals, which will achieve an improved coverage rate.
The differences between the properties of the marginal and test reliability estimators can be attributed to the differences with regard to the estimation of item response functions and the estimation of expected information functions. In Ogasawara (2002), it was shown that item response functions were possible to estimate accurately in spite of unstable item parameter estimates while the expected information functions were more affected by unstable item parameter estimates. Since the marginal reliability is calculated from the expected information functions while the test reliability only uses the item response functions, the difference in the accuracy of the two different reliability estimators is consistent with the results of Ogasawara (2002).
Since the marginal and test reliability coefficients measure different concepts, a direct comparison between them is not very meaningful. However, this study does indicate that the estimator of the test reliability coefficient has slightly higher sampling variability but lower bias than the estimator of the marginal reliability coefficient when using the same item parameters.
When calculating and reporting the IRT reliability coefficient estimates, it is important to note that the reliability coefficients are not invariant with regard to the latent distribution. This means that the reliability coefficients will be different for populations with different latent distributions even if the item parameters are invariant. If the distribution parameters of the population are known, the results of this paper can incorporate these without any changes to the derivations presented by suitably changing the approximation of the integrals needed. If the distribution parameters have been estimated, for example by using parametric (Mislevy, 1984), nonparametric (Bock & Aitkin, 1981), semi-parametric (Woods, 2006) or multiple group (Muthén & Lehman, 1985) methods, an extension to the methods in this article is required where the derivatives with respect to the distribution parameters are derived.
Footnotes
Appendix
With the GPCM, for each item
where
and, for
With the 3-PL model, for each item
and
where
and
The derivatives
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: Tao Xin declared funding by the National Natural Science Foundation of China (Grant No. 31371047).
