Abstract
a-Stratified computerized adaptive testing with b-blocking (AST), as an alternative to the widely used maximum Fisher information (MFI) item selection method, can effectively balance item pool usage while providing accurate latent trait estimates in computerized adaptive testing (CAT). However, previous comparisons of these methods have treated item parameter estimates as if they are the true population parameter values. Consequently, capitalization on chance may occur. In this article, we examined the performance of the AST method under more realistic conditions where item parameter estimates instead of true parameter values are used in the CAT. Its performance was compared against that of the MFI method when the latter is used in conjunction with Sympson–Hetter or randomesque exposure control. Results indicate that the MFI method, even when combined with exposure control, is susceptible to capitalization on chance. This is particularly true when the calibration sample size is small. On the other hand, AST is more robust to capitalization on chance. Consistent with previous investigations using true item parameter values, AST yields much more balanced item pool usage, with a small loss in the precision of latent trait estimates. The loss is negligible when the test is as long as 40 items.
Keywords
Among the greatest advantages of computerized adaptive testing (CAT) is efficiency: CAT is often able to achieve measurement goals much more quickly than paper-and-pencil (P&P) tests. Simply by matching item difficulty with examinee ability (
However, the MFI method has some disadvantages. By selecting the “best” items in a narrow statistical sense, some items, usually items with high discrimination parameters, are chosen frequently, whereas others are used rarely or not at all. This can lead to several serious problems. For example, the distributions of item exposure rates can be highly skewed, which has important implications for test security and the costs of item development. Moreover, Chang and Ying (2008) demonstrated that when highly discriminating items are administered early in the test, the final
A novel solution to these issues is to employ an a-stratified item pool in CAT, first proposed by Chang and Ying (1999) and later modified in Chang, Qian, and Ying (2001). This latter method is called a-stratified CAT with b-blocking (AST), where a and b are the item discrimination and difficulty parameters, respectively. The method first rank orders items in the pool according to their b parameters. Items with similar b parameters are grouped to form a “b-block,” and the items within each block are rank-ordered based on their a parameters. The low-a stratum is formed by pooling the low-a items from all b-blocks, and the high-a stratum is formed by pooling the high-a items. This way the item pool can be divided into K strata. The strata are ordered by increasing levels of average item discrimination, but the distribution of difficulty is roughly equivalent across strata. Then from the kth stratum (k = 1, 2, . . . , K), mk items are administered by minimizing the distance between
But there is another layer of complexity to the issue: In practice, the true values of item parameters are unknown, and only estimates of the item parameters are available. The optimization scheme in MFI was found to result in an overrepresentation of large calibration errors among the selected items (van der Linden & Glas, 2000). Or as Hambleton, Jones and Rogers (1993) put it, “Among items with generally high discrimination parameters (meaning in terms of the true values), those that are overestimated (i.e., with positive estimation errors) are often the ones selected for a test” (p. 145). This phenomenon is referred to as “capitalization on item calibration error” (van der Linden & Glas, 2000) or “capitalization on chance” (Cheng & Verkuilen, 2008; Olea, Barrada, Abad, Ponsoda, & Cuevas, 2012; Patton, Cheng, Yuan, & Diao, 2013; Veldkamp, 2012). In particular, the MFI method tends to select items with spuriously high discrimination estimates, resulting in spuriously low standard errors (SEs) for
In this study, we would like to examine the performance of AST in the presence of calibration error for the first time. AST uses low-a items in the beginning of the test, and within each stratum, item selection is made with respect to item difficulty alone, so one might expect the effects of capitalization on chance to be mitigated. Additionally,
Thus, the goal of our study is to compare the performance of AST against that based on MFI. In practice, the MFI method is usually coupled with exposure control techniques, for example, the S–H method. Also, the performance of the pure MFI method in the presence of calibration error has been well-studied (van der Linden & Glas, 2000). Therefore, in this study we compare the performance of the AST against two hybrid item selection methods: MFI in conjunction with S–H exposure control and MFI with randomesque exposure control. Various magnitudes of item calibration error are introduced by manipulating the calibration sample size. The dependent measures of interest are
Exposure Control Methods
A simple method of exposure control is called the randomesque method (Kingsbury & Zara, 1989). It is a method that is intuitively easy to understand and straightforward to implement. Rather than selecting the “best” item according to the item selection criterion (such as MFI), this method identifies the d most informative items at the current ability estimate and randomly selects one among them for administration. This way the exposure of the most informative items is reduced, particularly if d is large. Because the most informative items also tend to be those with spuriously large discrimination estimates, this method can be expected to reduce capitalization on chance.
An alternative is the S–H method of exposure control (Sympson & Hetter, 1985). It is “the most popular method of item-exposure control in computerized adaptive testing” (van der Linden, 2003, p. 249). The S–H method conducts a probability experiment to determine whether to administer an item once it is selected. If S
j
represents the selection of item j for a randomly sampled examinee and A
j
represents the administration of the item, the S–H method forces
Method
We used 520 items from the retired item bank of a large-scale achievement test. The goal of this study is to manipulate the amount of calibration error and examine the performance of the AST method in the presence of varying amounts of calibration error. Two IRT models are considered, the two parameter logistic (2PL) model and the three parameter logistic (3PL) model. Descriptive statistics of the parameter values are shown in Table 1; the distribution of item difficulty is roughly symmetric with a mean of 0.16, and the distribution of discrimination has a slight positive skew. Also, difficulty and discrimination have a medium positive correlation (r = .45), which is commonly found in the literature (e.g., Chang et al., 2001). These item parameters served as true parameter values and were used to generate item responses. For the 2PL model, the c parameters are set to 0. The true items are denoted as
Descriptive Statistics of True Item Parameter Values.
The 520 items were calibrated with each of three different sample sizes (N = 2,500, 1,000, or 500) using BILOG-MG-3 (Zimowski, Muraki, Mislevy, & Bock, 2003). Item calibration was performed with both the 2PL and 3PL models given each sample size. So in total there are six conditions (2 models × 3 sample sizes), and for each condition we obtained 50 sets of converged item parameter estimates. For each replication, dichotomous item responses under the 2PL or 3PL model were generated with each of three different sample size with true θs sampled from N(0, 1). For calibration of the 2PL model, no prior is specified. Calibration of the 3PL model typically requires a large sample size (de Ayala, 2009; Jones, Smith, & Talley, 2006); with small sample sizes the estimation may not converge or may lead to unstable estimates of the pseudo-guessing parameter. For calibration of the 3PL, therefore, the default prior (beta(5, 17), with mean of .2) for the c parameters is used. Even with the default prior for the c parameters, not all replications converged for calibration of the 3PL model. It took 55 replications for one condition (N = 500) to reach 50 converged set of item parameter estimates. All others conditions needed only 50 replications.
The average root mean squared error (RMSE) of parameter estimates for each calibration sample size under each IRT model (across replications), computed as the square root of the average mean square error, is shown in Table 2. For a given value of N, a given item parameter (discrimination or difficulty), and a given IRT model, the RMSE was computed (see Equation 3 for the computation of RMSE but replace
Root Mean Squared Error of Item Parameter Estimates (Averaged Across 50 Replications).
Each calibrated bank was then employed in three different CAT scenarios. In each scenario, item selection and examinee scoring were conducted with respect to item parameter estimates, whereas responses were generated using the true parameters of the corresponding items. We also included a baseline condition in which the true item parameters were used for item selection, response generation, and scoring. This corresponds to the hypothetical situation in which the examinee sample size is infinite (i.e.,
CAT Scenarios
In all CAT simulations in this study, item responses were generated based on the true item parameters, but item selection and ability estimation were performed based on the item parameter estimates. The simulated CATs had a fixed length of 15 or 40 items and employed one of three item selection methods: the AST method, the MFI_SH method, and the MFI_R method. For the AST method, we rank-ordered the bank with respect to estimated difficulty and partitioned the bank into 130 b blocks of four items each. The items within each block were then assigned to one of four a-strata, from lowest to highest. In this way, each a-stratum exhibited a similar distribution of item difficulty. Examinees were administered roughly equal numbers of items from each a-stratum, beginning with the a-stratum of lowest discrimination and proceeding to the a-stratum of highest discrimination. So for the 15-item test, 4 items were administered from each of the first three strata, and 3 items were administered from the last. For the 40-item test, 10 items were administered from each stratum. Within a given a-stratum, an item was selected by minimizing the discrepancy between estimated difficulty and the current
For the MFI_SH method, the MFI criterion was paired with S–H exposure control. The maximum exposure rate for every item in the bank was set at r = 0.2. For each of the 50 calibrated banks, exposure control parameters were estimated by administering the CAT to 1,000 simulated examinees with
Regardless of the item selection method, maximum likelihood was used to estimate ability. An initial ability estimate was obtained as follows: the item bank (or the first a-stratum for AST) was ordered with respect to estimated difficulty; the 20 easiest items, 20 most difficult items, and 20 items of medium difficulty were identified; and one item was randomly selected for administration from each group. For those examinees who responded incorrectly or correctly to all three items, a provisional ability estimate of −4 or 4 was used, respectively, until a finite
Dependent Measures
Latent Trait Recovery
To obtain results conditional on specific values of
and the 50 bias estimates were averaged. To obtain the empirical standard error, we first computed the sample variance of ability estimates for the 200 examinees with a given true
where
To obtain results that reflect a more realistic population of examinees, an additional scoring sample consisted of 1,000 examinees with
and the square root of the average of the 50 MSEs was computed. Pearson correlation coefficients between
Capitalization on Chance
We introduce relative test efficiency as a measure of the degree to which capitalization on chance occurred. Test efficiency (TE) is defined as the average value of (estimated) test information:
where
If no capitalization on chance occurs, then we expect RTE to be close to 1.0. When RTE deviates substantially from 1, it suggests that capitalization on chance occurs. In the context of CAT, it usually means capitalizing on spuriously large discrimination parameter estimates, which in turn leads to an RTE that is substantially larger than 1. This was computed using results from the “conditional” scoring sample. That is, we computed RTE within each group of 200 examinees at a given true
Item Pool Usage
To examine item pool usage, data from the “marginal” scoring sample were used. To quantify item pool usage, we used a χ2 statistic proposed by Chang and Ying (1999), which measures the variability of item exposure rates:
where M is the total number of items in the pool (here M = 520),
Results
Latent Trait Recovery
Conditional Results
Figure 1 displays the conditional bias of

Conditional bias of ability estimates for each test length and item selection method (2PLM, conditional results).
In the bottom row of Figure 1, where test length = 40, however, the AST method catches up quickly. The bias under the AST method does not change much with the calibration sample size N, showing that when the test is long enough, the AST method is robust against small calibration sample size. For the MFI-based methods, on the other hand, the reduction in bias from lengthening the test is not as substantial. With the MFI method, even with a test of 40 items, the separation among different sample sizes is still visible at the low and high ends of
Next, Figure 2 displays the standard error of

Empirical standard error of ability estimates for each test length and item selection method (2PLM, conditional results).
Similar to Figures 1 and 2, Figures 3 and 4 presents the bias and standard error of

Conditional bias of ability estimates for each test length and item selection method (3PLM, conditional results).

Empirical standard error of ability estimates for each test length and item selection method (3PLM, conditional results).
The SE curves in Figure 4 compared to those in Figure 2 are all higher. The separation among the curves from different sample sizes is more conspicuous with the 3PL model. In Figure 2, there is no clear separation when the test length is 40. In Figure 4, the separation is still evident at the low range of
Marginal Results
For the results based on the “marginal” scoring sample, the RMSE of
Ability Recovery for Each Test Length and Item Selection Method (2PLM, Marginal Results).
Similarly, the RMSE of
Ability Recovery for Each Test Length and Item Selection Method (3PLM, Marginal Results).
Capitalization on Chance
Figures 5 and 6 display relative test efficiency for each combination of test length and item selection method under the 2PL and 3PL models, respectively. Again, relative efficiency is the ratio of test information computed from estimated item parameters over that from the true item parameter values. Under MFI_SH (as well as under MFI_R, though these results are not shown as they are essentially very similar to results of MFI_SH), Figure 5 shows that, as expected, test information based on item parameter estimates is spuriously high relative to that based on the true item parameter values—this confirms that capitalization on chance indeed occurred. Furthermore, this overestimation of test information increases as the size of the calibration sample decreases, suggesting that capitalization on chance is more severe when the calibration sample is smaller. This makes sense because when the calibration sample size is small, the resulting item parameter estimates will contain larger errors. In particular, the most informative items (when evaluated with item parameter estimates) have large discrimination estimates that tend to contain positive calibration errors; the size of these errors tends to be larger for smaller samples. As the inverse of test information is used to approximate the variance of

Relative efficiency for each test length and item selection method (2PLM, conditional results).

Relative efficiency for each test length and item selection method (3PLM, conditional results).
Item Pool Usage
Last, item pool usage results under the 2PL model are shown in Table 5 for each combination of test length and item selection method. As expected, for a given test length, the AST method achieves more uniform item exposure rates than both MFI-based methods. The advantage is substantial. The effect of calibration sample size on item pool usage, on the other hand, is negligible. Between the two MFI-based methods, the difference is small but an interesting pattern emerges. When the test is short, item pool usage is the worse with the MFI_SH method, indicated by the larger
Item Pool Usage for Each Test Length and Item Selection Method (2PLM, Marginal Results).
Note. Underexposed items have exposure rates <.02, and over-exposed items have exposure rates >.20.
By breaking down the unbalanced item pool usage into overexposure and underexposure, we can easily see that for the MFI-based methods, underexposure is a big issue. The advantage of the AST method is more evident when the test is longer. When the test length is 40, only about 1% of the items are overexposed and only about 1% of the items are underexposed under AST, while under the MFI-based methods, about 14% to 15% of items are overexposed and a staggering 50% to 58% of items are underexposed. This means that over half of the item pool is largely dormant when the MFI-based methods are used. These findings concerning item pool usage are consistent with previous studies (e.g., Yi, 2002) that used true item parameters instead of item parameter estimates.
Similarly, results of item pool usage under the 3PL model are displayed in Table 6 for each combination of test length and item selection method. The same trends as discussed above in Table 5 are observed.
Item Pool Usage for Each Test Length and Item Selection Method (3PLM, Marginal Results).
Note. Underexposed items have exposure rates <.02, and overexposed items have exposure rates >.20.
Discussion
Though our statistical training tells us that larger samples are always better, for the calibration of items in some operational testing programs this is simply not an option. Indeed, because the costs associated with adding new items to an existing bank can be quite high, many test developers cut costs by limiting the size of the calibration sample (van der Linden & Pashley, 2010). Thus, it is important to study the effects of calibration sample size on measurement outcomes and to identify test designs that are robust to calibration error. In particular, our goal was to contrast the a-stratification method against item selection methods based on the principle of maximizing Fisher information. It has been shown that the MFI-based methods are prone to capitalization on chance (van der Linden & Glas, 2000), meaning that items with high discrimination parameter estimates, which tend to contain positive errors, are favored. Our simulations demonstrate that MFI, even when combined with exposure control techniques such as Sympson–Hetter or randomesque, still tends to capitalize on spuriously large discrimination estimates, leading to spuriously high test information. This effect is exacerbated when the calibration sample size decreases. In contrast, the a-stratification method leads to “observed” test information that is much closer to the “true” test information, and the method is quite robust to changes in the calibration sample size. This is one of the most important findings here. Additionally, item pool usage under the a-stratification method is much better than that under the MFI-based methods across all conditions.
Previous studies have demonstrated that with the a-stratification method, ability recovery tends to suffer a bit relative to the MFI method when item parameters are known. The results here demonstrate that this is also true when the item parameters are estimated, especially when the test is short. When the test gets longer, the a-stratification method quickly catches up. With a test of 40 items, the a-stratification method is only slightly worse in terms of measurement precision (on multiple indices the difference is in the second decimal place) but leads to much better item pool usage. Taken together, we find that the a-stratification method is largely robust to capitalization on chance, and the effect of calibration sample size on its measurement precision and exposure control is very small. With a 40-item test, the loss of measurement precision from the a-stratification method is negligible but its gain in balancing item pool usage is substantial.
The discussion above applies to both the 2PL and 3PL models. Between these two models, the latter is more susceptible to capitalization on chance, and the effect of calibration sample size is more conspicuous. When the test is long, regardless of the IRT model, the a-stratification method manages to achieve much better item pool usage and reduce susceptibility to small calibration sample size, while reaching essentially equivalent level of measurement precision.
It is important to note that the generalizability of our results is somewhat limited. Operational CAT programs often impose nonstatistical constraints on item selection such as content balancing. Content balancing may reduce the risk of capitalization on chance by considering only a subset of the item bank (i.e., those items that correspond to content areas not yet covered by the test). Note that the a-stratification method has been extended successfully to incorporate content balancing (e.g., Cheng, Chang, & Yi, 2007; Yi & Chang, 2003). Therefore, it is definitely possible and worthwhile to compare the a-stratification method when combined with content balancing against the conventional MFI-based algorithms with content control. In addition, there are also many other exposure control techniques available, including the restrictive maximum information method which freezes items that are overexposed (Revuelta & Ponsoda, 1998), the progressive method which adds a random component that receives progressively less weight than the usual information component (Revuelta & Ponsoda, 1998), and the flexible a-stratified method where increasing numbers of items are selected from strata with higher a parameters (Deng, Ansley, & Chang, 2010). Many hybrid strategies have also been proposed, such as the progressive restricted strategy (Revuelta & Ponsoda, 1998), the combination of a-stratification and S–H strategy (Leung, Chang, & Hau, 2002), and a-stratification with freezing (Parshall, Harmes, & Kromrey, 2000). For a review of exposure control strategies developed up to the year 2005, please see Georgiadou, Triantafillou, and Economides (2007). Since 2005 many new methods have been proposed, for example, the online S–H procedure, where exposure parameters are updated on the fly rather than being obtained through iterative simulations beforehand (Chen, Lei, & Liao, 2008), and the random component method (Barrada, Olea, Ponsoda, & Abad, 2008), which subsumes the progressive method and the proportional method (Segall, 2004) as special cases. It is beyond the scope of the current study to review and compare all available exposure control strategies. The AST method was chosen as the focus of this study because of its pre-alignment and stratification of items based on the error-prone discrimination parameter estimates. It is of interest in future studies to follow-up with newly developed exposure control strategies and evaluate their performance in the presence of calibration error.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: The project is supported by the 2010 Innovation R&D Grant from CTB/McGraw-Hill.
