Abstract
Variable-length computerized adaptive testing (VL-CAT) allows both items and test length to be “tailored” to examinees, thereby achieving the measurement goal (e.g., scoring precision or classification) with as few items as possible. Several popular test termination rules depend on the standard error of the ability estimate, which in turn depends on the item parameter values. However, items are chosen on the basis of their parameter estimates, and capitalization on chance may occur. In this article, the authors investigated the effects of capitalization on chance on test length and classification accuracy in several VL-CAT simulations. The results confirm that capitalization on chance occurs in VL-CAT and has complex effects on test length, ability estimation, and classification accuracy. These results have important implications for the design and implementation of VL-CATs.
Keywords
The fundamental goal of computerized adaptive testing (CAT) is to match items with examinee ability using an intelligent item selection algorithm. By tailoring the test to the examinee, this procedure yields tests that are more efficient than traditional paper-and-pencil tests. Another layer of test adaptation is possible by using a flexible termination rule that allows the test length to vary across examinees. A test that is adaptive in terms of both item selection and test length is often referred to as a variable-length CAT (VL-CAT). Although many termination rules have been proposed in the literature, the choice of rule depends on the purpose of the test. For example, if the goal is to achieve uniform measurement precision for all examinees, the test may end when the standard error (SE) of the ability estimate is sufficiently small (Thissen & Mislevy, 2000). If instead the goal is to classify examinees, the test may end when the confidence interval (CI) for ability does not include a predetermined cut point (Thompson, 2007).
In the framework of item response theory (IRT), item selection, ability estimation, and test termination all ultimately depend on the values of the item parameters. However, only estimates of the item parameters are available in practice, and because item selection involves optimization, capitalization on chance may occur. This phenomenon has been demonstrated in the context of fixed-length CAT (Olea, Barrada, Abad, Ponsoda, & Cuevas, 2012; van der Linden & Glas, 2000). Specifically, adaptive item selection criteria tend to choose items with spuriously large discrimination estimates, which in turn yield spuriously large values of test information and spuriously low SEs of ability estimates. Because flexible termination rules often depend on the SE of the ability estimate, capitalization on chance may also influence test length and classification accuracy in VL-CAT. Thus, the goal of this study is to examine the effects of capitalization on chance on VL-CAT outcomes.
The rest of the article is organized as follows. First, the authors describe popular termination rules used in VL-CAT. Next, they review literature on the effects of item calibration error on adaptive item selection and ability estimation. Finally, they present the results of a simulation to demonstrate the effects of capitalization on chance on VL-CAT, followed by a discussion of practical implications.
Termination Rules in VL-CAT
As mentioned previously, several termination rules have been proposed for VL-CAT, but the choice of rule depends on the purpose of the test. To achieve uniform measurement precision for all examinees, a straightforward rule is to end the test when the SE of the ability estimate (denoted by
Regardless of the purpose of testing, each of these termination rules depends on the SE of the ability estimate. An important question is whether capitalization on chance might affect SEs, and whether the SEs might influence VL-CAT outcomes, such as test length and classification accuracy. To explore this question, the authors next discuss adaptive item selection and the problem of capitalization on chance.
Item Selection and Capitalization on Chance
Central to the goal of CAT is an item selection criterion that matches examinee ability with a function of the item parameters. Perhaps the most popular criterion is to maximize Fisher information at the current ability estimate. If the probability of a correct response is modeled by the two-parameter logistic model (2PLM),
where aj and bj are item discrimination and difficulty, respectively, Fisher information is computed by
When |bj − θ| is small, Ij is increasing in aj for the range of ajvalues encountered in practice. So, if there are many items located near θ, this criterion tends to prefer items with the largest discrimination values. Under the three-parameter logistic model (3PLM),
where cj is the pseudo-guessing parameter, Fisher information is computed by
Unlike the 2PLM, information under the 3PLM is no longer maximized at θ = bj. However, information is still increasing in aj (for the range of aj values encountered in practice) for a fixed value of cj when bj is close to θ, and therefore also tends to prefer items with the largest discrimination values (van der Linden & Glas, 2000).
However, the true item parameter values are unknown in practice, so items are chosen on the basis of their parameter estimates, and capitalization on chance may occur. Specifically, the largest a estimates in an item pool tend to be spuriously large: the sum of large true a values and large, positive calibration errors (van der Linden & Glas, 2000). Because Fisher information is largely determined by the value of item discrimination, the maximum value of information evaluated with respect to item parameter estimates tends to be larger than that evaluated with respect to the true parameter values (Hambleton & Jones, 1994; Hambleton, Jones, & Rogers, 1993). This problem of capitalization on chance is exacerbated when item calibration errors tend to be large, for example, when the ratio of calibration sample size to the number of model parameters is small (Hambleton & Jones, 1994). The problem also depends on the ratio of test length to the conditional size of the item pool; when there are fewer items to choose from, it is less likely to systematically choose those items with spuriously large discrimination estimates (van der Linden & Glas, 2000). In addition to its effect on Fisher information, capitalization on chance has important practical effects on ability estimation, which the authors discuss in the following section.
Calibration Error and Latent Trait Estimation
In treatments of ability estimation, it is customary to assume that the item parameter values are known (e.g., Baker & Kim, 2004; Lord, 1983). Under this assumption, the SE of
which incorporates an additional term that depends on Σ, the asymptotic covariance matrix of item parameter estimates, and
In the context of CAT with unknown item parameters, estimation of SE(
Capitalization on chance also has implications for the bias of
Purpose of the Study
To our knowledge, two studies have examined the effects of capitalization on chance in fixed-length CAT (viz., Olea et al., 2012; van der Linden & Glas, 2000). However, the situation becomes more complicated when a variable-length termination rule is used. For example, if the CSE rule is used and the SE of
Thus, the goal of this study is to examine the effects of capitalization on chance on θ recovery, test length, and classification accuracy in simulated VL-CAT scenarios. In particular, the CSE and ACI rules are used because of their simplicity and popularity. To examine the effects of different magnitudes of item calibration error, the authors manipulated the size of the calibration sample, and because the degree of capitalization on chance also depends on model complexity, they also conducted simulations with the 2PLM or 3PLM as the true model.
Method
Pools of “true” 2PLM and 3PLM parameters for 400 items were obtained from a retired item pool of a large-scale achievement test. Each pool was calibrated using sample sizes of 2,500, 1,000, or 500, yielding three sets of item parameter estimates that differ with respect to the average magnitude of calibration errors. Each calibrated pool was then “administered” to a future scoring sample in two VL-CAT scenarios that differed with respect to the termination rule (CSE or ACI). To mimic reality, item selection and examinee scoring were conducted with respect to item parameter estimates, whereas responses were generated using the true parameters of the corresponding items. The authors also included a baseline condition in which the true item parameters were used for item selection, response generation, and examinee scoring. This corresponds to the hypothetical situation in which the calibration sample size is infinite (i.e., N = ∞), and the item parameter values are known exactly.
Construction of Item Pools
To create the pools of true item parameters, 400 items were sampled without replacement from a retired pool of 540 items previously calibrated with the 3PLM. These items comprised the 3PLM pool, and the 2PLM pool was created by setting the c parameters to zero. Note that because the 2PLM pool was created in this way, the authors changed not only the model but also the pool information function. Thus, results from the two pools are not directly comparable. Descriptive statistics of the parameter values are shown in Table 1, and histograms of the parameters are shown in Figure 1. The distribution of item difficulty is roughly symmetric with a mean of 0.17, and the distributions of discrimination and pseudo-guessing parameters have a slight positive skew. Discrimination and difficulty exhibit a medium positive correlation (ra.b = .48), whereas discrimination and difficulty each exhibit a small negative correlation with the guessing parameter (ra,c = −.15 and rb,c = −.18, respectively).
Descriptive Statistics of True Item Parameter Values

Histograms of true item parameter values
Next, each pool was calibrated with sample sizes of N = 2,500, 1,000, or 500. Rather than generate response data and estimate the item parameters, parameter estimates for a given item were drawn from their asymptotic normal sampling distribution, a method used in similar investigations (Spray & Reckase, 1987; van der Linden & Glas, 2000). This was done to avoid convergence problems during item calibration, which would almost surely occur during the calibration of 400 items with only 500 examinees. First, the authors computed the asymptotic covariance matrix Σ of ML item parameter estimates assuming a standard normal distribution of ability (see Thissen & Wainer, 1982). This method does not require observed data and uses Fisher information so that Σ is block diagonal with all inter-item covariances equal to zero. Second, to simulate the calibration of a given item, a random sample was drawn from a two- or three-variate normal distribution with mean vector
Item Parameter Recovery (MSE)
Note: MSE = mean squared error; 2PLM = two-parameter logistic model; 3PLM = three-parameter logistic model.
The MSE of 3PLM parameter estimates is based on 399 of the 400 items; the easiest item in the pool exhibited very poor recovery.
VL-CAT Scenarios
Regardless of termination rule, the purpose of each VL-CAT simulation was to score examinees and classify them as passing or failing. (Although the main purpose of the CSE rule is to obtain accurate ability estimates, the authors also classified examinees to obtain a measure of the practical consequences of calibration error on ability estimation.) For accurate classifications, it is desirable for the peak of information in the item pool to be located near the cut point. On inspection of the plots of information in the 2PLM and 3PLM pools (evaluated with respect to the true parameters), the cut was chosen to be at θ = 0.5. However, they also conducted simulations with the cut located at θ = 1.5; this corresponds to a situation in which the goal is to identify high-ability examinees. Regardless of cut location, the classification decision was made by comparing the final ability estimate with the cut point.
For all simulations, ML was used to estimate ability. An initial ability estimate was obtained as follows: The item pool was ordered with respect to estimated difficulty; the 20 easiest items, 20 most difficult items, and 20 “middle” items were identified; and 1 item was randomly selected for administration from each group. For those examinees who responded incorrectly or correctly to all 3 items, the next item was chosen to maximize Fisher information at θ = −4 or 4, respectively, until a finite
Under the CSE rule, the SE of
to show that σ
e
= .316 corresponds to a reliability coefficient (
Together, this resulted in four fully crossed factors: IRT model (2PLM or 3PLM), termination rule (CSE or ACI), cut location (θ0 = 0.5 or 1.5), and calibration sample size (N = ∞, 2,500, 1,000, or 500).
Dependent Measures
The scoring sample consisted of 500 examinees at each θ value from −2 to 2 in increments of 0.25. Each of the 25 calibrated pools was administered to a separate scoring sample, and the dependent measures were averaged across the 25 replications. In this way, the outcomes at each θ value reflect two sources of error: random responses to items (i.e., measurement error) and capitalization on calibration error via item selection. To measure the degree to which capitalization on chance occurred, the authors computed relative test efficiency. First, they define test efficiency as the average information per item administered, computed as follows for variable-length tests (Huo, 2009):
where
The authors expect this ratio to be greater than one and to increase as the calibration sample size N decreases.
To evaluate ability recovery, the authors computed the bias and standard deviation of ability estimates at each true θ value. To evaluate the effect of capitalization on chance on test outcomes, they also computed the average test length and percentage of accurately classified examinees at each θ value. Finally, all data generation, simulations, and analyses were performed in R (R Core Team, 2012) with codes written by the first author.
Results
Capitalization on Chance
First, the authors wanted to determine whether the maximum information criterion capitalized on item calibration error. For those simulations using the CSE termination rule, the top row of Figure 2 displays plots of relative efficiency for each combination of IRT model and calibration sample size. (Although not shown, these plots look very similar under the ACI rule.) In all conditions, relative efficiency is greater than one, demonstrating that when items were chosen on the basis of their parameter estimates, test information evaluated with respect to item parameter estimates is (on average) greater than that evaluated with respect to the true parameter values. Moreover, the discrepancy between estimated and true test efficiency increases as N decreases, which reflects that the smaller calibration samples yielded larger calibration errors. As expected, values of estimated test efficiency are spuriously high due to capitalization on spuriously high discrimination estimates, but this is not always the case. Under the 3PLM, examinees with θ < −1 were administered items with underestimated difficulty. This occurred because there are few very easy items in the pools (see Figure 1), but after calibration, low-ability examinees were supplied additional items with difficulty estimates containing negative calibration errors.

Test efficiency (CSE termination rule)
Next, the authors wanted to compare estimated test efficiency across the different sample sizes; these results are shown in the bottom row of Figure 2. (Note that there is no distinction between “estimated” and true efficiency in the baseline condition.) Regardless of N, estimated efficiency is greater for positive θ values because this is where the most discriminating items are located (recall that the true a and b values are positively correlated). At a given value of θ, estimated efficiency tends to increase as N decreases; this means that the items chosen in the N = 500 condition, for example, appear to be more informative than the items chosen in the baseline condition. This implies, in turn, that the SEs of ability estimates in the N = 500 condition are spuriously small.
CSE Termination Rule
When the CSE termination rule is used, what are the implications of capitalization on chance for test length? The top row of Figure 3 displays the average test length for each combination of IRT model and calibration sample size. As expected, spuriously low SEs under small N caused the test to end prematurely, and consistent with the plots of estimated test efficiency (see the bottom row of Figure 2), the effect of N on test length is proportional to the degree of capitalization on chance. This is apparent in the bottom row of Figure 3, which displays the difference in average test length relative to the baseline condition (e.g., test length for N = 500 minus test length for N = ∞). Under the 2PLM, the effect of N is relatively uniform across the range of ability. In contrast, the effect of N under the 3PLM is quite large for θ > 0 but negligible for θ < 0.

Average test length (CSE termination rule)
Next, the authors examined the effect of capitalization on chance on ability recovery. The top row of Figure 4 displays bias under each IRT model, and the bottom row displays the empirical SE (i.e., the standard deviation of ability estimates). As expected, ability estimates in the baseline condition are close to unbiased, regardless of model. But as N decreases, the magnitude of bias tends to increase. Concerning the empirical SEs, the effect of N under the 2PLM is negligible, and under the 3PLM, the empirical SE increases as N decreases. These results are at odds with the information-based SEs, which suggest that ability estimates under small N are more accurate than those under large N. But in reality, ability estimates under small N are equally or even less accurate than those under large N.

Ability recovery (CSE termination rule)
Finally, the authors were interested in the effect of capitalization on chance on classification accuracy. For those conditions with θ0 = 0.5, the left half of Table 3 displays percentages of correctly classified examinees. Because nearly all examinees with θ values far from the cut were correctly classified, Table 3 focuses on a smaller range of θ values about the cut. Regardless of IRT model, the effect of N is quite small. Because the classification decision is made by comparing the final ability estimate with the cut, these results can be explained in terms of ability recovery. Specifically, the effect of N on the bias and empirical SE of
Percentages of Correctly Classified Examinees (CSE termination rule)
Note: CSE = conditional standard error; 2PLM = two-parameter logistic model; 3PLM = three-parameter logistic model.
Difference is computed by subtracting N = ∞ results from N = 500 results.
ACI Termination Rule
Under the ACI termination rule, the effect of capitalization on chance on test length exhibits a different trend. For θ0 = 0.5, the top row of Figure 5 displays the average test length for each combination of IRT model and calibration sample size, and the bottom row displays the difference in average test length, relative to the baseline condition. In contrast with the CSE rule, test length depends on the cut location; accordingly, tests are quite long near the cut but much shorter at extreme θ values. Moreover, the effect of N at a given θ value is quite small. This result was unexpected; the authors know that when N is small, the SE of

Average test length (ACI termination rule, θ0 = 0.5)
Next, Figure 6 displays ability recovery results for θ0 = 0.5; the top row displays bias under each IRT model, and the bottom row displays the empirical SE. Similar to the CSE results, the effect of N on the empirical SE is negligible under the 2PLM, whereas under the 3PLM, the empirical SE is inversely related to N. For all values of N,

Ability recovery (ACI termination rule, θ0 = 0.5)
Finally, the left half of Table 4 displays percentages of correctly classified examinees when θ0 = 0.5. Similar to the CSE results, the effect of N is quite small, regardless of IRT model. Again, this is because the effect of N on the bias and empirical SE of
Percentages of Correctly Classified Examinees (ACI termination rule)
Note: ACI = ability confidence interval; 2PLM = two-parameter logistic model; 3PLM = three-parameter logistic model.
Difference is computed by subtracting N = ∞ results from N = 500 results.
Discussion
The goal of this study was to examine the effects of capitalization on item calibration error in VL-CAT. Consistent with previous research on fixed-length CAT, the authors found that the maximum information criterion capitalized on item calibration error, yielding spuriously high values of test information, particularly when the calibration sample was small. Unsurprisingly, this resulted in spuriously short tests when the SE of the ability estimate was used as the test termination criterion. Relative to when the item parameters were known, the average reduction in test length was as much as three items. In contrast, using the 95% CI for θ as the termination criterion appeared to safeguard against spuriously short tests, particularly when the cut was located at θ = 0.5, where the test information function peaks. Although small calibration samples yielded spuriously narrow CIs, variability in the location of the intervals prevented tests from ending prematurely. Regardless of termination rule, the effect of calibration sample size on classification accuracy was quite small when the cut was located at θ = 0.5. This occurred because the empirical bias and variance of ability estimates for examinees near the cut were similar for all values of N. However, classification accuracy under small N was reduced by as much as 10% relative to the baseline condition when a more extreme cut point was used (i.e., θ0 = 1.5). This larger effect was primarily due to the effect of N on the bias of ability estimates.
Although the authors have demonstrated the potential effects of capitalization on chance, they have not addressed ways to reduce its effects. One option is to use cross-validation to produce two sets of item parameter estimates. In this method, items are selected with respect to one set of estimates, and the other set is used for ability and SE estimation. van der Linden and Glas (2001) found that this method effectively reduced the effects of capitalization on chance, though they note that this solution may not be ideal when the calibration sample is already small (e.g., 250 examinees or less). Also by imposing constraints on item selection, such as content constraints and exposure control, which are already implemented by many operational testing programs, the effect of capitalization on chance may be reduced. The reason is that these constraints reduce the dependence of item selection on purely statistical criteria. For example, van der Linden and Glas (2000) demonstrated that implementing Sympson–Hetter exposure control in fixed-length CAT decreased the effects of capitalization on chance relative to when no exposure control was used.
It should be emphasized that the results are particular to the simulated conditions. The effects of capitalization on chance on VL-CAT outcomes depend on many factors that vary widely among operational testing programs. First, maximum Fisher information is but one of many possible item selection criteria (see van der Linden & Pashley, 2010). However, some research has suggested that different criteria, though conceptually different, tend to select very similar sets of items (e.g., Chen, Ankenmann, & Chang, 2000). Furthermore, van der Linden and Glas (2000) demonstrated that the effects of capitalization on chance in fixed-length CAT were very similar for several different item selection criteria. Second, there are a number of variable-length termination rules that the authors did not implement. Two attractive alternatives are the minimum information and predicted SE reduction rules (Choi et al., 2011). Whether or not a target SE for
In summary, the authors have demonstrated that capitalization on chance via item selection can have practically important consequences on VL-CAT outcomes. Specifically, if the SE of the ability estimate is used as the stopping criterion, the test may end prematurely when calibration errors are large. Although the goal of uniform measurement precision is a worthy one, this procedure may result in spuriously short tests and ability estimates that are less accurate than their SEs suggest. The current results also suggest that capitalization on chance may have large effects on classification accuracy when the cut point is located at an extreme θ value, regardless of the termination rule. An extreme cut point may be used to, for example, identify the highest or lowest performing examinees in a sample. Making classification decisions at a point where items are scarce is problematic when the item parameters are known. When only item parameter estimates are available, the accuracy of these classifications may suffer even more.
Footnotes
Acknowledgements
The authors would like to thank CTB/McGraw-Hill for supporting this research, as well as two anonymous reviewers for their insightful comments.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This research was supported by a 2010 CTB/McGraw-Hill Innovation Research and Development grant.
