Abstract
In this article, the effect of the upper and lower asymptotes in item response theory models on computerized adaptive testing is shown analytically. This is done by deriving the step size between adjacent latent trait estimates under the four-parameter logistic model (4PLM) and two models it subsumes, the usual three-parameter logistic model (3PLM) and the 3PLM with upper asymptote (3PLMU). The authors show analytically that the large effect of the discrimination parameter on the step size holds true for the 4PLM and the two models it subsumes under both the maximum information method and the b-matching method for item selection. Furthermore, the lower asymptote helps reduce the positive bias of ability estimates associated with early guessing, and the upper asymptote helps reduce the negative bias induced by early slipping. Relative step size between modeling versus not modeling the upper or lower asymptote under the maximum Fisher information method (MI) and the b-matching method is also derived. It is also shown analytically why the gain from early guessing is smaller than the loss from early slipping when the lower asymptote is modeled, and vice versa when the upper asymptote is modeled. The benefit to loss ratio is quantified under both the MI and the b-matching method. Implications of the analytical results are discussed.
Keywords
When the items are scored dichotomously, such as right/wrong, dichotomous item response theory (IRT) models are often used to model the relationship between the underlying latent trait θ and a respondent’s score. According to the widely used three-parameter logistic model (3PLM; Birnbaum, 1968), the probability for an examinee with latent trait level θ to receive score uj = 1 on the jth item (uj = 1 for a correct response and 0 for an incorrect one) is
where the parameters
As early as 1981, Barton and Lord discussed adding an upper asymptote, d, to the 3PLM to accommodate unfortunate slips of highly capable examinees, similar to the lower asymptote accounting for lucky guessing from low-ability examinees. The resulting model is
A more general model, the four-parameter logistic model (4PLM), further allows the upper asymptotes to differ across items
For item j, the item parameter vector is
Chang and Ying (2008) provided important theoretical results on how the step size between adjacent latent trait estimates is influenced by n, the number of items administered, and the discrimination parameter a. Their analytical results show that when highly discriminating items are used early in the test (i.e., when n is still small), the one-step adjustment or step size can be very large. Therefore, they projected that
if the examinee misses a number of initial items and the test length is short to moderate, then he or she may not be able to regain a score comparable to the true θ, even though he or she responds well to the rest of the items (p. 444).
Their simulation study further supported this assertion and showed that the a-stratify method, by using low-a items early in the test, offers a more robust algorithm for item selection in computerized adaptive testing (CAT).
The objective of this article is to extend the analytical results in Chang and Ying (2008). The extension happens in three ways. First, Chang and Ying focused on the role of the discrimination parameter in step size. In this article, the role of the upper and lower asymptotes is examined. The authors show analytically that these asymptotes can reduce the step size between adjacent interim ability estimates in CAT. Previous simulation studies (e.g., Rulison & Loken, 2009) alluded to the effects of asymptote parameters but analytical results are warranted. Second, the results of the present study are no longer limited to the maximum information method for item selection. General results, as well as results specifically under the maximum information method and the b-matching method, are obtained. Third, relative step size between modeling versus not modeling the asymptotes is derived. The authors also explain analytically the asymmetry, that is, larger gain from early guessing than loss from early slipping when the upper asymptote is included in the model, and vice versa when the lower asymptote is included in the model.
The rest of the article is organized as follows. First, the authors discuss technical aspects of CAT under the 4PLM, the model with both lower and upper asymptotes. The discussion focuses primarily on ability update and item selection, which are instrumental to the step size. Then, the authors derive the step size between successive ability updates under the 4PLM and the two models it subsumes, the model with the lower asymptote only (3PLM), and the model with upper asymptote only (3PLMU). The step size and relative step size are derived under both the maximum information method and the b-matching method. Implications of the analytical results are discussed at the end.
CAT Based on the 4PLM
Barton and Lord (1981) fit Model 2 to data of large-scale aptitude tests, and the conclusion was that the change in ability estimates was negligible. Since then, the discussion of adding an upper asymptote to the IRT model was rare, until very recently when there seems to be revived interest in this topic. There are three main reasons for such rekindled interest. First, the upper asymptote was found valuable in fitting psychological testing data (Loken & Rulison, 2010; Reise & Waller, 2003; Waller & Reise, 2009). Second, failing to account for unfortunate slips early in a test due to carelessness or warm-up effects may lead to negative consequences in CAT (Rulison & Loken, 2009; Yen, Ho, Laio, Chen, & Kuo, 2012). Last but not the least, in Barton and Lord (1981), the upper asymptote was fixed across items. Recent advances in computational statistics have made it straightforward to estimate an upper asymptote for each item (Loken & Rulison, 2010), that is, Model 3.
Liao, Ho, Yen, and Cheng (2012) and Yen et al. (2012) discussed computerized adaptive testing that is built upon the Barton and Lord (1981) model, where the upper asymptote is fixed across items. Below, it has been shown that traditional item selection and ability estimation methods can be easily modified to be used in CAT built on the more general 4PLM in Model 3.
There are four important components to a CAT (Cheng & Keng, 2009): (a) the start rule, that is, what item or items to give to examinees to start the test with. Oftentimes, the first items are chosen randomly in the bank or in a certain difficulty range of the bank; (b) the ability estimation method that updates
Ability Estimation
Assuming local independence, the likelihood function given a person’s response vector
where
Obtaining the
The Newton–Raphson iterative procedure can be used to solve the above equation, which involves the first- and second-order derivative of the log likelihood (Embretson & Reise, 2000). The first- and second-order derivatives under the 4PLM are given in Online Appendix A.
A popular alternative is the Expected a Posterior (EAP) estimate,
Computationally, the integration is transformed into a summation of weighted likelihoods at quadrature points:
where sk is one of the quadrature points, nq is the number of quadrature points, w(sk) is the weight associated with the quadrature point under the prior distribution, and
Item Selection
Once the ability estimate is updated, the CAT program selects among eligible items in the bank the most suitable one as the next item to administer. The most widely known item selection method in CAT is the maximum Fisher information method (MI; Lord, 1980). Eligible items in the bank are rank ordered according to their item information at the most recent
It is straightforward to compute the item information at
The MI method under the 2PLM or the 3PLM has been criticized for favoring highly discriminating items, because the information function includes the term
Another criticism of the MI method is related to early slipping in CAT. Chang and Ying (2008) showed mathematically that the adjustment step size between
As kindly pointed out by a reviewer, given the well-known heavy reliance on highly discriminating items of the MI method, other item selection algorithms should be considered. Here, the case where the (t+ 1)th item is chosen through b-matching, that is, the match-ability-with-difficulty method (Hulin, Drasgow, & Parsons, 1983), is considered. Note that such strategy ignores the discrimination parameter in item selection and therefore does not favor highly discriminating items.
Rulison and Loken (2009) argued that the 4PLM helps reduce the negative bias caused by early slipping in CAT, and demonstrated it using simulation studies given the constrained 4PLM of Barton and Lord (1981) under the MI method. The goal of this study is to show the effects of the asymptotes analytically under both the MI and the b-matching method.
Theoretical Results on Step Size Between Successive Interim Ability Estimates
Chang and Ying (2008) mathematically derived the update step size between
According to Chang and Ying (2008), the step size between
where a(t+ 1) is the discrimination parameter of the (t+ 1)th item. θ* is a point between
Similar to Chang and Ying (2008), it can be shown that the step size under the 4PLM can be approximated by
where
where
To show that
Therefore, the denominator (z1+z2+z3) > 0. When the numerator and denominator are both positive, the ratio must be positive. So
Next, it is shown that
It is clear from Equation 7 that the size of the discrimination parameter still has large influence on the early step sizes in CAT under the 4PLM. Meanwhile, the step sizes are moderated by
Below, the effect of c and d in isolation has been examined. When the model only includes the lower asymptote c, that is, when d = 1, the 4PLM reduces to the regular 3PLM. When the model only includes the upper asymptote d, that is, when c = 0, the 4PLM reduces to the following:
This model is referred to as 3PLMU, that is, the three-parameter model with an upper asymptote.
Effect of the Lower Asymptote on Step Size
To isolate the effect of the lower asymptote, the step size under the 3PLM, which is given in Equation 6, is examined. Chang and Ying (2008) showed that
Let
Clearly,
(Relative) step size under the maximum information method under the 3PLM
Assume that t items have been administered and the current ability estimate is
and the probability of answering the the (t+ 1)th item correctly is expected to be
Following Equation 6, the step size between
where the subscript MI indicates that the item is selected under the MI method. In general, the items preferred under the MI method have large discrimination and small guessing parameters. In other words,
Other things being equal, what is the effect by simply including the lower asymptote on the (t+ 1)th item? To find out the effect, a is fixed and the effect of c is isolated. Note that all t items that have been administered are the same. The difference only lies in the (t+ 1)th item, whether the

Relative step size between modeling versus not modeling the lower (left panel) or the upper (right panel) asymptote.
(Relative) step size under b-matching under the 3PLM
If the item pool is rich enough,
and the subscript BM indicates that the item is selected under the b-matching method. From Equation 13, it is clear that under the b-matching method, if the CAT is operating under an IRT model that is not 1PLM, then the discrimination parameter still plays an important role in the step size. Fortunately under the b-matching method, there is no heavy reliance on the highly discriminating items early in the test, so the step size is not expected to be very big in the beginning and to decrease as test progresses.
It is probably of more interest to the readers how the lower asymptote affects step size, when lucky guessing happens. Equation 13 shows that having the lower asymptote makes the step size smaller, because
Interestingly, having a lower asymptote parameter helps reduce the step size to a larger degree under the MI method than the b-matching method, as shown by the case when c = .25. In fact, for the same item, the ratio of step size under the MI versus the b-matching strategy is
Asymmetry: Early guessing versus early slipping
It can also be shown analytically why under the 3PLM and MI method, the benefit of early lucky guesses is not as much as the loss of early slips, as demonstrated by simulations in Rulison and Loken (2009). The rationale is as follows. Early on in the test when the test information
Formally, the step size under the MI method when a correct answer from lucky guess is registered is given in Equation 12. However, from Equation 6, the step size when an incorrect answer is given to the last item due to slipping is,
Consequently, the ratio of the step size under guessing versus slipping is
If the b-matching method is used for item selection, the step size under lucky guessing is given in Equation 13. The step size under slipping is,
The ratio is therefore
Effect of the Upper Asymptote on Step Size
The model that contains only the upper asymptote is the 3PLMU model, shown in Equation 8. Following Chang and Ying (2008), it can be shown that the step size under the 3PLMU is approximately,
where
When early slips happen, the step size is
(Relative) step size under the MI method under the 3PLMU
Under the 3PLMU, the next item to be selected should satisfy (if the item pool is rich enough)
and the probability of answering the next item correctly is expected to be
Following Equation 16, the step size between
Note that the step size is always positive as it quantifies the magnitude, not the direction. When slipping happens,
However, the product of
(Relative) step size under b-matching under the 3PLMU
If the item pool is rich enough, under the b-matching method
Equation 19 also shows that having an upper asymptote smaller than 1 rather than fixing it at 1 makes the step size smaller, because
Asymmetry: Early guessing versus early slipping under the 3PLMU
Analogous to the results regarding the lower asymptote, having an upper asymptote parameter between 0 and 1 helps reduce the step size to a larger degree under the MI method than the b-matching method, as shown by the case when c = .25. This is because other things being equal, the ratio of step size under the MI versus the b-matching strategy is
Because of such complementary effect, the positive bias related to early guessing is expected to be larger in magnitude than the negative bias related to early slipping under the 3PLMU. The step size under the MI method given an incorrect answer from slipping is provided in Equation 18. Meanwhile, the step size when a correct answer is given from guessing is,
The ratio of the step size under slipping versus guessing is
If the b-matching method is used for item selection, the step size under slipping is given in Equation 19. The step size under guessing is
In summary, the upper asymptote and lower asymptote have complementary effects. Under the 4PLM, the positive bias is smaller than that under the 3PLMU, and the negative bias is smaller than that under the 3PLM. It is also expected that given any IRT model, the a-stratified design helps reduce the bias caused by early slipping or guessing.
Conclusion and Discussion
By re-examining and extending the theoretical results in Chang and Ying (2008), the authors show in this article analytically that under both the MI method and the b-matching method, large discrimination parameter has a magnifying effect on the step size. The adverse effect of early slipping or guessing is therefore exacerbated under the MI method, which favors highly discriminating items early in the test. The adverse effect should be smaller under the b-matching method, which does not show favoritism to highly discriminating items. More interestingly, the preference of the MI method to the low-c and high-d items also exacerbates the adverse effect of early slipping or guessing.
Analytical results provided in this paper further prove that having the upper or lower asymptote parameters rather than fixing them at 1 or 0 (i.e., essentially leaving them out of the model) helps make the step size smaller when early slipping or guessing happens, respectively. The relative step sizes are derived under both the MI and the b-matching method. Notably, the reduction in step size is more pronounced under the MI method than under the b-matching method.
Such analytical results indicate that the effect of early misbehaviors in CAT can be moderated by the inclusion of proper parameters in the IRT models. In fact, models explicitly accounting for both slipping and guessing are not uncommon, for example, the deterministic inputs, noisy “and” gate (DINA) and deterministic inputs, noisy “or” gate (DINO) model in the literature of cognitive diagnosis (Rupp, Templin, & Henson, 2010). The use of the 4PLM in operational CAT, however, requires further investigation on the fit of the 4PLM to real testing data, especially pretesting data used for item calibration, as well as the precision of the asymptote parameter estimates. As noted by the reviewers, obtaining precise estimates of the c parameter in the 3PLM is known to be a challenge unless the sample size is large and the sample covers the lower end of the ability continuum adequately. The same applies to the estimation of the upper asymptote parameter d. The sample size required to accurately calibrate the 4PLM is even larger and there needs to be enough slipping happening in the data. Recent discussions of the calibration of the 4PLM such as Loken and Rulison (2010) adopted Markov Chain Monte Carlo (MCMC) method, which means priors are imposed on these parameters. It is still an open question how to pick appropriate choices of the priors and the magnitude of influence of informative priors on the estimation outcomes. Overall, this article provides required technical details to implement a CAT based on the 4PLM when item parameters are available, and show theoretical advantages of including proper asymptote parameters in the model to combat guessing or slipping. However, the authors advise readers to exercise caution when they decide on which IRT model to use. Model fit and precision of parameter estimates need to be evaluated.
In addition, the analytical results are derived under the MI method and b-matching method. When other mechanisms are introduced in the item selection of CAT, for example, certain exposure control strategies, the step sizes in Equations 12, 13, 18, and 19 will no longer apply and neither will the relative step sizes. But the general formula (Equations 6, 7, and 16) are still valid. If the add-on mechanism alleviates the reliance on highly discriminating items early in the test, the adverse effect of early guessing or slipping can be reduced compared with MI method. For example, the a-stratified design should help reduce the negative bias of early slipping under the 3PLM or the positive bias of early guessing under the 3PLMU, while leading to much more balanced item pool usage than the MI method. The a-stratified method was initially proposed to control the exposure of highly discriminating items and promote the use of low-discrimination items in CAT. It groups items into several strata according to their discrimination parameters, from low-a stratum to high-a stratum. In the early stage of a CAT, items can only be chosen from the low-a stratum and gradually moves on to higher strata (Chang, Qian, & Ying, 2001; Chang & Ying, 1999). Such “ascending a” principle is shown effective for balancing item exposure in Chang and Ying (2008) and Cheng, Chang, Douglas, and Guo (2009). Because the step size is proportional to the a-parameter under the 4PLM and the smaller models it subsumes, using low-a items early in the test helps reduce the effect of early slipping and early guessing.
One last thing to emphasize is that the analytical results in this article apply to maximum likelihood ability estimates. Rulison and Loken (2009) used simulation study to show how the upper asymptote helps reduce negative bias associated with early slipping, by contrasting the posterior distribution of ability given the same response pattern under the 3PLM versus under the constrained 4PLM model by Barton and Lord (1981). The analytical results in this article can be extended to Bayesian estimates of ability, for example, maximum a posteriori (MAP) estimates. This will be pursued in the future.
Footnotes
Acknowledgements
The authors would like to thank the action editor, Dr. Dan Bolt, and two anonymous reviewers for their insightful suggestions and comments.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: The authors thank the Institute for Scholarship in the Liberal Arts (ISLA) at the University of Notre Dame for a Small Grant for Research and Creative Work awarded to the first author
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
