Abstract
Under the theory of sequential design, compound optimal design with two optimality criteria can be used to solve the problem of efficient calibration of item parameters of item response theory model. In order to efficiently calibrate item parameters in computerized testing, a compound optimal design is proposed for the simultaneous estimation of item difficulty and discrimination parameters under the two-parameter logistic model, which adaptively focuses on optimizing the parameter which is difficult to estimate. The compound optimal design using the acceptance probability can provide ability design points to optimize the item difficulty and discrimination parameters, respectively. Simulation and real data analysis studies showed that the compound optimal design outperformed than the D-optimal and random design in terms of the recovery of both discrimination and difficulty parameters.
Keywords
Introduction
Computerized adaptive testing (CAT) is a tailored test that adapts to ability level of each examinee (Chang & Ying, 2007). Compared with a pen-and-paper or linear test, CAT has the advantages of high accuracy, short test, and real-time report. Quellmalz and Pellegrino (2009) report in Science that many large-scale assessment programs and formative assessments have implemented online testing or CAT (Chang, 2004; Chang et al., 2016; Chang & Ying, 2007; Quellmalz & Pellegrino, 2009; van der Linden & Glas, 2000; Yan et al., 2014). In order to take full advantages of CAT, it is indispensable for CAT to have a large-scale item bank with already calibrated parameters of item response theory models (Berger et al., 2019; Huebner, 2010; Reckase, 2010; Stocking, 1994; Wainer et al., 2000; Weiss & Kingsbury, 1984; You et al., 2010). In the process of continuously building, maintaining, and expanding an item bank of CAT, online calibration of item parameters of new test items is required (Chang & Lu, 2010; Stocking, 1994). Online calibration of item parameters is accomplished by “seeding” new items into the adaptive test sessions with previously calibrated items (Wainer et al., 2000; Wainer & Mislevy, 1990). Because it is an economical and effective way to collect item response data and expand item bank, online calibration (Chang & Lu, 2010; Chen, 2016; Chen et al., 2017; Chen & Wang, 2016) has long been concerned by researchers in CAT.
Online calibration method and online calibration design are two important aspects of online calibration (Chen & Xin, 2014). The online calibration method aims to estimate the item parameters of the new items, while online calibration design is intended to design the way the new items are assigned to test-takers. As early as the 1990s, several online calibration methods have been proposed under the unidimensional item response theory (UIRT) model, such as Method A and Method B (Stocking, 1988), and the marginal maximum likelihood estimate with one EM cycle (OEM; Wainer & Mislevy, 1990). Method A fixes the estimated abilities obtained from the responses to the operational items during the field test and applies the conditional maximum likelihood estimation to estimate item parameters of new items based on the collected item responses and the fixed abilities. Method B fixes the estimated abilities to estimate item parameters of new items and anchor items and uses the scale transformation obtained from the initial estimates and re-estimates for the anchor items to place the item parameters of new items on the scale of item pool. As a variation of the OEM method, the marginal maximum likelihood estimate with multiple EM cycles (MEM) has been proposed in later studies (Ban et al., 2001, 2002). In recent years, online calibration is an ongoing topic in the field of CAT. You et al. (2010) proposed the conditional maximum likelihood estimation method, which is an alternating iteration algorithm for estimating discrimination and difficulty parameters. The OEM and MEM was extended from dichotomous item response theory models to the generalized partial credit model (Zheng, 2016). For considering the measurement errors in ability estimation, Chen (2016) applied a full function maximum likelihood estimation (FFMLE) and an estimator which exploited the consequences of sufficiency (ECSE) to correct the estimation error of ability before using Method A for online calibration. He et al. (2017) introduced the modified Lord’s bias-correction method to correct the deviation of maximum likelihood ability estimates in Method A, namely, MLE-LBCI-Method A. Moreover, Method A and its variations have been generalized to multidimensional item response theory (MIRT; Chen, 2017; Chen & Wang, 2016; Chen et al., 2017; Hassan & Miller, 2024). Method A will be used in our study, because it is the relatively simplest and most straightforward method among all available calibration methods in UIRT (Chen & Wang, 2016; He et al., 2017).
Beside online calibration methods, two kinds of online calibration designs are developed in the UIRT. The first kind requires a static examinee pool, where examinees with suitable ability values will be selected to calibrate new items (He et al., 2020). This kind is suitable for a parallel testing situation (Ul Hassan & Miller, 2019). The design points or examinees’ abilities in a static examinee pool can be obtained from the D-optimal design, which is based on a determinant criterion of item parameter information matrix (Berger, 1992; Chang & Lu, 2010; Jones & Jin, 1994; Ul Hassan & Miller, 2019). The second kind of examinee pool is dynamic, because examinees are allowed to leave immediately after their operational test sessions (He & Chen, 2020) and every time only an examinee reaches a seeding location (Zheng, 2016). van der Linden and Ren (2015) proposed an reformed D-optimal design (denoted as D-VR design) for online calibration based on Bayesian optimality criteria. The alternative two-point D-optimal design (denoted as D-Tp) was introduced for selecting the new item with the greatest increase in the determinant among all new items to the current examinee (Ren et al., 2017). He and Chen (2020) proposed a new online design (denoted as D-c design) to select the new item for which the current examinee would yield the greatest decrease in the determinant of the covariance matrix of item parameter estimates. Moreover, a lot of online designs were developed based on the excellence degree criterion (He et al., 2020), the maximum
It is not difficult to find out from the above studies that random design and D-optimal design have been adapted into static and dynamic scenarios. Random design has the advantage of being easy to implement, but its accuracy is not as good as optimal design. The basic goal of optimal design is to select as few control variables as possible to estimate unknown parameters of the model (Kitsos, 2013; Silvey, 1980). For example, the maximum Fisher information (MFI) method in CAT (Lord, 1971) is a fully sequential D-optimal design method for selecting item parameter design points in order to accurately estimate the ability of examinees. Sequential design allows for iteratively selecting informative experiments based on current knowledge about the parameters (e.g., Cao et al., 2024). Thus, CAT is more efficient (e.g., shorter test length) and accurate (e.g., higher precision for all examinees) than a linear test. Similar to the MFI method used in CAT, instead of items, the optimal calibration design is to select the most suitable ability design points for a new item according to a certain criterion, so as to solve the problem of item parameter calibration with a limited number of examinees.
Now in the framework of online calibration, the D-optimal design still has some problems. From the relevant study, it can be seen that the D-optimal design has the problem of large error in the estimation of discrimination parameters of high discrimination items and difficulty parameters of low discrimination items (Chang & Lu, 2010). The high estimation error of discrimination parameters in the small sample calibration conditions leads to the overestimation of the precision of the ability estimates in CAT, which is called as the error capitalization on chance (COC) problem (Olea et al., 2012; Patton et al., 2013). Although cross-validation techniques are regarding as possible solutions (van der Linden & Glas, 2000), these methods only suppress symptoms without addressing underlying causes. There is still a need to develop parameter estimation procedures for reducing the estimation errors of the item parameters (Olea et al., 2012).
To fundamentally tackle the problem from its root, it is necessary to propose optimal designs and calibration methods to improve the accuracy of item parameter estimates. The main purpose of the study is to explore an optimal design for online calibration, optimizing the parameter which is difficult to estimate. The new design is first introduced under the static examinee pool. Then, the new design is extended to the dynamic examinee pool from the idea of the D-Tp, because it has been generalized from the static examinee pool to the dynamic examinee pool and performs very well (He & Chen, 2020; Ren et al., 2017).
The rest of this article is organized as follows. First, the UIRT model and online calibration method used in this article are described, followed by detailed introduction of design for a static examinee pool and a dynamic examinee pool, respectively. Then, simulation studies and real data analysis are conducted for illustrating the performance of the new optimal design. Finally, we conclude with a discussion. The detailed results for the static examinee pool are given in the appendix.
Method
Two-Parameter Logistic Model
The item response function of the two-parameter logistic model (2PLM) can be written as
Online Calibration Method
As described above, online calibration methods mainly include Method A, OEM, and MEM. Since this study focuses on online optimal design, Method A as a simple and effective online calibration method was used and introduced below.
Let
Since maximizing the logarithm of likelihood function with respect to item parameter vector
The iteration formula of the Newton-Raphson algorithm is
Using the information matrix
For 2PLM, the first-order derivative and information matrix in equation (6) can be written as respectively
Design for a Static Examinee Pool
Random Design
Random design, that is, completely randomized design, refers to random assignment of examinees to new items in the field test for online calibration. In other words, examinees are randomly assigned or selected to answer new items, and the position of new items are also randomly seeded in a field test. The random design is very simple to implement, but it does not lead to either an optimal match between item’s difficulty and examinee’s ability or the adaptive of CAT.
D-Optimal Design
As can be seen from Section “Online Calibration Method,” the online calibration of item parameters is a non-linear problem. For the non-linear model or problem, because optimal design relies on unknown parameters, sequential D-optimal design (Chang & Lu, 2010) and optimal Bayesian adaptive design (van der Linden & Ren, 2015) are usually used. That is, the calibration of the unknown parameters is performed while collecting the sample data, and then the optimal design point is selected according to the current estimation of the unknown parameters, and the process is repeated until a certain termination condition reached.
The D-optimal design is one of the commonly used designs under the fully sequential design. The performance of the D-optimal design for CAT online item calibration has been investigated on test items from the synthesized data and National Assessment of Educational Progress (Chang & Lu, 2010). Parameter calibration is one of the basic tasks of statistics or measurements. To accurately calibrate item parameters of a new item with a limited number of examinees, it is necessary to select examinee ability or design point according to certain criteria, so that the design point can minimize estimation error of item parameters for the new item. For the current sample size of
D-Tp Design
Suppose that the current estimates of item parameters are known, 2PLM can be regarded as a logistic regression model. For the two-point design, the previous studies suggested that the D-optimal design points should satisfy
Equation (9) is slightly different from equation (10), and the former is a one-point design and the latter is a two-point design.
Compound Optimal Design
The design is based on the diagonal elements of Fisher information matrix to find the design points based on the following two optimality criteria
One solution for a specific variable
Since the following three equations are established
It should be noted that although the sum of these two matrices is a diagonal matrix,
There are two optimality criteria in equations (11) and (12). Obviously, the corresponding designs do not dominate each other, in that one design will yield a higher efficiency with respect to one criterion but not the other. An approach to simultaneously maximize two criteria was proposed in Eccleston and Whitaker (1999). The design point is accepted with a probability. It is similar to the acceptance probability which refers to the likelihood of a candidate solution being accepted based on its quality and often used in simulated annealing algorithms or Metropolis–Hastings algorithms (Givens & Hoeting, 2013). This idea is implemented in our study to develop a new way of combining these two opposing criteria, which can be regarded as a compound design criterion (McGree et al., 2008). Thus, the two-points are accepted with the following probability
Design for a Dynamic Examinee Pool
Let
Random Design
A new item is randomly selected from the set of
D-optimal Design
For the current active examinee
D-Tp Design
The two-point D-optimal design selects the new item that satisfies (He & Chen, 2020; Ren et al., 2017)
Compound Optimal Design
First, the two design points are calculated by equation (19) for each item
Simulation Studies
Two simulation studies were conducted to investigate the performance of four designs of random design (R), the single point D-optimal design (D), the two-point D-optimal design (D-Tp), and the compound optimal design (C) under the static examinee pool and dynamic examinee pool, respectively.
Design
According to the design of related studies (Chang & Lu, 2010; He & Chen, 2020; Ren et al., 2017), three factors were considered here, including item parameters, sample size, and ability estimate error. The simulation design mainly comes from the study of Chang and Lu (2010). They have applied the similar design to investigate the performance of D-optimal design for online calibration via variable length CAT. For the static examinee pool, the values of the discrimination parameter were a = 0.5, 1, 1.5, 2, and 2.5, and three values of the difficulty parameter were only considered, b = 0, 1, and 2, where b = −2 and −1 need not to consider since these cases are symmetric to 1 and 2, respectively. The full factorial design formed a total of 15 item parameter combinations. For the dynamic examinee pool, the difficulty of −1 and −2 (i.e., b = −1 and −2) were considered in the study, and formed a total of 25 item parameter combinations. The sample size was divided into three levels,
We did not conduct simulations within a CAT framework or a computerized test, because many factors (such as item bank size, item parameter distribution, item selection, ability estimation, seeding location, test length, or stopping rule) influence the standard error of ability estimates, which is difficult to be manipulated. Thus, we directly manipulated four levels of the standard error of ability estimates based on the approach of Chang and Lu (2010) to create four scenarios that cover a range of outcomes, from ideal to realistic. The standard deviation of the ability estimation error contains three fixed levels, σ = 0, 0.2, and 0.5, where σ = 0 denoted the true or simulated ability without adding error to estimate item parameters, σ = 0.2 denoted a predetermined standard error criterion for variable length CAT (Choi et al., 2011), whose corresponding information of the ability estimation was 25, and σ = .5 denoted the condition of the fixed large error of ability. When information function of an item pool with uniform difficulty parameters is relatively flat (Babcock & Weiss, 2012), or the optimal item pool with sufficient numbers of good quality items that are most informative at a series of ability levels (He & Reckase, 2014), the variable length CAT can achieve a predetermined level of measurement precision and yield near equivalent measurement precision across the entire range of ability (Choi et al., 2011). If the operational item pool has difficulty parameters centered around a specific range of the ability scale (e.g., −1 < θ < 1) and has fewer items appropriate for administration for extreme ability levels, the operational CAT can yield better measurement precision for examinees at the middle of the ability scale than at the extremes of the ability scale. The study of He and Reckase (2014) shows that conditional error of ability estimates for the retied operational item pool has the shape of a parabolic bowl with its bottom at origin (He & Reckase, 2014). Because conditional standard errors of ability estimates for CAT are almost in the range of 0.2 and 0.65 on an interval
Procedures
The Static Examinee Pool
The specific steps for estimating item parameters of each new item in four designs were as follows: Step 1 Choose one combination of item parameters in turn to produce results for each combination. Step 2 Obtain the initial estimation of item parameters. First 50 ability values, denoted as Step 3 Select the ability design points based on the random design or the optimal designs, respectively. For the random design, each element of estimated ability vector Step 4 Update the estimates of item parameters. The
The Dynamic Examinee Pool
The specific steps for estimating item parameters for all new items in four designs were as follows: Step 1 Simulate the abilities of the examinees and their item responses on the new items. The number of new items is Step 2 Obtain the initial estimation of item parameters. In the initial stage, each new item should receive Step 3 Select the current active examinee with the estimated ability Step 4 Select the new item Step 5 Let
Evaluation Criteria
The bias, absolute deviation (Abs), and mean square error (MSE) of item parameters estimator were used to measure the accuracy of item parameters’ estimates. Their calculation formulas are given by
Results
Estimation Precision of Item Parameters for Four Designs Under Different Ability Errors.
Estimation Precision of Item Parameters for Four Designs Under Different Sample Sizes.
Bias of Item Parameters for Four Designs With Sample Size of 120 and Ability Error of 0.2.
Note. The value closest to zero was highlighted in bold.
Abs of Item Parameters for Four Designs With Sample Size of 120 and Ability Error of 0.2.
Note. The minimum value was highlighted in bold.
MSE of Item Parameters for Four Designs With Sample Size of 120 and Ability Error of 0.2.
Note. The minimum value was highlighted in bold.
MSE of Item Parameters for Four Designs With Sample Size of 120 and Conditional Standard Error of
Note. The minimum value was highlighted in bold.
Second, the impact of sample size on the estimation precision of item parameters for four online designs was considered. Table 2 presents the mean squared errors across two factors of item parameters and ability estimation errors. It was found that the mean square error for both discrimination and difficulty parameters became smaller as the sample size increased for the compound optimal design. The sample size could effectively improve the precision of difficulty parameter estimates for four designs. While for the random design and two D-optimal designs, increasing the sample size did not improve the precision of discrimination parameter estimates.
Tables 3–5 show the BIAS, ABS, and MSE under the condition that the sample size is 120 and the standard error of the estimated ability is 0.2. Similar results from other conditions were not shown here. It can be seen from the mean for each column or the penultimate row of Tables 3–5 that for four online designs, the means for BIAS, ABS, and MSE of the item parameter estimation from the compound optimal design were almost the smallest, except that random design has smallest BIAS of difficult parameter than other designs. The performance of the D-Tp and D-optimal designs was very similar and was better than random design. On average, it indicates that the compound optimal design was superior to the other three designs. The variability of the estimation precision for different true item parameters was applied to examine the robustness of four online designs. From the standard deviation for each column or the last row of Table 5, we found that true item parameters had a small impact on the MSE for the compound optimal design. Thus, the compound optimal design performed quite robustly under a variety of item parameters.
From Table 5, for the precision of the estimate of discrimination parameter, the compound optimal design produced smaller MSE (
Real Data Analysis for TIMSS Items
In the real data analysis, item responses, item parameters, and abilities are taken from a set of International Mathematical and Scientific Trends Research (TIMSS; Foy et al., 2013). The data comes from a 0–1 scored data set in the CDM package written in R language (Robitzsch et al., 2022). The file name of the data set is “data.timss11.G4.AUT.rda,” which is the fourth-grade Austrian math test data in the 2011 TIMSS. The total sample size is 4,668, which is the same as the number of examinees shown in the related document published on the website (https://timss.bc.edu). The published TIMSS test consists of a total of 175 Items. Codes of 174 items contained in the data set are identical to codes published on the website, while one published item, identified as “M051061Z,” was not included in the R dataset.
The way of the study of You et al. (2010) was used to simulate the procedures of online calibration. The 2PLM was used to fit all item responses in the dataset. The Bilog-MG software was employed to estimate item parameters and ability parameters. Since the TIMSS test was administrated by use of a balanced incomplete block (BIB) design, there is a large number of missing item responses. In the estimation of item parameters, missing item responses were treated as not presented, an option available in Bilog-MG (Du Toit, 2003; Finch, 2008). Only 13 items did not fit 2PLM well because the p-values of the chi-square statistics for these items are less than 0.05. The item parameter estimates were taken as the true values of item parameters. The mean of discrimination and difficulty were 1.002 and −0.0673, respectively. By default, expected posterior estimates of ability from Bilog-MG were regarded as the ability estimates for examinees taking the CAT or the field test, which would be used in the online calibration. Online calibration of item parameters of the 174 items was very similar to the procedures under the dynamic examinee pool in the simulation study, except that the simulation of item responses and the ability estimates were not necessary, because real item responses and the ability outputs from Bilog-MG were available. The active examinee was sequentially selected in the order of the row of item response matrix. It should be noted that measurement errors were not needed to add in the estimation ability because there were errors in the ability outputs. For the design of online calibration of item parameters, the ability design points were generated by one of the four online designs under the dynamic examinee pool and the corresponding item responses without missing of examinees were then selected. The estimation of item parameters and the selection of item responses were repeated until the number of item responses reached 130. Because the sample size is 4,668, the number of items is 174, and 5 items are seeded to each examinee, the mean sample size for each item is
Results of Empirical Study for Four Online Designs.
Note. The minimum value was highlighted in bold.
Conclusion and Discussion
Online calibration of new items in the operational CAT provided an economical and efficient way for maintaining and expanding of item bank. Through the analysis of existing related literatures, it was found that the D-optimal design has some advantages but it results a large estimation error of discrimination parameters for high discrimination items or difficulty parameters for low discrimination items. Moreover, previous studies have shown that the frequent use of high discrimination parameter items with large estimation errors in CAT has the potential effects of capitalization on chance (Patton et al., 2013), which ultimately seriously affects the precision of CAT ability estimation. Therefore, it is very important to investigate a new design for online item calibration under the 2PLM.
The contribution of our study to the literature has two aspects: First, we propose the compound optimal design with two optimality criteria for online calibration of item parameters of new items under the 2PLM. The design using the acceptance probability adaptively selects design points for optimizing the parameter which is difficult to estimate. The new design improves the estimation precision of item difficulty and discrimination parameters in simultaneous. Second, the compound optimal design is suitable for both the static examinees pool and the dynamic examinees pool. The static examinees can be obtained from group-based testing or a whole group (e.g., school class or school grade) of test-takers undergoing assessment concurrently (Bengs et al., 2021). The concurrent test-taking aligns with usual classroom testing or state testing programs, such as the VERA (Vergleichsarbeiten). The VERA are comparison tests that students take in the 3rd and 8th grades (VERA-3 and VERA-8) in Germany. In recent years, several German federal states decided to implement the computer-based testing of the VERA tests (Wagner et al., 2022). For the situation, a static examinee pool is available when there are many students attending computer-based testing in parallel on the operational item pool. In addition, the static examinees can be sampled based on ability-related background variables, such as grades or courses taken in school (Berger et al., 2019), because the background variables are usually related to the ability (Eggen & Verhelst, 2011; Mislevy & Wu, 1996). Furthermore, if the number of students attending computer-based testing in parallel is small, a dynamic examinee pool is available for sequential designs in the testing period.
The simulation study shows that the compound optimal design can achieve an efficient estimation of the discrimination and difficulty parameters at the same time. In other words, the compound optimal design can control the estimation error of discrimination and difficulty parameters simultaneously. With the same sample size, the compound optimal design outperforms the D-optimal design and random design in terms of the precision of item parameters. The compound optimal design also performs well in real data analysis. Thus, the compound optimal design can obtain the simultaneous estimation of item difficulty and discrimination parameters, which can be used for online calibration of item parameters.
The advantage of the compound optimal design is that the practitioner can focus on the estimation of important item parameters that is vital for item selection and ability estimates. It can directly control the estimation error of item parameters through the selection of the ability design points by the definition of the acceptance probability. Whether complex methods for the simultaneous maximization of two objective functions in the optimization of designs can help to improve the performance of the designs which require further study. The fixed sample size, the 2PLM, and one online estimation method were considered in the study. Other optimal design criteria, optimal design method, and stopping rule should be further studied for the replenishment of item bank. For example, we should extend the compound optimal design and compare it with the optimal Bayesian adaptive design (He & Chen, 2020; Ren et al., 2017; van der Linden & Ren, 2015). For the 2PLM, if one of item parameters (i.e., difficulty parameter) has fast achieved a relatively small estimation error, then it will leave a large number of examinees to estimate other item parameter (i.e., discrimination parameter). Thus, the stopping rule needs further study. The 2PLM was used in the study, while optimal designs under other models in the framework of item response theory need to be further considered. Besides, whether optimal designs and online estimation methods would be interacted with each other needs to be further discussed. In the study, ability estimation was only used for online calibration of item parameters on new items. It is worth paying attention to how to estimate and optimize the item parameters of the new item with the help of other auxiliary information, such as response times and related characteristics of test items (He et al., 2021; Kang et al., 2020; Li & Kuang, 2023). How to generalize the new design to multidimensional item response theory (MIRT) is an interesting problem, because some studies have extended related methods to the MIRT (Chen, 2017; Chen et al., 2017; Chen & Wang, 2016).
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was funded by the National Natural Science Foundation of China (grant numbers 62067005, 62267004, and 62467003).
