Abstract
The within-item characteristic dependency (WICD) means that dependencies exist among different types of item characteristics/parameters within an item. The potential WICD has been ignored by current modeling approaches and estimation algorithms for the deterministic inputs noisy “and” gate (DINA) model. To explicitly model WICD, this study proposed a modified Bayesian DINA modeling approach where a bivariate normal distribution was employed as a joint prior distribution for correlated item parameters. Simulation results indicated that the model parameters were well recovered and that explicitly modeling WICD improved model parameter estimation accuracy, precision, and efficiency. In addition, when potential item blocks existed, the proposed modeling approach still demonstrated good performance and high robustness. Furthermore, the fraction subtraction data were analyzed to illustrate the application and advantage of the proposed modeling approach.
Keywords
Introduction
Currently, a number of cognitive diagnosis models (CDMs) have been developed (see, for example, Rupp, Templin, & Henson, 2010). The deterministic-inputs, noisy “and” gate (DINA) model (Junker & Sijtsma, 2001) is the most popular one because of its interpretability. Different parameterizations of the DINA model (e.g., DeCarlo, 2011; de la Torre, 2011; Henson, Templin, & Willse, 2009; Junker & Sijtsma, 2001; von Davier, 2014) and estimation algorithms (e.g., Culpepper, 2015; de la Torre, 2009) have been proposed. However, the potential within-item characteristic dependency (WICD; Fox, 2010) has been ignored by these modeling approaches and estimation algorithms, which may underestimate the real-world complexity. The WICD means that dependencies exist among different item characteristics/parameters within an item. More specifically, the guessing and slipping parameters in the DINA model are not independent of each other. A correlation could be used to quantify such a relationship between the guessing the slipping parameters in the DINA model.
The correlation among item parameters in the DINA model was supported by many empirical studies. For example, a correlation coefficient of −.892 was found between the guessing and slipping parameters based on the examination for the certificate of proficiency in English data in a study of Templin and Hoffman (2013); a correlation coefficient of −.802 was found for the revised Purdue spatial visualization tests—visualization of rotations data in Culpepper (2015); a correlation coefficient of −.686 was found for the FCAT (Florida Comprehensive Assessment Test) 2003 ninth grade mathematics data in Lee, de la Torre, and Park (2012); furthermore, a correlation coefficient of −.415 was found for the fraction subtraction data in de la Torre (2009). It can be seen that the estimated guessing and slipping parameters in the DINA model were negatively correlated based on the real data from reading, mathematics, and cognitive psychological tests although the WICD was ignored in the modeling and estimation process. Thus, this study aims at accounting for WICD in the DINA model within the Bayesian framework.
The rest of the article starts with a review of the DINA model. Then, it demonstrates a way to account for WICD in the DINA model, which is estimated with the Bayesian approach. Three simulation studies evaluating the feasibility of the proposed modeling approach are followed. Finally, an empirical example is presented to illustrate the application of the proposed approach for WICD.
Standard Bayesian DINA Modeling
Let
where pni is the probability of person n with attribute profile c correctly responding to item i; α
ck
is the kth element of attribute profile c (c = 1,. . . ., C; k = 1, . . . ., K), where α
nk
= α
ck
= 1 if person n masters the kth attribute of attribute profile c, and α
nk
= α
ck
= 0 otherwise;
Let π
c
= P(
In the Bayesian framework, the posterior distribution of the model parameters is proportional to the product of their prior distributions and the likelihood of item response data, which is given by,
where P(
where P(gi) and P(si) are the prior of item guessing and slipping parameters, respectively. In general, a monotonicity restriction (gi < (1 –si)) can be imposed on Equation 4. The monotonicity restriction imposed a within-item scale/range dependence upon slipping and guessing parameters, but it is different from WICD.
As the guessing and slipping parameters are on the probability scale, the Beta distributions are often specified as their prior distributions,
where
Bayesian DINA Modeling Incorporating the WICD
Although a priori independence assumption does not imply the parameter independence in the posterior distribution (Levy & Mislevy, 2016), the WICD is not explicitly modeled. To account for the WICD, a multivariate prior distribution could be specified for the item parameters (Mislevy, 1986). In the item response theory framework, some studies (e.g., Glas & van der Linden, 2003; van der Linden, 2007) have explored to account for WICD using the multivariate normal distribution. The same method could be applied to the DINA model to account for WICD. In this study, a bivariate normal distribution is employed as a joint prior distribution of item parameters, as follows:
where logit(x) = log (x / (1 –x)), such transformation frees up the restriction of range for the original gi and si parameters (i.e., 0-1);
where ρβδ is the correlation coefficient between the logit-transformed item parameters.
Overall, the main difference between the standard and the proposed Bayesian DINA modeling approaches is the difference in the priors of the item parameters, that is, a truncated bivariate uniform prior versus a joint bivariate normal prior distribution. For illustrative purposes, the proposed modeling approach is denoted as the W-DINA and the regular one is denoted as the U-DINA throughout this article. It is noted that the W-DINA and the U-DINA have the same item response function but the former uses a more informative prior for item parameters than the latter to account for the WICD.
Bayesian Parameter Estimation
The Bayesian Markov chain Monte Carlo (MCMC) method was used to estimate model parameters. In this study, the JAGS (version 4.3.0; Plummer, 2015) were used to automate the estimation process. The JAGS uses a default option of the Gibbs sampler (Gelfand & Smith, 1990). JAGS code for the W-DINA and the U-DINA are provided in Table S1 in online supplements.
To begin with, given local independence,
In addition, in the W-DINA, the prior of the item parameters is assumed to be a joint bivariate normal distribution as follows,
with the hyper priors specified as µδ ~ N(–2.197, 2), µβ ~ N(–2.197, 2),
Simulation Study
Three simulation studies were conducted. The purposes of Study 1 are two-fold. First, to examine whether parameters of the W-DINA can be recovered accurately; second, to evaluate the consequences of ignoring WICD, in which the data were simulated from the W-DINA, but analyzed with the U-DINA. However, the method of generating item parameters in Study 1 yields values in a subset of the identified space. It is, therefore, expected that an informative prior would outperform a uniform prior. Thus, the main purpose of Study 2 is to evaluate the performance of the W-DINA when the item parameters were simulated from a uniform distribution of the whole identified space. In addition, in practice, items in the test may be grouped into several item blocks. For instance, content constraints may be taken into consideration in test construction (Yi & Chang, 2003). It is possible that items in the same item block have more similar characteristics. Thus, Study 3 mainly investigated the impact of item blocks on parameter estimation of the W-DINA.
Simulation Study 1
Design and data generation
In Study 1, five independent variables were manipulated including (a) the degree of the WICD (ρβδ) at two levels of −0.3 and −0.9; (b) the values of µβ and µδ (

K-by-I Q matrix for Study 1.
Analysis
Thirty replications were implemented in each simulated study condition. For each replication, two Markov chains with random starting points were used and 10,000 iterations were run for each chain. The first 5,000 iterations in each chain were discarded as burn-in. Finally, 10,000 iterations were retained for model parameter inferences. The potential scale reduction factor (PSRF; Brooks & Gelman, 1998) was computed to assess the convergence of each parameter. In our studies, the PSRFs were generally less than 1.01, suggesting good convergence in the specified settings. 1 To evaluate parameter recovery, the bias and root mean square error (RMSE) of the item parameter estimates were computed. The pattern correct classification rates (PCCRs) were computed to evaluate the classification accuracy. In addition, the number of iterations for convergence, the effective sample sizes (Kass, Carlin, Gelman, & Neal, 1998), and the overall computing time 2 for the same total number of iterations were employed to evaluate the difference in computational efficiency between the two modeling approaches. The ratio of the computing time for the U-DINA over that for the W-DINA was reported.
Results
Figures 2 and 3 present the recovery of item parameters for the W-DINA and the U-DINA in conditions with a test length of 30 items. Due to page limit, the results for conditions with 15 items were presented in Figures S1 and S2 in online supplements. First, the W-DINA and the U-DINA are similar in the patterns of the item parameter recovery. For the si parameter, more required attributes lead to higher bias and RMSE. By contrast, the estimation of the gi parameters was less biased and more accurate than the si parameters, and the variability of the gi parameter estimates was slightly smaller for items that required more attributes. As expected, the bias and RMSE both decreased as sample size increased. The mean quality of items slightly affects the recovery of item parameters; its impact was more evident in conditions with short test length, and the impact on the U-DINA was more noticeable than that on the W-DINA.

The recovery of the guessing parameters across I = 30 conditions in Study 1.

The recovery of the slipping parameters across I = 30 conditions in Study 1.
Regarding the difference between the two modeling approaches, the recovered gi parameters between them were quite similar; however, the bias and RMSE of the W-DINA were slightly smaller than those of the U-DINA, and such advantages increased as sample size decreased. By contrast, the recovery of the si parameters for the W-DINA was more accurate and more precise than the U-DINA, and such advantages also increased as sample size decreased. The W-DINA has outperformed the U-DINA in terms of the si parameters in items with multiple attributes. Such results were within expectation because the recovery of the si parameters in items with multiple attributes was worst in the U-DINA, which means the room for potential improvement was the greatest. Likewise, the W-DINA also slightly outperformed the U-DINA in terms of the gi parameters in items with fewer attributes.
Obviously, the sample size and the strength of WICD had more impact on the difference between these two modeling approaches. The differences for the conditions with a small sample size were significantly larger than those with a large sample size. A possible explanation is that when the sample size is large enough, response data provide enough information to assure better accuracy in item parameter estimation. Thus, the influence of the prior distribution decreases. In addition, the W-DINA led to higher parameter estimation accuracy as the strength of WICD increased. Essentially, modeling the WICD borrow information among item parameters in estimation. That is, this joint distribution is an augmented approach in utilizing extra information in the gi parameters when estimating the si parameters.
For the recovery of attribute profiles see Figures S3 and S4 in online supplements. First, the patterns in the recovery of attribute profiles for the W-DINA were similar to those for the U-DINA. In almost all the conditions, the PCCR decreased as the magnitude of the WICD increased. A larger number of items and higher item mean quality improved the attributes recovery for both models. Regarding the difference between the two modeling approaches. For a given condition, the W-DINA was slightly better than the U-DINA. Although the advantages in the PCCR of the W-DINA were less than 0.015, the advantages increased as the test length increased, the extent of the WICD increased, the mean quality of items decreased, or the sample size decreased.
For the number of iterations needed to reach convergence, the multivariate potential scale reduction factor (MPSRF) was computed (Brooks & Gelman, 1998) for each replication. As similar patterns were found in different conditions, Figure 4 presents the MPSRF in several conditions for illustration. First, for a given condition, the pattern of the MPSRF for the W-DINA was quite similar to that for the U-DINA, which means that imposing the WICD does not evidently affect the convergence of the posterior distribution. Second, test length affects convergence, whereas the sample size, the strength of the WICD, or the mean quality of items does not. Although employing a burn-in of 1,000 iterations is adequate for both models, a burn-in of 5,000 iterations was employed in this study to be conservative.

Plots of Brooks–Gelman MPSRF for selected conditions in Study 1.
In addition, Figure 5 plots the mean, minimum, and maximum effective sample sizes of each item parameter across 30 replications to provide evidence regarding the degree of autocorrelation in the Markov chains. Larger effective sample sizes are associated with smaller autocorrelation. As similar patterns were found in different conditions, only limited conditions were presented. The effective sample size was computed from two chains with all iterations, resulting in an expected value of 20,000. The results indicated that imposing WICD has no significant effect on effective sample sizes.

The mean, minimum, and maximum effective sample sizes for item parameters in selected conditions in Study 1.
The sample size had some effect on the ratio of overall computing time of the U-DINA over the W-DINA, and the ratios were almost the same for conditions with the same sample size. The mean ratio is 1.418 (SD = 0.011) for conditions with a sample size of 1,000 and 1.313 (SD = 0.004) for conditions with a sample size of 200, respectively. Consequently, the W-DINA is more efficient in parameter estimation than the U-DINA.
Overall, both the W-DINA and the U-DINA parameters were well recovered. Explicitly, modeling the WICD can further improve model parameter estimation accuracy, precision, and efficiency, especially when the sample size is small.
Simulation Study 2
Design and data generation
In Study 2, two independent variables were manipulated including (a) the test length (I) at two levels of 15 and 30 and (b) the sample sizes (N) at two levels of 200 and 1000. The Q-matrices in Figure 1 were still used. The true attribute profile of each person was still randomly chosen from 32 possible patterns with equal probability. To shed light on the potential limitations or the performance in boundary conditions of the proposed modeling approach, item parameters were generated from si ~ U(0, 1) and gi| si ~ U(0, 1 –si). When item parameters were uniformly generated, it means the items were developed without careful screening for quality though such study conditions may rarely exist in practice.
Results
Table 1 summarizes the recovery of item parameters. The disadvantages of the W-DINA were around 0.002 in mean absolute bias and were less than 0.01 in mean RMSE, respectively. Table 2 presents the recovery of attribute parameters. The disadvantages of the W-DINA were less than 0.01 in PCCR across all conditions. Note that such a low PCCR was expected because of unrealistically low item quality. Overall, the W-DINA can provide a comparable recovery of model parameters to the U-DINA in the simulated boundary conditions.
Recovery of Item Parameters in Study 2.
Note. I = test length; N = sample size; MA_Bias = mean absolute value of bias across all items; M_RMSE = mean value of root mean square error across all items; W = W-DINA; U = U-DINA; DINA = deterministic inputs noisy “and” gate.
Pattern Correct Classification Rate in Study 2.
Note. I = test length; N = sample size; W = W-DINA; U = U-DINA; DINA = deterministic inputs noisy “and” gate.
Simulation Study 3
Design and data generation
In Study 3, two independent variables were manipulated including (a) the number of potential item blocks (B) at two levels of 2 and 3 and (b) the sample sizes (N) at two levels of 200 and 1,000. Let subscript b (b = 1, . . ., B) denote the bth item block. Considering the operability, items were assumed to be grouped by content domains containing different sets of attributes. There are six attributes in total measured by 30 items. The corresponding Q-matrices are presented in Figure 6. In Q1, there are two item blocks, and three attributes (Kb = 3) were assessed by 15 items (Ib = 15) in each item block. In Q2, there are three item blocks and two attributes (Kb = 2) were assessed by 10 items (Ib = 10) in each item block. For both Q-matrices, item parameters in each item block were separately generated from,
where ρβδ, b is −0.9, −0.8, and −0.6 in three item blocks, respectively. The true attribute profile of each person was randomly chosen from all possible patterns (C = 64) with equal probability. Note that a common

K-by-I Q matrices for simulation Study 3.
Results
Note that the results of the conditions with different Q-matrices should be viewed separately because they have different structures and characteristics. Figures 7 and 8 present the bias and RMSE of the estimated item parameters for the W-DINA across all conditions. Virtually, all conditions had unbiased estimates, except a few items in N = 200 conditions had small bias less than ± 0.03. Although the RMSE for item parameter estimates increased as the sample size decreased, the largest RMSE was smaller than 0.06. The PCCRs were around 0.63 for Q1 (i.e., B = 2) and 0.75 for Q2 (i.e., B = 3), respectively. Higher PCCRs in Q2 than those in Q1 could be attributed to the fact that the Q2 has a simpler structure than Q1. For instance, 24 out of the 30 items require single attribute in Q2, whereas 12 out of the 30 items require single attribute in Q1. In addition, the W-DINA worked slightly better than the U-DINA with potential item blocks (see details in Tables S2 and S3 in online supplements). Overall, item blocks have little effect on parameter estimation of the W-DINA.

Recovery of the guessing parameters across all conditions in Study 3.

Recovery of the slipping parameters across all conditions in Study 3.
An Empirical Example
Data Description and Analysis
To further demonstrate the application of the proposed modeling approach, a fraction subtraction data from de la Torre (2009) were analyzed. The data set contained a total of 536 people responding to 15 items measuring five required attributes. The total number of possible attribute profiles is 32. The Q-matrix can be found in de la Torre (2009). The W-DINA and the U-DINA were fitted to this data set. The analysis was specified in the same way as the previous simulation study. The deviance information criterion (DIC; Spiegelhalter, Best, Carlin, & van der Linde, 2002) was computed for each model to evaluate the model data fit.
Results
Figure 9 shows the plots of the Brooks–Gelman MPSRF. Compared to the MPSRF for the U-DINA, the MPSRF for the W-DINA assessed the convergence of two additional elements,

Plots of Brooks–Gelman MPSRF in the fraction subtraction data.
Table 3 presents the estimated item parameters. According to the IDI
i
, the quality of these 15 items was high. Most of the estimates of the U-DINA were larger than those of the W-DINA, which was consistent with the patterns of the bias found in Simulation Study 1. In addition, all the standard errors of the U-DINA were equal or larger than those of the W-DINA, which was consistent with the patterns of RMSE found in Simulation Study 1. Table 4 presents the estimated
Estimated Item Parameters for the Fraction Subtraction Data.
Note. DINA = deterministic inputs noisy “and” gate; IDI = item discrimination index.
Estimated Item Mean Vector and Variance and Covariance Matrix for the Fraction Subtraction Data.
Conclusion and Discussion
To explicitly model WICD, this study proposed a modified Bayesian DINA modeling approach, in which a bivariate normal distribution was employed as a joint prior distribution for item parameters. Simulation results indicated that the model parameters for the W-DINA were well recovered and that explicit modeling of WICD can further improve model parameter estimation. The W-DINA outperforms the regular U-DINA (i.e., more accurate, precise, and efficient model parameter estimates, especially for small sample size conditions) in most normal testing conditions. Even under boundary conditions, the W-DINA can provide a comparable recovery of model parameters to the U-DINA. In addition, the proposed modeling approach with a common bivariate normal distribution also worked well when potential item blocks presented. Furthermore, the fraction subtraction data were analyzed to illustrate the application of the proposed modeling approach. Overall, given the results of the simulation and empirical studies, the advantage of explicitly modeling WICD was evident. Although the improvement in parameter estimation accuracy and precision decreased as the sample size increased, the proposed modeling approach still improved the efficiency in parameter estimation.
In general, the use of a multivariate normal distribution between item parameters serves as an augmentation approach that the estimation of one item parameter could borrow information from the other item parameter. Such an augmentation approach for item parameter estimation is quite similar to some proposed augmentation approaches for person parameter estimation (e.g., de la Torre, 2008b; Wang, Chen, & Cheng, 2004). In addition, in many simulation studies related to the DINA model, the guessing and slipping parameters were generated independently, which may not be the case in reality. The item parameter generation approach in this study could be a possible approach to generate correlated within-item parameters in future simulation studies. Furthermore, according to the results, the advantage of using the proposed model was more evident in item parameters when the sample size is small. The practical gains of using the proposed model would be reflected in some applications where items are precalibrated with sufficient precision and low item exposure rate, for example, the cognitive diagnostic computerized adaptive testing (Cheng, 2009).
The work represented in this article is an initial attempt to better understand the WICD in CDMs. Despite promising results, some limitations still exist. Although only the DINA model was used in this study, the proposed modeling method can be easily extended to the Deterministic Input Noisy Output “OR” (DINO) model (Templin & Henson, 2006), because the DINA and the DINO models share a “dual” relation (Köhn & Chiu, 2016). However, unlike the DINA and the DINO models that contain 2I item parameters, different items in more general CDMs (e.g., de la Torre, 2011; Henson et al., 2009; von Davier, 2008) can have different numbers of parameters. Thus, employing a common multivariate normal distribution for all items does not seem feasible. How to consider the WICD in generalized CDMs is an interesting topic. In addition, for potential item blocks existed, it is valuable to include multiple multivariate normal distributions in further exploration. In such cases, the idea of item-family (Glas, van der Linden, & Geerlings, 2016) can be imposed. Currently, the W-DINA only focuses on dichotomous items, but it is meaningful and practical to extend it to polytomous response items (e.g., Ma & de la Torre, 2016; von Davier, 2008). Although the WICD issue was explored in the full Bayesian estimation framework in this study, the WICD issue also exists in current non-Bayesian algorithms for the U-DINA (e.g., the MMLE/EM (Marginal Maximum Likelihood Estimation/ expectation maximization) approach, see Figures S5 and S6 and Tables S4 and S5 in online supplements), which need to be further studied. Furthermore, in the empirical example, the Q-matrix does not satisfy identifiability conditions (Xu & Zhang, 2016), which may yield biased diagnostic classifications. Thus, further analyses are needed based on a modified Q-matrix (Chen, Culpepper, Chen, & Douglas, 2018; de la Torre & Chiu, 2015). Finally, it should be noted that although a multivariate normal distribution for item parameters is assumed to account for WICD in this study, caution should be exercised when claiming universal validity for such assumption.
Supplemental Material
Online_Appendix – Supplemental material for Bayesian DINA Modeling Incorporating Within-Item Characteristic Dependency
Supplemental material, Online_Appendix for Bayesian DINA Modeling Incorporating Within-Item Characteristic Dependency by Peida Zhan, Hong Jiao, Manqian Liao and Yufang Bian in Applied Psychological Measurement
Footnotes
Acknowledgements
The authors thank Dr. Hua-Hua Chang, Dr. Jimmy de la Torre, and two anonymous reviewers for their constructive comments on earlier drafts of this article. In addition, the first author would also like to express his heartfelt thanks to Dr. Wen-Chung Wang for his selfless guidance in the past few years.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This study was partly supported by Natural Science Foundation of Zhejiang Province, China (Grant No. LY16C090001).
Supplemental Material
Supplemental material is available for this article online.
Notes
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
