Abstract
The Asymptotic Classification Theory of Cognitive Diagnosis (ACTCD) developed by Chiu, Douglas, and Li proved that for educational test data conforming to the Deterministic Input Noisy Output “AND” gate (DINA) model, the probability that hierarchical agglomerative cluster analysis (HACA) assigns examinees to their true proficiency classes approaches 1 as the number of test items increases. This article proves that the ACTCD also covers test data conforming to the Deterministic Input Noisy Output “OR” gate (DINO) model. It also demonstrates that an extension to the statistical framework of the ACTCD, originally developed for test data conforming to the Reduced Reparameterized Unified Model or the General Diagnostic Model (a) is valid also for both the DINA model and the DINO model and (b) substantially increases the accuracy of HACA in classifying examinees when the test data conform to either of these two models.
Cognitive diagnosis models (CDMs; Rupp, Templin, & Henson, 2010) of educational test performance decompose an examinee’s overall ability into a set of specific discrete skills, called attributes, each of which he or she may or may not have mastered, thereby providing a detailed description, or attribute profile, of his or her strengths and weaknesses in the ability domain of the test. The entire set of possible attribute profiles for a given test defines classes of intellectual proficiency to which examinees can be assigned.
Current methods of fitting CDMs to educational test data generally use maximum likelihood estimation (MLE) procedures such as Expectation Maximization (EM) or Markov chain Monte Carlo (MCMC) to estimate model parameters that are then used to assign examinees to proficiency classes. These procedures often encounter difficulties in practice. For example, to obtain reliable parameter estimates, MLE procedures typically require large samples of examinees that may not be available in small- or medium-sized testing programs. In addition, the iterative MLE procedures are sensitive to the starting values, and do not guarantee optimal solutions despite often consuming considerable amounts of computer time. But perhaps most important, MLE procedures are vulnerable to the misspecification of the CDM supposedly underlying the data (recall that the true model is never known). Model misspecification causes examinees to be assigned to proficiency classes to which they do not belong.
In response to these difficulties, a number of researchers (Ayers, Nugent, & Dean, 2008; Chiu, 2008; Chiu & Douglas, 2013; Chiu, Douglas, & Li, 2009; Park & Lee, 2011; Willse, Henson, & Templin, 2007) have explored the potential of nonparametric classification techniques as heuristic or approximate methods for assigning examinees to proficiency classes. (A heuristic uses clever computational shortcut strategies to obtain a solution that is very close, if not identical, to the optimal solution.) Software for implementing these techniques can be developed from efficient cluster analysis programs that are readily available in the major statistical packages.
The Asymptotic Classification Theory of Cognitive Diagnosis (ACTCD; Chiu et al., 2009) provides the theoretical foundation for using hierarchical agglomerative cluster analysis (HACA) as a heuristic for assigning examinees to proficiency classes for educational data conforming to the Deterministic Input Noisy Output “AND” gate (DINA) model (Junker & Sijtsma, 2001; Macready & Dayton, 1977).
This article extends the ACTCD to item responses conforming to the Deterministic Input Noisy Output “OR” gate (DINO) model (Templin & Henson, 2006). It also demonstrates that an addition to the statistical framework of the ACTCD, augmented attribute sum-score profile (Chiu, 2008), which was originally developed to accommodate item responses conforming to the Reduced Reparameterized Unified Model (Reduced RUM; Hartz, 2002; Hartz & Roussos, 2008) or the General Diagnostic Model (GDM; von Davier, 2005, 2008), is valid also for the DINA model and the DINO model and increases the accuracy of HACA in classifying examinees when the item responses conform to either of these two models. A brief review of relevant definitions and key technical concepts precedes the theoretical development. The results of two simulation studies using item responses generated from the DINO model and the DINA model and an application of the studied methods to a real-world data set are then reported to provide empirical support for these extensions of the ACTCD. Finally, the “Discussion” section summarizes some of their practical implications.
Technical Background
Models for Cognitive Diagnosis
Let Yij
be the observed response of examinee i, i = 1, . . ., N, to binary item j, j = 1, . . ., J. Consider N examinees who belong to M distinct latent classes of intellectual proficiency. CDMs constrain the relation between the observed item response and the latent variable, so that the mastery of cognitive attributes characteristic for distinct latent proficiency classes determines the observed response (correct or incorrect) to a test item. Suppose that K latent binary attributes constitute a certain ability domain; there are then 2
K
distinct attribute profiles composed of these K attributes representing M = 2
K
distinct latent proficiency classes. Note that an attribute profile for a proficiency class can consist of all zeroes, because it is possible for an examinee not to have mastered any attributes at all. Let the K-dimensional vector,
Consider a test of J items assessing ability in the domain. Each individual item j is associated with a binary attribute profile that specifies or constrains the particular skills required for answering it correctly. Given K attributes, there are at most 2
K
− 1 distinct item-attribute profiles; item-attribute profiles that consist entirely of zeroes are inadmissible, because they correspond to items whose answers require no skills at all. The entire set of constraints specifying the associations between J items and K attributes constitutes the Q-matrix,
Nonparametric Classification Using HACA Adapted for Cognitive Diagnosis
For input into HACA, Chiu and collaborators (Chiu, 2008; Chiu et al., 2009) aggregated each examinee’s set of item responses,
Many techniques exist for the nonparametric classification of a set of objects (such as the rows of a matrix). The principal objective shared by all of these techniques is to identify maximally homogeneous clusters that are maximally separated. (The classic reference is Hartigan, 1975.) To adapt HACA for cognitive diagnosis requires first transforming the N×K matrix of the examinees’ attribute sum-score profiles into an N×N symmetric matrix of inter-examinee Euclidean distances. Popular HACA algorithms include single-, complete-, and average-link clustering (Johnson, 1967) and Ward’s (1963) minimum-variance method. The link algorithms all sequentially merge or agglomerate examinees (or groups of examinees) closest to each other at each step into an inverted tree-shaped hierarchy of nested clusters that represents the relationships between examinees. The inter-examinee distances are updated after each merger to reflect the latest status of examinee/cluster cohesion as input for the next agglomeration step; the specific method of updating these distances distinguishes the various link algorithms. Ward’s (1963) method differs from the link algorithms in that it updates the within-cluster sum of squared errors rather than the inter-examinee distances.
The ACTCD Developed for the DINA Model
The development of the ACTCD was inspired by a search for legitimate nonparametric classification heuristics for item responses conforming to the DINA model (Junker & Sijtsma, 2001; Macready & Dayton, 1977), a conjunctive non-compensatory CDM. The conjunction parameter η
ij
indicates whether examinee i has mastered all the attributes needed to answer item j correctly; η
ij
is defined as
The ACTCD (Chiu, 2008; Chiu et al., 2009) consists of three lemmas, each of which specifies a condition necessary for a consistency theorem to hold. (For formal proofs, consult Chiu et al., 2009.) Like MLE procedures, heuristic classification techniques using attribute sum-score profiles as input require that the Q-matrix be complete. A Q-matrix is said to be complete if it allows identification of all possible attribute profiles. This definition translates into the formal expression tailored to the DINA model,
Lemma 1 of the ACTCD states that, for item responses conforming to the DINA model,
Let
Proof (Modified):
(⇒) Assume that some
(⇐) Assume that K of the J rows of
Let
Thus, Lemma 2 justifies using
If a finite mixture model with M latent classes underlies the item responses, then Lemma 3 of the ACTCD establishes that complete-link HACA accurately assigns examinees to their true proficiency classes, provided that the resulting classification hierarchy is “cut” at M clusters (Chiu et al., 2009).
Building on these three lemmas, the consistency theorem of classification states that, for item responses conforming to the DINA model, the probability that HACA assigns examinees correctly to their true proficiency classes based on
Extension of the ACTCD to the DINO Model
The DINO model (Templin & Henson, 2006) is a disjunctive CDM. Define the disjunction parameter
Lemmas 1 and 2 of the ACTCD for the DINA model specify regularity conditions that must be satisfied by any CDM for the consistency theorem of classification to hold. Hence, extension of the ACTCD to the DINO model requires proving that it is covered by these two lemmas. These proofs rely on the “duality” of the DINO model and the DINA model (Liu, Xu, & Ying, 2011, 2012).
The Duality of the DINA Model and the DINO Model
As Y. Liu et al. (2011) discovered and proved, the DINA model and the DINO model are technically identical under certain transformations of (a) the examinees’ attribute profiles, (b) their observed item scores, and (c) the model parameters. This means that one model can be expressed in terms of the other and both models can be fitted by the same software. (As an aside, note that the characterization of the special relation between the DINA model and the DINO model as “dual” deviates from the well-defined meaning of this term in operations research; for details, consult Papadimitriou & Steiglitz, 1998.) The proof of the duality of the DINA model and the DINO model presented here is a modification of the original proof by Y. Liu et al. (2011) tailored to the ACTCD.
Transformation of Examinees’ Attribute Profiles
Consider the attribute profile
where
(for emphasis, the slipping and guessing parameters of the DINO model are henceforth denoted by ssj and ggj , respectively).
Transformation of Examinees’ Observed Item Response Scores
Let
because
Transformation of the Model Parameters
In replacing
by setting gj = ssj and sj = ggj . This completes the proof.
Results Based on Equation 2
First, the observed item responses,
Proof:
(⇒) The proof of Lemma 1 for the DINA model demonstrated that if
(recall that
(⇐) Assume that
Proof:
Lemma 2 for the DINA model states
and
In summary, Lemma 1 (DINO) and Lemma 2 (DINO) can be used with Lemma 3 to prove that the consistency theorem of classification holds when the item responses conform to the DINO model. The proof is omitted because it is identical to that given by Chiu et al. (2009).
The Augmented Attribute Sum-Score Profile for the DINA Model and the DINO Model
The augmented attribute sum-score profile,
Note that
Does the enhancement of well-separated proficiency classes by
Proof:
Lemma 2 in Chiu et al. (2009) states that, if the item responses conform to the DINA model and
Lemma 2 (DINO:
In summary, the augmented attribute sum-score profile,
Simulation Studies
Two simulation studies were conducted to provide empirical support for the theoretical developments of the ACTCD proposed for the DINO model and the DINA model by demonstrating that (a) the consistency theorem of classification is valid for item responses conforming to the DINO model as well as for those conforming to the DINA model—that is, the accuracy of HACA in assigning examinees to their true proficiency classes increases with the number of test items; and that (b) use of the augmented attribute sum-score profile,
Study 1
Experimental Design and Simulation of Item Responses
Examinee item responses conforming to the DINO model and to the DINA model were simulated according to the method described by Chiu et al. (2009). The experimental design included five variables: (a) the number of examinees, N = 100, 500; (b) the number of attributes, K = 3, 4; (c) the number of items, J = 20, 40, 80; (d) the distribution (discrete uniform, multivariate normal) underlying the attribute profiles; and (e) the distribution of the slipping and guessing item parameters, sj
and gj
(continuous uniform
Each examinee’s attribute profile was generated from either a discrete uniform distribution (i.e., the attribute profile for each proficiency class,
The simulated item responses, Yij
, were sampled from a Bernoulli distribution with π
ij
defined by the IRF of the respective model and the slipping and guessing item parameters, sj
and gj
, drawn from the continuous uniform distributions specified previously. The values of the slipping and guessing parameters reflect the degree of error perturbation of the item responses, with
Completely crossing the levels of all five variables yielded for the DINO model and the DINA model experimental designs each with 2 × 2 × 3 × 2 × 2 = 48 cells. Twenty-five data sets were generated for each cell, for a total of 2,400 simulated data sets.
Each simulated examinee’s set of item responses was aggregated into an attribute sum-score profile and an augmented attribute sum-score profile to serve as input to complete-link, average-link, and Ward’s method HACA; single-link HACA was not used because its well-known susceptibility to “chaining” (observations being added one-by-one to a cluster, with the resulting clusters resembling long strings or chains) makes it an inappropriate choice in this context. All three clustering methods are implemented in the hclust routine in R. The cluster tree produced by each of the three clustering methods was “cut” at level M = 2 K , the true number of proficiency classes.
Evaluation of the Clustering Results
For simulated item responses, the true proficiency-class memberships of examinees are known and provide a standard for evaluating the accuracy of the results of the cluster analyses (i.e., to which extent examinees are assigned to their true proficiency classes). However, because the clustering methods do not label the clusters representing the proficiency classes with their respective attribute profiles (i.e., HACA cannot identify the different proficiency classes in terms of
Study 1: Results
Tables 2 and 3 (see the online appendix) present the results of the first simulation study for each of the two distributions (discrete uniform, multivariate normal) used to generate examinees’ attribute profiles for item responses conforming to the DINO model; Tables 4 and 5 (see the online appendix) present analogous results for item responses conforming to the DINA model. Each table is organized, so that each of its rows reports the simulation results for a specific combination of the variables N, J, and K, and the distribution of the slipping and guessing parameters. The columns of each table contain the simulation results for each of the three clustering methods when either
The simulation results for the DINO model and the DINA model are very similar, so they are summarized together here.
In accordance with the ACTCD, a larger number of test items generally leads to a more accurate assignment of examinees to their true proficiency classes by all three clustering methods, regardless of whether
The performance of all three clustering methods in assigning examinees to their true proficiency classes is better when
The distribution underlying the attribute profiles appears to have a differential effect on the performance of the clustering methods: With one exception, the performance of Ward’s method HACA is better when attribute profiles are underlain by a discrete uniform distribution rather than by a multivariate normal distribution, and with one exception, the performance of average-link HACA is better when attribute profiles are underlain by a multivariate normal distribution rather than by a discrete uniform distribution.
When either the number of attributes or the level of error perturbation increases, the accuracy of the assignment of examinees to their true proficiency classes deteriorates—this is most obvious under those conditions where J = 20.
The number of examinees seems to have no effect on the accuracy of the assignment of examinees to their true proficiency classes.
When the level of error perturbation is moderate, or the number of test items is large, the ARIs computed for Ward’s method HACA and average-link HACA are fairly comparable with those obtained using the MLE method.
The results obtained using the MLE method appear much less sensitive to variations in the number of attributes, the number of test items, the distribution of attribute profiles, and the level of error perturbation. This result was expected, of course, because the MLE method for the DINO model or the DINA model should outperform nonparametric heuristic classification techniques when used to analyze item responses generated from the corresponding model.
Study 2
The design of Study 2 is identical to that of Study 1, except for one modification: Only attribute profiles with an underlying multivariate normal distribution were considered. Thus, completely crossing the levels of all five variables yielded for the DINO model and the DINA model experimental designs each with 2 × 2 × 3 × 1 × 2 = 24 cells. Twenty-five data sets were generated for each cell, for a total of 1,200 simulated data sets. Each simulated examinee’s set of item responses was aggregated into an attribute sum-score profile and an augmented attribute sum-score profile to serve as input to the same clustering methods, as were used in Study 1. The recovery of the true proficiency classes was assessed by the ARI. Recall that the main purpose of Study 2 was to compare the performance of HACA with the exact EM-based MLE method when the true model underlying the data is unknown and has been misspecified (i.e., data truly conforming to the DINA model are fitted with the DINO model and vice versa).
Study 2: Results
The results of Study 2 are summarized in Tables 6 and 7 (see the online appendix), which follow the exact same lay-out like the previous tables. However, now, the last column in each table, labeled DINA-EM and DINO-EM, respectively, presents the results when the simulated item responses were analyzed based on the misspecified CDM using the corresponding MLE method. Each entry in the tables reports the average (mean) ARI for the 25 data sets generated for a cell in the experimental design. The results regarding the performance of the different HACA methods match those from Study 1 and are, therefore, not further commented on. The most remarkable finding of Study 2 concerns the performance of the HACA methods in comparison with the EM-based MLE methods when the CDM has been misspecified: For all 48 cells, the heuristic HACA methods outperformed the exact EM-based MLE methods in the recovery of the true proficiency classes as measured in terms of the average ARI scores.
Practical Application
The real-world fraction–subtraction data set (K. K. Tatsuoka, 1984) is one of the most thoroughly studied in cognitive diagnosis research (e.g., de la Torre, 2008, 2009; de la Torre & Douglas, 2008; DeCarlo, 2011; Mislevy, 1996; C. Tatsuoka, 2002). A subset of this data set consisting of the responses of 536 middle school students to a collection of 15 test items (de la Torre, 2008) is used here to compare the ability of the three clustering methods to assign examinees to proficiency classes when
Twenty-four proficiency classes were identified using DINA-EM; these 24 classes are used here as the standard of comparison for evaluating the proficiency-class assignments produced by the three clustering methods. The results are reported in Table 9 of the online appendix. The ARI for complete-link HACA when
Still, the relatively low ARI scores obtained for the HACA clustering methods when used with the fraction–subtraction data call for some additional explanations. First, note that with five attributes, there are 25 = 32 possible proficiency classes. However, the Q-matrix of the fraction–subtraction data is incomplete. Thus, not all of the 32 different proficiency classes are identifiable. (Recall that completeness of the Q-matrix is a universal requirement of cognitive diagnosis. If the Q-matrix is not complete, then the identification of all 2
K
= M different proficiency classes is impossible regardless of whether MLE methods or nonparametric classification techniques are used.) It can be shown that the Q-matrix of the fraction–subtraction data allows only the identification of 10 out of the 32 proficiency classes. But the benchmark for computing the ARI scores for evaluating the performance of
Discussion
The ACTCD developed by Chiu et al. (2009) proves that for educational test item responses conforming to the DINA model, the probability that HACA assigns examinees to their true proficiency classes approaches 1 as the number of test items increases. This article proves that the ACTCD also covers item responses conforming to the DINO model.
HACA requires as input a statistic for an examinee’s attribute profile,
The empirical performance of the two statistics for
What recommendations can be made to the educational practitioner? Recall that originally nonparametric classification techniques were proposed as a heuristic alternative to the exact MLE procedures for assigning examinees to proficiency classes because these methods often encountered difficulties in practice and suitable computer software was mostly unavailable. But recently, Robitzsch et al. (2014) have developed the package CDM that provides an implementation of the EM algorithm for fitting the DINO model and the DINA model in R. Thus, educational practitioners should use these exact MLE methods whenever possible because they typically outperform HACA in assigning examinees to proficiency classes under regular conditions—that is, when the sample sizes of examinees are sufficient and the CDM underlying the data has been correctly identified.
However, nonparametric classification techniques can be useful in situations where MLE methods fail or are difficult to implement—most notably, when the true CDM underlying the data is unknown. The second simulation study demonstrates that if the CDM has been misspecified, then across all experimental conditions, the three HACA methods obtain substantially better recovery of examinees’ true proficiency-class membership than the exact EM-based MLE procedure using the wrong model. One could argue, of course, that the (most likely) correct model could always be identified by fitting two different CDMs to the data set in question, and then choose the one with the better fit. But what to do if test item responses conform to multiple CDMs (a scenario that was not considered here due to space limitations)? In this situation, nonparametric classification techniques can offer a solution to the task of assigning examinees to proficiency classes.
Finally, it should be recalled that the examinee clusters obtained from nonparametric classification methods serve as proxies for the proficiency classes. But nonparametric classification methods cannot estimate the attribute profiles underlying the clusters. Hence, the clusters must be interpreted or labeled—that is, their underlying attribute profiles must be reconstructed from the chosen input data, which can be tedious if the number of examinees is large. For
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
