Abstract
The fit of cognitive diagnostic models (CDMs) to response data needs to be evaluated, since CDMs might yield misleading results when they do not fit the data well. Limited-information statistic M2 and the associated root mean square error of approximation (RMSEA2) in item factor analysis were extended to evaluate the fit of CDMs. The findings suggested that the M2 statistic has proper empirical Type I error rates and good statistical power, and it could be used as a general statistical tool. More importantly, we found that there was a strong linear relationship between mean marginal misclassification rates and RMSEA2 when there was model–data misfit. The evidence demonstrated that .030 and .045 could be reasonable thresholds for excellent and good fit, respectively, under the saturated log-linear cognitive diagnosis model.
Introduction
Cognitive diagnostic models (CDMs) are developed to provide fine-grained skills information for well-directed interventions. In fact, the chosen CDM, as a statistical model, needs to be evaluated regarding its validity with absolute model–data fit testing, in case the results are misleading when the model does not fit the data well (Rupp, Templin, & Henson, 2010). Generally speaking, the existing absolute model–data fit tests can be divided into item and test level. While item-level statistics have been the focus in recent work (e.g., Chen, de la Torre, & Zhang, 2013; de la Torre & Lee, 2013; Kunina-Habenicht, Rupp, & Wilhelm, 2012), test-level goodness-of-fit evaluation has been relatively stagnant due to the fact that the full-information statistics χ2 and G2 are virtually useless in practice (e.g., Maydeu-Olivares, 2013) and other suitable alternatives remain underdeveloped.
The Pearson χ2 test statistic and the likelihood ratio test statistic G2 as full-information statistics are routinely used in the commercial software Mplus 7 (Muthén & Muthén, 2012). Both of them are computed from all possible response patterns (full contingence table). χ2 and G2 can be effective when all expected frequencies are large (the usual rule of thumb is that the expected frequency should exceed 5 in each cell). However, when the number of items is large or the number of respondents is small, so that the contingency table for possible responses is sparse, then χ2 and G2 cannot be trusted to test for lack of fit because the empirical Type I error rates are often different from the expected values under their reference asymptotic distributions (e.g., Maydeu-Olivares & Joe, 2005). Although alternatives to χ2 and G2 have been proposed to assess test-level fit (Rupp et al., 2010), much work remains to be done. For example, the Monte Carlo resampling technique (Templin & Henson, 2006) needs to obtain a stable empirical p value with multiple simulated data sets, which leads to an excessive amount of time. And the Bayesian posterior predictive model checking method (Sinharay, 2006; Sinharay & Almond, 2007) was found to be conservative and computationally intensive.
In practice, the limited-information method (Reiser, 1996; Reiser & Lin, 1999) is practicable since it uses the summary characteristics of the full contingence table (Cai & Hansen, 2013; Cai, Maydeu-Olivares, Coffman, & Thissen, 2006; Maydeu-Olivares & Joe, 2005). The limited-information statistics use only low-order marginal information in the contingency table to evaluate the model–data fit. The up-to-order r marginal information is used to develop the limited-information statistics Mr. The results showed that the M2 statistic using the univariate and bivariate margins is enough for routine applications (Cai et al., 2006; Maydeu-Olivares & Joe, 2005). The model frequently does not exactly fit the data in practice; therefore, it is important to assess how well the model reflects reality (Brown, 2006; Hu & Bentler, 1998). Further, Maydeu-Olivares and Joe (2014) proposed up-to-order r root mean square error of approximation (RMSEA) fit indices RMSEA r to assess the approximate goodness of fit. RMSEA r can be estimated by using Mr, and Maydeu-Olivares and Joe (2014) recommended using RMSEA2 to assess the approximate goodness of fit of the models for routine applications. Jurich (2014) conducted a small-scale simulation study to examine the statistical properties of M2 under the log-linear cognitive diagnosis model (LCDM; Henson, Templin, & Willse, 2009); the results showed that M2 had good Type I error control under the null conditions and high power to detect model misspecifications. However, no studies have systematically examined the performance of M2 and RMSEA2 in assessing model–data fit and degrees of model misfit, under CDMs.
This article aims to extend M2 and RMSEA2 in item response theory to those of CDM by taking the LCDM as an example. Firstly, we introduce LCDM and its marginal likelihood as the basis of computing CDMs’ goodness-of-fit statistics. Secondly, we describe the idea of full- and limited-information model–data fit and approximate goodness-of-fit index. Thirdly, we present the results of simulation studies conducted to systematically evaluate the statistical properties of M2 and the performance of RMSEA2. Finally, as an empirical illustration, we present an analysis of fraction subtraction data originally reported by Tatsuoka (1990). This data set has been a commonly used example in cognitive diagnosis literature (e.g., DeCarlo, 2010; de la Torre, 2009, 2011; de la Torre & Douglas, 2008; Henson et al., 2009).
LCDM and Its Marginal Distribution
Suppose that a CDM test of J dichotomous scored items is administrated to N respondents, in which K binary attributes are diagnosed. Let the ith examinee’s response vector be denoted by Bernoulli variables
Statistically, the LCDM is essentially an item-based conditional probability of a correct response of the ith examinee for item j
where
where
Further, assuming local independence, the likelihood function for examinee i, conditional on the attribute mastery pattern
The marginal probability of the ith examinee’s response vector can be written as:
In the above expression,
To accommodate this sum-to-one constraint, the following expression is used:
and one identifying constraint has to be imposed on the model parameter
The model-predicted probability
The Limited Information as an Alternative of Full Information
Full-Information Statistics
The exact model fit is used to determine whether CDMs can precisely predict the phenomena provided by the data (Sinharay & Almond, 2007). In other words, the model-fit statistics compare the following model-predicted probabilities and the observed proportions associated with the contingency table:
Here,
The aim of the exact model-fit statistics is to test the simple or composite null hypothesis. We focused on the composite context because it is encountered more frequently in psychological and educational assessments. The composite null hypothesis of the full information is:
where
compare the contingency table using full information from all response patterns; degrees of freedom of χ2 and G2 are
According to Maydeu-Olivares and Joe (2014), once the model parameters
A 90% confidence interval (CI) for RMSEA J is (Browne & Cudeck, 1993):
where
Limited-Information Statistics
By contrast, limited-information statistics were developed as an alternative to address the frequently encountered sparseness phenomenon by compressing the contingency table. Specifically, the H cells of marginal probabilities associated with the full response patterns are significantly reduced to smaller ones. Essentially, the limited-information statistics are used as a statistical concept of lower-order margins.
Symbolically, let the first-order marginal probabilities
is the marginal probability of correctly answering the jth item.
contained in
For example, if 3 items were designed for a diagnostic test, then the lower-order marginal probabilities would be:
When the model parameters are estimated from the data under the necessary regularity assumptions on the model, the ML estimator
where
Limited-information statistics Mr based on up-to-order r marginal residuals can be written as follows (Maydeu-Olivares & Joe, 2005):
where
Evaluate the Fit of CDMs Using M2 Statistic
In CDMs, it is often the case that both the item parameters and the probabilities of attribute mastery pattern are estimated simultaneously from the data using ML procedure. To test the composite null hypotheses, the limited-information statistic M2 based upon up-to-order two residuals has been recommended by researchers (Cai & Hansen, 2013; Maydeu-Olivares & Joe, 2005). Therefore, the composite hypothesis is:
After the model parameters have been estimated from the data, the model-predicted marginal probabilities of up-to-order two response patterns are:
The observed proportions associated with the univariate and bivariate model-predicted marginal probabilities are:
The asymptotic covariance matrix
where
and
are the marginal probabilities of correctly answering the a th, b th, and c th item, respectively:
are the marginal probabilities of correctly answering the ath and bth item pair and the cth and dth item pair, respectively;
is the marginal probability of correctly answering the ath item, bth item, and cth item simultaneously:
is the marginal probability of correctly answering the ath item, bth item, and cth item, and dth item simultaneously.
The
where
The generic entries of
and
The elements of
and
According to Equation 16, M2 for CDMs can be constructed:
where
And a 90% CI for RMSEA2 is:
here,
Simulation Studies
A series of Monte Carlo simulations were conducted using R software (R Core Team, 2014). R code for computing M2 and RMSEA2 of the LCDM framework and generalized deterministic inputs, and the noisy “and” gate model (de la Torre, 2011) framework are all available upon request from the first author. For each simulation, two sample sizes, N = 1,000 and N = 5,000, were considered. There were 500 replications in each condition. The data were all generated using the saturated LCDM model. Table 1 summarizes the true item parameters. Table 2 presents the summary of the true probabilities of responding correctly to items that required one, two, or three attributes for examinees that had mastered zero, one, two, and three attributes.
True Item Parameters
Note. Kj is the number of the required attributes by the jth item.
True Item Response Probabilities
Note. Kj is the number of required attributes by the jth item.
The purpose of the first study was to investigate the performance of χ2, G2, and M2 under the null condition of correct model specifications. Simulation 1 was conducted to examine the Type I error rates of χ2, G2, and M2, when the test lengths were relatively small and to examine the Type I error rates of M2, when the test lengths were relatively large. In the second study, simulation 2 was conducted to illustrate the power of M2, when the Q-matrix of the model was misspecified. The purpose of the third study was to explore the performance of RMSEA2 under more comprehensive Q-matrix misspecification conditions. Simulation 3 was conducted to examine the performance of RMSEA2 systematically.
Type I Error Rates
In this section, a simulation was performed to examine the empirical Type I error rates of χ2, G2, and M2. We compared the Type I error rates of the three statistics under nonsparse and sparse conditions. In order to test the robustness of the M2 statistic, a multidimensional normal distribution was specified for the attributes; the mean vectors were randomly chosen from the uniform distribution
Simulation conditions
For this simulation, three factors were manipulated: Four Test Lengths × Two Sample Sizes × Three Statistics. (a) Test length: 6, 15, 30, and 50 items; (b) sample size: 1,000 and 5,000 examinees; (c) statistic: χ2, G2, and M2. When J = 30, the number of cells in the contingency table equaled more than 1 billion, so χ2 and G2 were not considered in conditions with J = 30 and J = 50. There were a total of 16 data generating conditions that were considered.
For conditions of 6 items, there were two latent attributes. The Q-matrix for 6 items is presented in Table 3. For the 15, 30, and 50 item conditions, there were five attributes, and the Q-matrices are presented in Appendix Table A1. For each Q-matrix, the number of items for measuring each attribute was equal, and three attributes were required by items at most.
Q-Matrix for J = 6, K = 2
Results
Since we assumed that the attributes were moderately or highly correlated, some of the expected frequencies of the response patterns were certainly smaller than
Type I Error Rates for Nonsparse Contingency Tables (J = 6, K = 2)
Note. N = the sample size; df = the degrees of freedom; Var/2 = variance/2.
When 15 items were considered in a test, the number of all response patterns was
Type I Error Rates for Sparse Contingency Tables (J = 15)
Note. n = the sample size; df = the degrees of freedom; Var/2 = variance/2.
We also investigated the performance of M2, under large test length conditions. As is shown in Table 6, when J = 30 and J = 50, the Type I error rates of M2 were reasonably close to the nominal levels. Integrating Tables 4, 5, and 6, the empirical Type I error rates of M2 were found to be stable and accurate.
Type I Error Rates for Sparse Contingency Tables (J = 30, 50)
Note. J = the test length; N = the sample size; df = the degrees of freedom; Var/2 = Variance/2.
Power to Detect Model Misspecification
Some studies have revealed that there are many sources of model–data misfit, such as CDM misspecification and Q-matrix misspecification (Chen et al., 2013; de la Torre & Lee, 2013; Kunina-Habenicht et al., 2012). Kunina-Habenicht, Rupp, and Wilhelm (2012) have shown that if the data were generated by the saturated model, the exclusion of the interaction effects did not have a noticeable impact on the marginal correct classification rates; however, the misspecification of Q-matrix had a significant effect on classification accuracy. In this section, in order to illustrate the performance of M2, under the incorrect specification of the Q-matrix conditions, a simulation was conducted to examine the power of M2. In this simulation, the data were generated using LCDM and the Q-matrices are provided in Appendix Table A1.
Simulation conditions
The following four factors were manipulated in the simulation: Three Test Length Conditions × Two Q-matrix Misspecification Types × Two Attribute Correlation Conditions × Two Sample Sizes. (a) Test length: 15, 30, and 50 items; (b) Q-matrix misspecification type: the random design and the random balance design; (c) attribute correlation: .5 and .8; and (d) sample size: 1,000 and 5,000 examinees. A total of 24 data generating conditions were examined.
For Q-matrix misspecification, two levels were considered: The first was that 20% of the Q-matrix elements were randomly assigned by limiting the maximum number of required attributes to 3 and the minimum to 1, we will refer to this as a random design. In the second condition, 20% of the elements of the Q-matrix were misspecified; specifically, the overspecified items were chosen randomly from the items measuring only one attribute, and both under- and overspecified items were chosen randomly from items measuring three attributes, which we will refer to as a random balance design. The random balance design was driven by the Q-matrix misspecification in the previous studies (Chen et al., 2013), and the summary of the random balance design of Q-matrix misspecification is shown in Table 7. The misspecified Q-matrix was generated for each replication. Kunina-Habenicht et al. (2012) reported that in their mathematics and network engineering assessments, the correlations between attributes were all higher than .5 and, some even exceeded .8, and in their simulation studies with LCDM, the correlations between attributes were set to ρ = .5 and ρ = .8. In order to compare the performance of M2 under different attribute correlation conditions, we followed the simulation design used by Kunina-Habenicht et al. (2012), and the strength of the attribute tetrachoric correlations were set to ρ = .5 and ρ = .8 to reflect the low and high correlations between attributes.
The Random Balance Design of Q-Matrix Misspecification
Note. Kj is the number of required attributes by the jth item.
Results
The empirical rejection rates of M2 were 1 under the random balance design, and the rejection rates were all above .990 when J = 30 and J = 50. Thus, only the empirical power results of M2 for J = 15 under the random design are presented in Table 8. As shown in Table 8, the statistical power for ρ = .5 was higher than ρ = .8, and with the increase of the sample size, the power of M2 was also strengthened. In addition, the results revealed that M2 was a powerful tool for detecting model–data misfit.
The Empirical Rejection Rates of M2 for J = 15 Under the Random Design
Note. ρ = the attribute correlation. N = the sample size.
The Performance of RMSEA2
Overall goodness-of-fit index provides us the information regarding the degree of discrepancy between observed and expected responses. However, it is realistic to expect that CDMs might not always fit the data exactly in practice. As illustrated in the section on power to detect model misspecification, the statistic M2 has high power to detect model misspecification. Therefore, it is important to provide an effect-size estimator to evaluate whether model–data fit is good enough for practitioners or not. In order to systematically examine the performance of RMSEA2 in LCDM, Simulation 3 was conducted with particular attention to the relationship between the RMSEA2 and the mean marginal misclassification rates (MMMRs) under more comprehensive Q-matrix misspecification conditions. Specifically, three levels of the percentage of Q-matrix misspecification: 10% (small), 20% (medium), and 30% (large) and two types of Q-matrix misspecification: Random design and random balance design were manipulated.
Simulation conditions
The following five factors were manipulated in the simulations: (a) test length: 15, 30, and 50 items; (b) percentage of Q-matrix misspecification: 10%, 20%, and 30%; (c) Q-matrix misspecification type: the random design and the random balance design; (d) attribute correlation: .5 and .8; (e) sample size: 1,000 and 5,000 examinees. Since the required numbers of item were larger than the test length, the 30% random balance design condition was not included in Simulation 3. There were a total of 60 conditions.
Results
The primary focus was to examine the relationship between RMSEA2 and MMMR. As revealed by Figure 1, there is a positive relationship between RMSEA2 and MMMR. Figure 1 also demonstrated that most of the MMMR were greater than .160 when RMSEA2 ≥.045, and most of the MMMR were less than .05, when RMSEA2 ≤.030. Thus, we propose that .045 might be a reasonable cutoff criterion of good fit for LCDM, and .030 might be the proper cutoff for excellent fit. Figure 1 also reveals that the types of Q-matrix misspecification had a noticeable impact on the relationship between RMSEA2 and MMMR.

Scatter plot of root mean square error of approximation (RMSEA2) and mean marginal misclassification rates.
We also found that the MMMR could be well predicted from RMSEA2 values when a linear regression model was used. The coefficients of determination R2 for the random design and random balance design were .830 and .880, respectively. The correlation was higher in the former than that in the latter. This is probably because there was at least one attribute being correctly specified for each item in the latter.
In Figure 2, in order to systematically illustrate the relationship between MMMR and RMSEA2 under different conditions, MMMR is displayed as a function of RMSEA2 value, attribute correlation, test length, and Q-matrix misspecification type. The sample sizes did not have a noticeable impact on the wrong classification rates, thus they were not included in Figure 2. In terms of the attribute correlations, the classification inaccuracy and RMSEA2 were higher in conditions with ρ = .5 than that with ρ = .8. For Q-matrix misspecification types, the classification accuracy was higher under the random balance design than those under random design. When ρ = .8, the classification accuracy increased with test length. The same trend can also be observed in conditions with random balance design and ρ = .5. It’s important to point out that even though the rejection rates were 1, the classification accuracy was above .950, when J = 50, and ρ = .8, under the random balance design. We also found that the impact of the percentage of Q-matrix misspecification on the MMMR depended on test length, Q-matrix misspecification type, and attribute correlation.

Scatter plot of root mean square error of approximation (RMSEA2) as a function of attribute correlation, test length, and Q-matrix misspecification type.
Empirical Illustration
We used fraction subtraction data (Tatsuoka, 1990) to illustrate the usefulness of M2 and RMSEA2. Although many researchers have studied the data (e.g., DeCarlo, 2010; de la Torre, 2009, 2011; de la Torre & Douglas, 2008; Henson et al., 2009), none could come to an agreement on which model and the corresponding Q-matrix fit the data exactly; in this study, for illustrative purposes, we used M2 and RMSEA2 statistics to investigate the performance of CDMs among the deterministic inputs, noisy “and” gate model (DINA; de la Torre, 2009; Junker & Sijtsma, 2001), the compensatory reparameterized unified model (C-RUM; Hartz, 2002), and LCDM regarding fraction subtraction data.
A subset of this data set containing the responses of 536 middle school students to 10 fraction subtraction items was analyzed as an example (DeCarlo, 2010; Henson et al., 2009). The items and Q-matrix we implemented were proposed by Henson et al. (2009) and are presented in Table 9. We use the item numbering originally used by Henson et al. (2009). Item 8 was removed from the original data set, as Henson et al. (2009) mistakenly took the test Item 8
The Q-Matrix for Fraction Subtraction Data.
Note. We use the items numbers from Henson, Templin, and Willse (2009).
In order to resolve the identification problems and ensure the parameters were reasonable, the item parameters were all estimated by using MMLE algorithm, using flexMIRT (Cai, 2013), and a β (1.6, 1) prior distribution on the item parameters of LCDM and C-RUM was adopted (Bock, Gibbons, & Muraki, 1988). The item parameter estimates of LCDM, DINA, and C-RUM are shown in Table 10. Table 11 provides the values of M2 and RMSEA2 for these models.
Item Parameter Estimates for Fraction Subtraction Data
Note.
M2 and RMSEA2 Statistics for Fraction Subtraction Data Set
Note. RMSEA = root mean square error of approximation. CI = confidence interval. LCDM = log-linear cognitive diagnosis model. DINA = deterministic inputs, noisy “and” gate model. C-RUM = compensatory reparameterized unified model; df = the degrees of freedom. p Values < .001 are reported as .000.
According to Table 11, the p value of LCDM was .786, which suggests that the saturated LCDM provided a good fit to fraction subtraction data. However, the p values of DINA and C-RUM were significantly smaller than .001, which suggests that the two models did not fit the data well. Following the rules of thumb for RMSEA2 presented in the section on the performance of RMSEA2, the DINA and C-RUM did not have a close fit to the data. Among the three fitted models, LCDM was the best choice.
Discussion
Since model–data misfit may yield misleading information, assessing the overall goodness-of-fit for CDMs is of great importance for practitioners. The existing methods for evaluating overall model–data fit under a sparse contingency table have their pros and cons. By contrast, the M2 statistic developed here hold an obvious predominance in computation efficiency and accurate empirical Type I error rates. In this article, we have conducted simulation studies to compare the performance of χ2, G2, and M2, under the null condition of correct model specifications and the empirical evidence suggested that χ2, G2, and M2 can be used to test the overall goodness of fit when the contingency table was not sparse; however, only M2 had accurate Type I error rates when the contingency table was sparse. We also have examined the empirical behavior of M2 under the incorrect specification of the Q-matrix conditions, and simulation results suggested that M2 was a powerful tool to detect the misspecification of the model. So we recommend the use of M2 statistic for assessing the overall exact model fit in CDMs.
The overall goodness-of-fit index only provided information about whether the CDMs and the corresponding Q-matrices fit the data or not. It is also important to assess the goodness of approximation of CDMs. After examining the performance of RMSEA2 under different conditions, especially the relationship between RMSEA2 and the MMMR, we have found significant correlations between them. The results of simulation studies provided evidence that the cutoff values .030 and .045 might be reasonable criteria for excellent and good fit for LCDM. The results of the simulation study also revealed that the variables of test length, percentage of Q-matrix misspecification, Q-matrix misspecification type, and attribute correlation had a noticeable impact on the classification accuracy and RMSEA2 values. More importantly, although the rejection rates were very high, the classification inaccuracy and RMSEA2 were very low in conditions with J = 50 and random balance Q-matrix misspecifications. We thus propose the use of RMSEA2 statistic for assessing the approximate fit in CDMs.
As a widely used parameter estimation method in CDMs (de la Torre, 2011; Robitzsch, Kiefer, George, & Uenlue, 2015), MMLE may yield implausible parameter values for CDMs. When this occurs, it is necessary to impose a reasonable prior on the item parameters (Mislevy, 1986). In the empirical illustration, in order to obtain reasonable item parameter values, model parameters were estimated using flexMIRT, which provides a flexible and practical model parameter estimation procedure in CDMs, and a β (1.6, 1) prior was imposed on the item parameters of LCDM and C-RUM. It is important to point out that these item parameter estimates were still asymptotically normally distributed (Mislevy, 1986). By using the parameter estimates, we assessed the overall goodness of fit and goodness of approximation fit of DINA, C-RUM, and LCDM to fraction subtraction data. The value of M2 suggested that the saturated LCDM fit the data exactly, however, DINA and C-RUM did not have exact fit; and the values of RMSEA2 suggested that DINA and C-RUM did not have close fit to the data.
There are some limitations in the present study. For example, we only investigated the power of M2 under the Q-matrix misspecification condition. Further studies are needed to investigate the performance of M2 under different conditions. The source of misfit at the item level should be investigated when the model does not fit the data well (Liu & Maydeu-Olivares, 2014). Additionally, although the performance of RMSEA2 in LCDM have been investigated in this study, it is desirable to study the performance of RMSEA2 in different CDMs.
Footnotes
Appendix
The Q-Matrix for J = 15, 30, and 50, K = 5
| Item | α1 | α2 | α3 | α4 | α5 | Item | α1 | α2 | α3 | α4 | α5 |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | 1 | 0 | 0 | 0 | 0 | 26 | 1 | 1 | 0 | 1 | 0 |
| 2 | 0 | 1 | 0 | 0 | 0 | 27 | 1 | 0 | 1 | 1 | 0 |
| 3 | 0 | 0 | 1 | 0 | 0 | 28 | 1 | 0 | 1 | 0 | 1 |
| 4 | 0 | 0 | 0 | 1 | 0 | 29 | 0 | 1 | 1 | 0 | 1 |
| 5 | 0 | 0 | 0 | 0 | 1 | 30 | 0 | 1 | 0 | 1 | 1 |
| 6 | 1 | 1 | 0 | 0 | 0 | 31 | 1 | 1 | 0 | 0 | 0 |
| 7 | 1 | 0 | 0 | 0 | 1 | 32 | 1 | 0 | 0 | 0 | 1 |
| 8 | 0 | 1 | 1 | 0 | 0 | 33 | 0 | 1 | 1 | 0 | 0 |
| 9 | 0 | 0 | 1 | 1 | 0 | 34 | 0 | 0 | 1 | 1 | 0 |
| 10 | 0 | 0 | 0 | 1 | 1 | 35 | 0 | 0 | 0 | 1 | 1 |
| 11 | 1 | 1 | 1 | 0 | 0 | 36 | 1 | 1 | 1 | 0 | 0 |
| 12 | 1 | 1 | 0 | 0 | 1 | 37 | 1 | 1 | 0 | 0 | 1 |
| 13 | 1 | 0 | 0 | 1 | 1 | 38 | 1 | 0 | 0 | 1 | 1 |
| 14 | 0 | 1 | 1 | 1 | 0 | 39 | 0 | 1 | 1 | 1 | 0 |
| 15 | 0 | 0 | 1 | 1 | 1 | 40 | 0 | 0 | 1 | 1 | 1 |
| 16 | 1 | 0 | 0 | 0 | 0 | 41 | 1 | 0 | 1 | 0 | 0 |
| 17 | 0 | 1 | 0 | 0 | 0 | 42 | 1 | 0 | 0 | 1 | 0 |
| 18 | 0 | 0 | 1 | 0 | 0 | 43 | 0 | 1 | 0 | 1 | 0 |
| 19 | 0 | 0 | 0 | 1 | 0 | 44 | 0 | 1 | 0 | 0 | 1 |
| 20 | 0 | 0 | 0 | 0 | 1 | 45 | 0 | 0 | 1 | 0 | 1 |
| 21 | 1 | 0 | 1 | 0 | 0 | 46 | 1 | 0 | 1 | 1 | 0 |
| 22 | 1 | 0 | 0 | 1 | 0 | 47 | 1 | 0 | 1 | 0 | 1 |
| 23 | 0 | 1 | 0 | 1 | 0 | 48 | 1 | 0 | 0 | 1 | 1 |
| 24 | 0 | 1 | 0 | 0 | 1 | 49 | 0 | 1 | 1 | 0 | 1 |
| 25 | 0 | 0 | 1 | 0 | 1 | 50 | 0 | 1 | 0 | 1 | 1 |
Note. The 15-item test used Items 1–15 and the 30-item test used Items 1–30.
Acknowledgments
The authors would like to thank the editor, Dr. Dan McCaffrey, and several anonymous reviewers for their valuable comments and suggestions, which led to many improvements.
Authors’ Note
The views expressed here belong to the authors and do not reflect the views or policies of the funding agencies.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by the National Natural Science Foundation of China (Grant No. 31371047) and by the Fundamental Research Funds for the Central Universities (Grant No. SKZZX2013028).
