Abstract
This study investigates polytomous item explanatory item response theory models under the multivariate generalized linear mixed modeling framework, using the linear logistic test model approach. Building on the original ideas of the many-facet Rasch model and the linear partial credit model, a polytomous Rasch model is extended to the item location explanatory many-facet Rasch model and the step difficulty explanatory linear partial credit model. To demonstrate the practical differences between the two polytomous item explanatory approaches, two empirical studies examine how item properties explain and predict the overall item difficulties or the step difficulties each in the Carbon Cycle assessment data and in the Verbal Aggression data. The results suggest that the two polytomous item explanatory models are methodologically and practically different in terms of (a) the target difficulty parameters of polytomous items, which are explained by item properties; (b) the types of predictors for the item properties incorporated into the design matrix; and (c) the types of item property effects. The potentials and methodological advantages of item explanatory modeling are discussed as well.
Keywords
In item response theory (IRT), both person traits and item characteristics are modeled together as predictors in a statistical model with a nonlinear relationship between the item response probability and the predictors. Using explanatory item response model (EIRM; De Boeck & Wilson, 2004), one can explain person side or item side, or both sides of the item response data by incorporating person properties, item properties, and their interactions as predictors in the statistical model. Among them, a typical and well-known case of item explanatory IRT model to examine the item property effects is the linear logistic test model (LLTM; Fischer, 1973), in which the item difficulty parameters of a Rasch model are decomposed into weighted sums of elementary components related to item properties (Embretson & Reise, 2000). The LLTM accounts for the effects of item properties on the item difficulties by incorporating observable item properties such as item design variables, cognitive operations, item content features, and response formats. It is also used as a testing tool for hypothesized constructs in item generation. For example, when a specific test construction rationale that is hypothesized for item generation is developed to predict item difficulties, the LLTM can assess the appropriateness of that rationale and help enable automatized item generation (Reif, 2012). Moreover, the LLTM approach can be used for psychometric studies on measuring the effect of various testing conditions such as item presentation position, content-specific learning, speeded item presentation, and item response format as well as the measurement of changes (Kubinger, 2009; Poinstingl, 2009). Thus, this item explanatory approach can serve to provide informative feedback for enhancing test development, item generation, and cognitive assessment.
Despite such diverse uses of the LLTM approach in educational and psychological measurement practices, it has been less widely used. Historically, most of the early studies using the LLTM often appeared in German-speaking countries and were written in German (Kubinger, 2009). In particular, most of its applications are actually limited to dichotomous data (e.g., Fischer, 1973; Kubinger, 2009; Poinstingl, 2009). However, it is very common to have polytomous data in a wide range of educational, psychological, and sociological applications (De Boeck & Wilson, 2004). In this research, we will focus on ordered-category (or ordinal) responses for polytomous data because they are quite common achievement outcomes in measurement and assessment contexts. For instance, very often the scoring rubrics under a learning progression framework will have ordered categorical levels showing qualitatively different learning progression levels (e.g., Jin, Shin, Johnson, Kim, & Anderson, 2015). Also, we will focus on the adjacent-categories logit relationship between the ordered-category item responses and predictors of person ability and item characteristics. Adjacent-categories logit-based item response models can make person and item parameters separable, that is, parameter separability, and hence permit specifically objective comparisons of persons and items, that is, specific objectivity (Masters, 1982). They are appropriate where local comparison of response probabilities between ordered adjacent categories is of interest (Masters & Wright, 1997). Moreover, adjacent-categories logit ordinal regression models have a suitable parameterization for subjectively assessed scores in ordered responses (Johnson, 2007).
Our research aim is to investigate how the LLTM approach can be applied to polytomous data, particularly in terms of item explanatory extensions of a polytomous Rasch model (hereafter referred to as polytomous item explanatory IRT models) under a general statistical modeling framework, the multivariate generalized linear mixed modeling (MGLMM; Hartzel, Agresti, & Caffo, 2001). One may say that any item explanatory IRT model for dichotomous data could be simply applied to polytomous data. However, this is not the case conceptually and practically. Applying the LLTM approach to polytomous data is not that easy due to the complications of item parameterization in the statistical model as well as the difficulty in reparameterization incorporating the observed item properties. The seminal study by Fischer (1977) and its follow-up studies by Fischer and Parzer (1991) and Fischer and Ponocny (1994) investigated extensions of polytomous Rasch models using the LLTM approach, which are based on conditional maximum likelihood estimation rather than more commonly employed marginal maximum likelihood (MML) estimation. They utilized a normalization constant, basic parameters, and given weights or values of item properties for item parameterization; however, these parameters are complicated and difficult to interpret and even make it hard to calculate and interpret the step difficulties. Glas and Verhelst (1989) investigated extensions of a polytomous Rasch model based on MML estimation. Although they introduced a reparameterization to impose linear restrictions on the item parameters, their item parameters are not straightforward to interpret, and difficult translation into the original parametrization are necessarily required. Masters (1982) pointed out, “To my knowledge, this polytomous generalization of the linear logistic test model (Fischer, 1972) has not been applied.” (p. 153) The complications and difficulty in item parameterization and reparameterization may lead applications of the LLTM approach to polytomous data to be considerably less common than to dichotomous data. Such problems for polytomous item explanatory IRT models can be resolved using the MGLMM framework, which is flexible in statistical modeling based on MML estimation.
To achieve the aim of this research, given our focus on ordered-category responses and the adjacent-categories logits, a unidimensional polytomous Rasch model will be reviewed as a starting point through the following sections. Next, we will investigate how to develop it toward polytomous item explanatory IRT models using the LLTM approach under the MGLMM framework. And then, two empirical studies that focus on applying these models to the Carbon Cycle assessment data and to the Verbal Aggression data, respectively, will be demonstrated. In the empirical studies, we can see how the developed models work in practice and how the observed item properties explain and predict the difficulties of polytomous items.
Item Explanatory Extensions of a Polytomous Rasch Model
Partial Credit Model
Based on ordered-category responses and adjacent-categories logits, a polytomous Rasch model will be extended toward polytomous item explanatory IRT models. To reach those extensions, we will start with a basic model for the adjacent-categories logit-based item response models, the partial credit model (PCM; Masters, 1982). The PCM is a straightforward application of the Rasch model to polytomous data (Masters & Wright, 1997), in which the conditional probability that person p with ability
where
For explanatory convenience and intuitive interpretation of item parameters, the PCM can be described in terms of local comparison in the response probabilities for pairs of adjacent categories in a sequence in ordered-category responses (Masters & Wright, 1997). In the local comparison, a linear predictor element of the mth adjacent-categories logit (Tuerlinckx & Wang, 2004), from the adjacent category score
where
Additionally, using a twofold item parameterization adopted in the ConQuest software (Wu, Adams, Wilson, & Haldane, 2007), the step difficulty parameter
where
Item Explanatory Extensions With Item Properties
In the adjacent-categories logit-based item response models under the MGLMM framework, there are three types of predictors that can be included in a design matrix for the item side of the response data (Rijmen, Tuerlinckx, De Boeck, & Kuppens, 2003; Tuerlinckx & Wang, 2004): (a) item predictors, (b) step predictors, 2 and (c) item-by-step predictors. The item predictor is a predictor if and only if the elements of the corresponding column of the design matrix vary across items but are fixed as category scores for steps within each item for all persons, the step predictor is a predictor if and only if the elements of the corresponding column of the design matrix vary across steps but are constant across items for all persons, and the item-by-step predictor is a predictor if and only if the elements of the corresponding column of the design matrix vary across all items and steps but are constant across persons. Simply put, predictors in the design matrix could be indicators or explanatory properties. For the PCM with the original item parameterization, item-by-step predictors accounting for the step difficulties are indicators for each step of each item. For the PCM with the twofold item parameterization, item predictors accounting for the item locations are indicators for each item and item-by-step predictors accounting for the step deviations are indicators for each step of each item. Thus, the PCM is a traditional measurement approach of describing individual differences in items’ difficulties and/or persons’ abilities.
For item explanatory extensions of the PCM using the LLTM approach, however, predictors in the design matrix are item properties that have explanatory value for polytomous item characteristics beyond descriptive value. If the observed item properties are hypothesized to determine the item main effects and/or have explanatory value for unique characteristics inherent within items regardless of steps (e.g., item content features, reading passage length), they are incorporated as item predictors into the design matrix. If the observed item properties are hypothesized to account for the step main effects and/or have explanatory value for unique characteristics inherent within steps for all items (e.g., learning progression levels, scoring rubrics), they are incorporated as step predictors into the design matrix. If the observed item properties are assumed to account for the item-by-step interaction effects and/or have explanatory value for unique characteristics inherent within steps of items (e.g., cognitive operations, item exposure time), they are incorporated as item-by-step predictors into the design matrix.
However, the item properties that would be incorporated as step predictors are usually predetermined by test/item developers before items design. For example, the progression levels hypothesized in a learning progression framework are hard to be explained by external explanatory properties, because they are just indicators for the step main effects. Under the MGLMM framework, particularly, the number of adjacent-categories logits is determined by the number of response categories. The step main effects would not likely be explained by item properties, unless the number of response categories is reduced. Because of these concerns, in this research, we will leave out the case that polytomous item characteristics are accounted for by item properties incorporated as step predictors.
Polytomous Item Explanatory Item Response Theory Models
As discussed in the introduction, item parameterization and reparameterization incorporating the observed item properties for polytomous item explanatory IRT models are complicated and difficult. To avoid such complications and difficulties, it is helpful to clarify the target difficulty parameters of polytomous items, which are explained by item properties as well as the types of predictors as which the observed item properties are incorporated into the design matrix. Building on the PCM, there are two cases of polytomous item explanatory models: (a) the item location (overall item difficulty) parameters
By default, we will consider a two-facet 3 measurement situation (i.e., person and item) and the same set of item properties for each model. Item properties are regarded as subfacets within an item facet (Wang & Wilson, 2005), since they are unique characteristics inherent in a set of items.
Item Location Explanatory Many-Facet Rasch Model
First, consider the former case based on the PCM. Note that the item location parameter is interpreted as the overall item difficulty for each polytomous item. As was done in the LLTM, one may want to account for the overall item difficulties by incorporating item properties as item predictors into the design matrix. The twofold item parameterization for the PCM enables us to impose linear restrictions focused on the item location parameters and also to estimate the step deviation parameters for each item. Thus, the restricted item location parameters
so that
where
This polytomous item explanatory model is a variation of the many-facet Rasch model (MFRM; Linacre, 1989). By using the twofold item parameterization, the item location parameters are decomposed into a linear combination of the effects of subfacets (i.e., item properties) within the item facet. The step deviation parameters are additively estimated for each item, because each item can have a different number of steps across ordered response categories in its scale structure. Building on the original idea of the MFRM approach, this item explanatory extension of the PCM will be called the “item location explanatory MFRM.” Although this model is a variation of the polytomous MFRM, compared with factorial (ANOVA-style) notations, which are commonly used in the MFRM, the model specification in Equation 5 has a more general regression notation. While the original MFRM approach uses only categorical predictors to represent the effect of each facet (i.e., facet descriptive), this model can incorporate both categorical and continuous explanatory predictors inherent in the item facet to see the effect of item properties (i.e., item explanatory). For instance, this model is useful and suitable when one may want to see the effects of both passage lengths (continuous) and contents (categorical) on the overall item difficulties in a reading comprehension test. Moreover, the step deviation parameters can be estimated variously according to the scale structure in the MFRM (see Eckes, 2009), but they are estimated for each item in this model.
To show a simple illustration of this item location explanatory MFRM under the MGLMM, suppose we have two polytomous items with three response categories for each (
First, consider the linear predictor elements of the adjacent-categories logits for Item 1:
when
when
when
Second, consider the linear predictor elements of the adjacent-categories logits for Item 2:
when
when
when
Consequently, the linear predictor vector for person p of this MFRM variation corresponding to Equation 5 can be written as
Step Difficulty Explanatory Linear Partial Credit Model
The second case is an item explanatory extension of the PCM, in which the step difficulties are determined by item-by-step property effects. Compared with the previous approach, the effects of item properties are step-specific, that is, their estimates can vary across steps. For model specification, we can impose linear restrictions on the step difficulty parameters by incorporating item properties as item-by-step predictors into the design matrix. Thus, the restricted step difficulty parameters
so that
where
This polytomous item explanatory model is a variation of Fischer and Ponocny’s (1994) linear partial credit model (LPCM). Although less interpretable item parameters such as a normalization constant and basic parameters were employed, linear restrictions were imposed on the step difficulty parameters and the basic parameters were estimated from the values of item-by-step predictors. Building on the original idea of the LPCM, this item explanatory extension of the PCM will be called the “step difficulty explanatory LPCM.” Compared with the original parameterization for the LPCM, however, the model specification in Equation 8 employs the item-by-step property effect parameters
For a simple illustration of this step difficulty explanatory LPCM under the MGLMM, we will use the previous example for the MFRM variation, but now will consider decomposition of step difficulty parameters rather than decomposition of item location parameters. After using dummy coding for the categorical item property (Y is a reference), an intercept (
First, consider the linear predictor elements of the adjacent-categories logits for Item 1:
when
when
when
Second, consider the linear predictor elements of the adjacent-categories logits for Item 2:
when
when
when
Accordingly, the linear predictor vector for person p of this LPCM variation corresponding to Equation 8 can be expressed as
Data Analysis
Two empirical studies were conducted to show how the proposed polytomous item explanatory models under the MGLMM framework, that is, the item location explanatory MFRM and the step difficulty explanatory LPCM (hereafter referred to as MFRM and LPCM, respectively), work for practical item response data sets. These two item explanatory approaches are conceptually and functionally different in terms of the target difficulty parameters of polytomous items to be explained and the types of predictors for the incorporated item properties as well as the types of item property effects. To see how the two approaches are different in practice, we fit both models to the Carbon Cycle assessment data and to the Verbal Aggression data, respectively. In addition, since the PCM is a saturated model for polytomous Rasch models, it was fitted to each empirical data set to see how the two polytomous item explanatory models perform compared with the saturated model. To fit all the adjacent-categories logit-based item response models including the proposed models and the PCM under the MGLMM framework, we used the gllamm command in Stata, which is implemented with MML estimation (Rabe-Hesketh, Skrondal, & Pickles, 2004).
Regarding the goodness of fit of the models, the likelihood ratio (LR) test was conducted to compare the nested models. The two polytomous item explanatory models are nested within the PCM, and the LPCM is nested within the MFRM due to the restrictions on the item parameters. Three other goodness-of-fit indices were also reported for each model: the deviance (
Study 1: Application to the Carbon Cycle Assessment Data
Data and Item Properties
A subset of the Carbon Cycle assessment data, specifically pretest responses of the Math and Science Partnership Carbon student assessment data collected during the 2010-2011 academic year, was used in the first empirical study. Participants were middle and high school students from urban, suburban, and rural areas in five states. The Carbon Cycle assessment and items were developed based on a learning progression framework for carbon cycling in socioecological systems in science education (Jin et al., 2015; Mohan, Chen, & Anderson, 2009), by conducting an iterative process of designing, analyzing, and modifying assessments with interviews for secondary school students. The learning progression for carbon cycling in socioecological systems has four ordered levels of achievement corresponding to students’ progress toward more sophisticated reasoning about biogeochemical processes (Mohan et al., 2009), as shown in Table 1. Based on the learning progression framework, items were developed to ask students to answer forced choice questions and explain their choices, and their scoring rubrics were developed to score item responses into the four ordered achievement levels. If students’ answers are not related to the question, or are illegible or nonsensical, they are treated as missing.
The Learning Progression Framework for the Carbon Cycle Assessment.
The 13 Carbon Cycle items consist of a combination of three categorical item properties: (a) the Process property has three types of biogeochemical processes in carbon cycling that transform carbon in socioecological systems at multiple scales—cellular respiration (CR), photosynthesis (PS), and digestion/biosynthesis (DB); (b) the Progress property has four learning progress variables related to the carbon cycling process—large-scale systems (LS), micro-scale systems (MS), energy (EN), and mass (MA); and (c) the Format property has two item response formats—multiple choice with explanation (MC) and yes/no choice with explanation (YN). For example, the first item (BODYTEMP) can be identified as a combination of a CR predictor in the Process property, an EN predictor in the Progress property, and an MC predictor in the Format property. The text of the first item is presented in Figure 1.

The text of Carbon Cycle Item 1 (BODYTEMP).
Through item analyses for the data and consideration of item design, the items can be classified by the predictors of each item property as in Table 2. These predictors of the item properties are functioning as weights of the elementary components, which are gathered into a Q matrix in the LLTM approach (Kubinger, 2009; Poinstingl, 2009). Three categorical item properties were dummy coded to be incorporated in polytomous item explanatory models; the DB in the Process property, the MA in the Progress property, and the YN in the Format property served a reference for each item property.
The Carbon Cycle Items Composed of Three Item Properties.
In the initial data, several items had too few or zero responses for the highest level, which would result in poor or unfeasible estimation of the third-step difficulty parameters. To avoid such sparseness problem, item responses of the four-level categories (1, 2, 3, 4) were recoded into three category scores (0, 1, 2, 2). We can interpret the second-step difficulties as the relative difficulties as one goes from the Level 2 to the combined Levels of 3 and 4. Additionally, cases with less than three valid item responses were dropped from the analysis to enhance the quality of estimation. In total, 1,157 students’ responses on the 13 polytomous items were included in the analysis.
Empirical Results
According to the data analysis procedure, the PCM, the MFRM, and the LPCM were fit to the Carbon Cycle assessment data. Table 3 shows the results of the fitted models to see performance of the two polytomous item explanatory models that we proposed compared with the saturated model, the PCM.
Performance of the Many-Facet Rasch Model and the Linear Partial Credit Model on the Carbon Cycle Assessment Data.
Note. AIC = Akaike information criterion; BIC = Bayesian information criterion; q = The number of estimated parameters.
The estimated person variance (
In terms of the goodness of fit, both the MFRM and the LPCM appeared to fit worse than the PCM. The LR test comparing to the PCM was significant for both the MFRM, χ2(6) = 767.12, p < .001; and the LPCM, χ2(12) = 971.97, p < .001. This result was confirmed in that all goodness-of-fit indices of the two models were greater than the PCM. This is as expected because the item explanatory models commonly fit worse than the saturated model (Kubinger, 2009; Tuerlinckx & Wang, 2004). The methodological advantage of a smaller number of item parameters in the LLTM approach is at the cost of statistically lower goodness of fit.
When comparing between the two item explanatory models, both the AIC and BIC indicate that the MFRM showed a superior goodness of fit than the LPCM for the Carbon Cycle data. The LR test comparing the LPCM to the MFRM was also significant as χ2(6) = 204.85 (p < .001). This result makes sense because the step deviation parameters were freely estimated for each item in the MFRM and hence it had a larger number of the estimated parameters than the LPCM. Nonetheless, this statistical result does not mean that the LPCM is methodologically or practically inferior to the MFRM. They are conceptually and functionally different models in that the target difficulty parameters of polytomous items which are explained, the types of predictors for the incorporated item properties, and the types of item property effects are different.
In addition to the goodness-of-fit comparison, a graphical comparison is useful to see performance of the two polytomous item explanatory models, by examining agreement for the estimated and calculated step difficulties between the fitted models. The step difficulties

Graphical comparison of step difficulties between the models fitted to the Carbon Cycle assessment data.
The correlations show the same results to the graphical comparison. Correlations with the PCM were similar in the two item explanatory models (
To see the practical difference between the two polytomous item explanatory models, the effects of the three item properties on the overall item difficulties or the step difficulties were reported for each model. Table 4 shows the results of the item property effects in the item location explanatory MFRM. Recall that there are three categorical item properties (Process, Progress, and Format) in the design of the Carbon Cycle assessment. The three item properties incorporated as item predictors were taken into account to explain and predict the overall item difficulties (item locations). As we examined the item property effect parameters
Item Property Effects on the Carbon Cycle Items in the Many-Facet Rasch Model.
For the Process property, holding other properties constant, cellular respiration (
In addition, 13 step deviation parameters were estimated for each item. Accordingly, the constructed step difficulties for each observed combination item were calculated by using the estimated item property effects and step deviation parameters. For example, the second-step difficulty of Item 1 (BODYTEMP),
Table 5 shows the results of the item property effects in the step difficulty explanatory LPCM. The three item properties incorporated as item-by-step predictors were taken into account so that they could explain and predict the step difficulties for each step. Based on the estimated item-by-step property effect parameters
Item Property Effects on the Carbon Cycle Items in the Linear Partial Credit Model.
With regard to the effects of the three item properties on the step difficulties of the first step, keeping other properties constant, the cellular respiration process (
For the step difficulties of the second step, holding other properties constant, the cellular respiration process (
Furthermore, we can reconstruct the step difficulties for each item with the observed item property combination using weighted sums of the estimated item property effects. For example, the second-step difficulty of Item 1 (BODYTEMP),
Study 2: Application to the Verbal Aggression Data
Data and Item Properties
For the second empirical study, the Verbal Aggression data set (Vansteelandt, 2000) was used. Participants were first-year psychology students at a Dutch-speaking Belgian university. Twenty-four items were presented in Dutch, the native language of all the participants, to ask behavioral questions about verbally aggressive reactions to frustrating situations. A total of 316 persons responded to the 24 items. The item responses were three ordered-category responses (no = 0, perhaps = 1, and yes = 2) in the order of endorsing an item, which were used as polytomous data without dichotomization in this study.
All items were designed to have a stem, which describes a frustrating situation, and a verbal aggression response part, which describes how people could respond to the situation in question. In this item design, there were three experimental design factors:
(a) the Behavior Mode factor has two levels of modes—wanting (Want) and doing (Do),
(b) the Situation Type factor has two types of situations—situations in which someone else is to blame (Other-to-blame) and situations in which oneself is to blame (Self-to-blame), and
(c) the Behavior Type factor has three kinds of verbal aggressive behaviors, which represent the extent to which they ascribe blame and the extent to which they express frustration—cursing (Curse), scolding (Scold), and shouting (Shout).
For the verbal aggression response, one of the two behavioral modes could be combined with one of the three verbal aggressive behaviors. For example, combinations of these two factors generate six responses such as “I would curse” (for doing and cursing) and “I would want to scold” (for wanting and scolding).
For the item stem, four frustrating situations were used:
Other-to-blame A: A bus fails to stop for me. (Bus)
Other-to-blame B: I miss a train because the clerk gave me faulty information. (Train)
Self-to-blame A: The grocery store closes just as I am about to enter. (Store)
Self-to-blame B: The operator disconnects me when I used up my last 10 cents for a call. (Call)
These four frustrating situations were nested in the second design factor: two other-to-blame situations and two self-to-blame situations. The full item contains one of the four frustrating situations for the item stem, and one of the six combinations of two behavioral modes and three behavior types for the verbal aggression response to the situation. After considering two specific situations of the same type (A and B) as replications, in total 24 items were written to fit a 2 × 2 × 3 design with two replications within each cell. For example, the first item, “A bus fails to stop for me. I would want to curse” was designed and identified as a combination of Want (Behavior Mode), Other-to-blame (Situation Type), and Curse (Behavior Type).
The three design factors are regarded as categorical item properties. The Verbal Aggression items are classified by the predictors of each item property as in Table 6, which can be functioning as weights in a Q matrix in the LLTM approach. To incorporate the three categorical item properties into polytomous item explanatory models, they were dummy coded: the Want in the Behavior Mode property, the Self-to-blame in the Situation Type property, and the Shout in the Behavior Type property served a reference for each item property.
The Verbal Aggression Items Composed of Three Item Properties.
Empirical Results
The PCM and the two polytomous item explanatory models were fit to the Verbal Aggression data, according to the data analysis procedure. Table 7 shows the results of the fitted models: the MFRM, the LPCM, and the PCM.
Performance of the Many-Facet Rasch Model and the Linear Partial Credit Model on the Verbal Aggression Data.
Note. AIC = Akaike information criterion; BIC = Bayesian information criterion; q = The number of estimated parameters.
The estimated person variance (
When comparing the goodness of fit, the LR test (compared with the PCM) was significant for the MFRM, χ2(19) = 165.142, p < .001, and the LPCM, χ2(38) = 206.818, p < .001, meaning that both the MFRM and the LPCM fit worse than the PCM, as expected. This result was confirmed by the AIC: The values for the two models were greater than the PCM (12737.47 for the PCM, 12864.61 for the MFRM, 12868.29 for the LPCM). This makes sense because the smaller number of estimated parameters in the item explanatory models made model fit (accuracy) much worse in spite of them being more parsimonious models.
However, this was not supported by the BIC: The value for the PCM was greater than the others (13131.06 for the PCM, 13105.58 for the MFRM, 12956.64 for the LPCM). Although this conflicted result might seem surprising, it can be understood by considering the difference between the AIC and BIC. They differ in theoretical motivations, objectives, assumptions, and meaning of penalty terms (Kuha, 2004; Wagenmakers & Farrell, 2004). Basically, the BIC penalizes complexity more heavily for the number of freely estimated parameters, whereas the AIC penalizes a model with more parameters much less than the BIC (Kuha, 2004). In other words, the BIC favors parsimonious models to a greater extent than the AIC does. In Table 7, the 48 parameters for item effects in the PCM were reduced to 29 in the MFRM and 10 in the LPCM. In terms of the BIC, we could say that the two polytomous item explanatory models performed better than the PCM in this empirical study. In addition, in the Bayesian aspect of the BIC, it assumes that the true data-generating model is in the set of candidate models, and it measures the degree of belief that a certain model is the true model, which generates the observed data (Wagenmakers & Farrell, 2004). In this view, we could also say that the LPCM that had the lowest BIC was most likely to be the true data-generating model for the Verbal Aggression data. Thus, this result implies that the three-factor design with two replications worked well for item generation and the three item properties (design factors) had high explanatory value for the Verbal Aggression items.
Comparing the LPCM with the MFRM, the LR test was significant, χ2(19) = 41.68, p = .002, which indicates that the MFRM fit better than the LPCM to the Verbal Aggression data. The lower AIC value for the MFRM supports this finding, which make sense because the larger number of the estimated parameters in the MFRM made model fit (accuracy) much better than the LPCM. However, the LPCM had the lower BIC value, meaning that the LPCM performed better because it was much more parsimonious than the MFRM. In this empirical study, we couldn’t conclude that one was better than the other in terms of the goodness of fit. In fact, the two polytomous item explanatory models are different methodologically as well as practically.
In addition, a graphical comparison was conducted to examine agreement for the estimated and calculated step difficulties between the fitted models. The constructed step difficulties

Graphical comparison of step difficulties between the models fitted to the Verbal Aggression data.
The correlations show the LPCM had a slightly higher correlation with the PCM than the MFRM (
The effects of the three item properties on the overall item difficulties or the step difficulties were reported separately for each model to show the practical difference between the two polytomous item explanatory models. For the MFRM, three categorical item properties (Behavior Mode, Situation Type, and Behavior Type) from the item design factors were incorporated as item predictors to explain and predict the overall item difficulties (item locations) of endorsing an item at higher levels. Table 8 represents the results of the item property effects in the item location explanatory MFRM, and the item property effect parameters
Item Property Effects on the Verbal Aggression Items in the Many-Facet Rasch Model.
For the Behavior Mode, holding other properties constant, the Do mode made the overall item difficulty of an item 0.43 logits more difficult to endorse than the Want mode (
In addition, 24 step deviation parameters were estimated for each item. By using the estimated item property effects and step deviation parameters, the constructed step difficulties for individual items could be calculated. For example, the first-step difficulty of the first item (“A bus fails to stop for me. I would want to curse.”),
Table 9 shows the results of the item property effects in the step difficulty explanatory LPCM. The three item properties were incorporated as item-by-step predictors to explain and predict the step difficulties for each step. The estimated item-by-step property effect parameters
Item Property Effects on the Verbal Aggression Items in the Linear Partial Credit Model.
The effects of the three item properties on the step difficulties of the first step were interpreted as follows. When a person answered “perhaps” rather than “no,” keeping other properties constant, the Do mode made the items 0.54 logits more difficult to endorse than the Want mode (
To reconstruct the step difficulties for individual items, we can calculate them using weighted sums of the estimated item property effects. For example, the first-step difficulty of the first item (“A bus fails to stop for me. I would want to curse.”),
Conclusion and Discussion
We have investigated how to apply the LLTM approach to polytomous data under the MGLMM framework, given the ordered-category responses and the adjacent-categories logits. The two item explanatory extensions of the PCM—the item location explanatory MFRM and the step difficulty explanatory LPCM—were developed and specified with general statistical formulations. These two polytomous item explanatory IRT models were fit to the Carbon Cycle assessment data and to the Verbal Aggression data, respectively, and their explanatory values were demonstrated in that both models help us figure out how the observed item properties affect the overall item difficulties or the step difficulties in the empirical studies.
The first empirical study showed that the MFRM had a superior goodness of fit compared with the LPCM in terms of both the AIC and BIC, whereas in the second empirical study there was no uniform agreement between the two goodness-of-fit indices regarding the performance of the MFRM and the LPCM. In both studies, the two polytomous item explanatory models showed comparable performance in reconstructing the step difficulties using the estimated item property effects, but they demonstrated practical differences in interpreting the item property effects on the polytomous item difficulties. The MFRM could explain and predict the overall item difficulties (item locations) by the item properties incorporated as item predictors, and the LPCM could explain and predict the step difficulties by the item properties incorporated as item-by-step predictors. Thus, the explanatory and predictive values of polytomous item explanatory IRT models can go beyond traditional descriptive measurement, as they provide informative feedback for enhancing quality of item development in practice.
In fact, the two polytomous item explanatory models are methodologically and practically different in terms of the target difficulty parameters of polytomous items, which are explained by item properties, the types of predictors for the item properties incorporated into the design matrix, and the types of item property effects. Before fitting the model to the data, it is highly recommended for model selection to clarify which polytomous item difficulties should be explained by item properties as well as to examine what kind of predictors should be used for the item properties with consideration for the types of item property effects. When designing items, it is helpful to examine underlying hypotheses for the item properties or design factors. If an item property is assumed to affect switching levels of the construct as in the learning progression framework (see Table 1), the LPCM is useful to test that hypothesis. If one assumes that an item property relates to the item as a whole but not to specific levels, the MFRM is appropriate to see its influence on the overall construct. One recommendable and systematic way of designing items is to use Wilson’s (2005) four-building-blocks approach to constructing measures, which takes advantage of the principles of sound educational and psychological measurement. By using an iterative cycle of the measurement processes (construct mapping, item design, outcome space, and measurement model) one can develop items with articulating the underlying hypotheses.
Among the EIRM approaches, we have focused on explaining the item side only of the response data. There are huge potentials to extend polytomous item explanatory IRT models. For instance, although we have investigated the main effects of item properties in the empirical studies, the interaction effects between them can be taken into account as well. We can also consider adding person predictors and/or person-by-item predictors, which are extended with person or doubly explanatory models such as a latent regression and differential item functioning analysis. A future study could address a multidimensional extension of the models to examine interactions between item properties and individual persons when the item property effects vary over all persons, which are applications of the random-weights linear logistic test model (Rijmen & De Boeck, 2002) to polytomous data. In addition, when we are not sure of perfect explanation through the observed item properties, by allowing for random residuals, item error terms can be added to the polytomous item explanatory models, which can enhance predictions of the item property effects. This is an extension of the linear logistic test model with random item error (Janssen, Schepers, & Peres, 2004) to polytomous data (see Kim, 2018). Last, we have considered the two-facet situation (person and item), but this may not be the only consideration. A future study could concern other-facet situations (e.g., rater, task, and criteria) with different facet properties (see Eckes, 2009).
This research also sheds light on the methodological advantages of item explanatory modeling. In the polytomous item explanatory models we have specified, the step difficulty parameters or the item location (overall item difficulty) parameters are constrained to be a linear combination of the effects of item properties. These models allow to estimate a smaller number of item parameters than the saturated model to explain and predict the item effects, which is a methodological advantage in extracting essential and meaningful elementary components by incorporating item properties. They are useful to understand how the items generated by item design or test construction work, as well as to validate hypothesized constructs for item design or test construction. The models can also help us generate new items by composing specified combinations of the item property effects so that we can predict item locations or step difficulties of the newly developed items in a more scientific and systematic manner rather than an intuitive manner. It is promising that item explanatory models are widely used for polytomous data to measure the effects of various testing conditions such as raters, cognitive operations, and item exposure time, as has been done for dichotomous data (e.g., Fischer, 1973; Kubinger, 2009; Poinstingl, 2009).
In conclusion, we can say that the polytomous item explanatory IRT models can have a large effect in methodological foundations for the educational and psychological measurement, particularly with consideration for De Boeck and Wilson’s (2004) EIRM framework. Peter Drucker, a famous management thinker, said, “If we can’t measure it, we can’t manage and improve it.” Likewise, if we cannot explain the measurement, we cannot know how to improve it.
Footnotes
Acknowledgements
The authors would like to thank Sophia Rabe-Hesketh for her careful comments on model specification and thank anonymous reviewers for their helpful comments on an earlier draft. The authors would also like to thank Charles W. Anderson and Hui Jin for leading the Carbon Cycle research project and sharing a data set.
Authors’ Note
Jinho Kim is also affiliated with KU Leuven and ITEC, imec research group at KU Leuven, Kortrijk, Belgium.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: Collection of the Carbon Cycle assessment data was supported in part by grants from the National Science Foundation (NSF): Learning Progression on Carbon-Transforming Processes in Socio-Ecological Systems (NSF 0815993), and Targeted Partnership: Culturally relevant ecology, learning progressions and environmental literacy (NSF 0832173). Any opinions, findings, and conclusions or recommendations expressed in this material are those of the authors and do not necessarily reflect the views of the NSF.
