Abstract
Responses to items with formats in more than two ordered categories are ubiquitous in education and the social sciences. Because the putative ordering of the categories reflects an understanding of what it means to have more of the variable, it seems mandatory that the ordering of the categories is an empirical property of the assessments and not merely a property of the model used to analyze them. To provide an unequivocal interpretation of category ordering in rating formats, this article expands the original derivation of the polytomous Rasch model for ordered categories. To do so, it integrates a complex of mathematical relationships among response spaces from which a space of experimentally independent Bernoulli variables, characterized by Rasch’s simple logistic model, can be inferred. From this inference, the article establishes the necessary and sufficient evidence to test the hypothesis that the required ordering of the categories is an empirical property of the assessments. This expanded derivation, which exposes how Adams, Wu, and Wilson (2012) misconstrue the model and its implications, is intended to dispel the so-called disordered threshold controversy they claim exists.
Keywords
Introduction
The immediate motivation for this article is a response to Adams, Wu, and Wilson (AWW; 2012) who consider that the threshold order in the polytomous Rasch model (PRM), the special cases of which are the rating and partial credit parameterizations, is controversial. Because the focus is on the response of a single person to a single item, and the threshold structure within an item is identical in the two cases, this article uses the general term PRM rather than rating scale or partial credit model.
Many of the topics AWW canvas seem irrelevant to their title. However, with respect to the topic in their title, AWW assert that (a) the Guttman structure with a hypothesized threshold order is not a formal part of the PRM and that category ordering is independent of the empirical threshold parameter estimates, (b) the reason for reversed thresholds is that a particular category has a low frequency, and (c) claim that item fit to the model adds to their case that reversed threshold estimates imply no anomaly in the empirical ordering of the categories.
AWW’s first two assertions are rebutted in this article, and the test of fit on which they base their third assertion is shown to be irrelevant. As a result, the conclusion of this article, exactly the opposite of AWW’s, is that every item that has disordered threshold estimates shows an anomaly. Moreover, the implication is that a substantive, not a statistical, explanation needs to be sought to understand and correct the anomaly. The article uses an example provided by AWW to illustrate this conclusion and implication.
However, the purpose of the article is broader than simply dispelling AWW’s misunderstandings. The broader purpose is to provide a full appreciation of the set of coherent mathematical properties of the model that enable the empirical investigation of the ordering of categories. Therefore, the emphasis is on explaining the model from first principles, and within this explanation, to reply to AWW’s assertions. It is noted that the various algebraic forms of the model in AWW are not an issue, and the one chosen in this article is used for purposes of efficient exposition.
Assessments in ordered category formats, especially rating formats, are ubiquitous in educational and other social sciences and are used by analogy to measuring instruments of the natural sciences. Essentially, a response in a higher category is intended to reflect a higher location of the respondent on the variable of assessment, and symmetrically, it is intended that the higher the location on the variable, the more likely a response in a higher category.
Order among categories is a substantively severe constraint on the response structure. However, the order may not work empirically as intended. If it does not, then not only is the understanding of what it means to have more of the variable called into question, but also it may lead to incorrect decisions being made, ranging from areas of individual diagnoses, through policy decisions, to research interpretations. Therefore, following Fisher (1958), it seems mandatory that the ordering is a verifiable, empirical property of the categories, and not merely assumed. Thus, the PRM, which can be used to identify whether or not the ordering of categories is working as intended, is an important tool for analyzing responses in ordered categories and, in particular, rating formats.
Two distinct components are integrated in this article. The first specifies, a priori to data collection, category order as a hypothesis to be assessed; the second is the mathematical structure of the PRM that permits assessing this hypothesis. Following first Kuhn (1961/1977) that the function of measurement is to disclose anomalies, and second Rasch (1961) that his models are models for measurement, the term anomaly is used for evidence of any incompatibility between the data and category order in the PRM.
This article establishes that the empirical ordering of categories is equivalent to the empirical ordering of successive thresholds that define adjacent categories on a continuum. If dichotomous assessments at each of the thresholds were experimentally independent, then evidence of threshold order would be available readily. However, because there is just one response in more than two categories, dichotomous assessments at the thresholds are neither observable nor independent. Instead, if they can be considered at all, they must be latent and therefore inferred. This article explicates a complex of mathematical relationships of the PRM to show that inferences from ordered category formats can be made as if responses at the thresholds were experimentally independent.
As shown in the article, deriving the above result in the PRM requires a subtle distinction between the structure and the values of thresholds in the PRM, a distinction that AWW conflate.
The rest of the article is presented as follows. First I explain in detail, including with a simulated example, the role of the Guttman structure in the threshold and category order in the PRM. Then I demonstrate that AWW’s contention that the determinant of disordered thresholds is a low frequency in a related category in the whole sample is incomplete. Following this I explain the distribution of frequencies that leads to reversed thresholds and demonstrate that the test of fit that AWW use is irrelevant to implications of reversed thresholds. All derivations are consistent with the original identification of thresholds in the Rasch model for ordered categories (Andrich, 1978). Then I provide a discussion. Finally, I provide a conclusion anticipating further consideration of the role of tests of fit in the application of models in a measurement context.
Standard Ordered Category Formats and a Resolved Design
Although AWW refer to accounting for local dependence between items using the PRM, an application considered in Andrich (1983, 1985a), that case is not considered in this article. Instead, because AWW, as did the original explication of the thresholds in the Rasch model (Andrich, 1978), focus on a standard, rating format, this article does the same. However, AWW provide an example of partial credit, and their example is reexamined from the perspective of this article.
Andrich (1978) and AWW use only three ordered categories illustratively, and the latter use those of Fail, Pass, and Distinction in the assessment of proficiency. These same categories are used in this article. To make the interpretations concrete, Table 1 shows a specific example from Harris (1991) in the assessment of essay writing with respect to the criterion of setting, which is elaborated with descriptors. It is clear that the successive descriptors imply order of achievement.
Operational Definition for Judging Essays With Respect to the Property of Setting
Source: Adapted, with permission, from Harris (1991), p. 49. The labels of Fail, Pass, and Distinction have been added for this article.
Resolving a Rating Design Into an Experimentally Independent Design
The inferences from the PRM rest on the relationships among three probability spaces of Bernoulli variables introduced in Andrich (1978) and expanded on here. Unfortunately, AWW pay no attention to response spaces.
To motivate the role of responses spaces, a design with experimentally independent responses based on the format in Table 1 is constructed. This design is considered a thought experiment as in Andrich (1978). As part of their misunderstanding, AWW appear to take such a design literally.
As we have argued above, this is not the case; parameter disorder does not imply category disorder, unless a particular latent process is assumed; a process that can only hold when judges are involved in some judgment processes and even then the process is merely conjecture. (AWW, 2012, p. 24)
Although the design imagines judges making independent decisions, it is proved that the mathematical structure of the PRM is independent of any actual judging process. Because of continuing misunderstandings in AWW, this motivation, which was introduced in Andrich (1978) and elaborated in Andrich (2005, 2010), is expanded further.
Consider a design in which a single response in one of the

The resolved design constructed from the standard design
The first two rows locate the thresholds in the resolved design. Each threshold is defined as the point on the real number line at which an essay will have an equal probability of being declared successful and unsuccessful in meeting the standard. In this example, the two threshold values are notated as
As a direct consequence of the intended category order, the a priori specification is that the threshold at Distinction is more difficult than the threshold at Pass, that is,
It is stressed that in this design, it is the continuum that is dichotomized at two places (thresholds) with independent decisions made at each threshold. This specification is reinforced in Figure 1 by interpolating descriptors by “~” to indicate that the descriptors lead from one to the other on the continuum and by the incomplete ovals that envelop the descriptors on either side of each threshold. The independence of the decisions is reinforced by the resolution of the continuum for each decision. Rows 1 and 2 of Figure 1 are referred to in this article as a resolved design. This dichotomization of the continuum contrasts with the dichotomization of a single probability distribution in the graded responses model (Samejima, 1969).
Row 3 in Figure 1 reintegrates the location of the thresholds on a single continuum. This depiction is exactly that of locations of two independent dichotomous items when responses to them are characterized by Rasch’s simple logistic model (SLM; Rasch, 1960/1980).
Row 4 in Figure 1 repeats the design of an ordered category format of Table 1 in relation to the resolved design. This row is referred to in this article as the standard design. To reflect this design, the threshold values

TCCs for the two thresholds in the resolved design
Indeed, in relation to Andrich (1978), AWW state, It appears from Figure 3 that Andrich sees the tau parameters as thresholds in the sense of Thurstone (1928). Furthermore, he says that “

Integration of the standard design with probabilistic responses at the thresholds in the resolved design
The gross misrepresentation and misunderstanding of Andrich (1978, 2005) is puzzling. First, the division of the continuum in Andrich (1978), as in this article, is central to the identification of thresholds within the general form of the polytomous Rasch model then known in the form of Andersen (1977). This is the only form that retains sufficient statistics characteristic of Rasch models. Second, the 1978 article explicitly contrasts the newly defined thresholds from Thurstone’s in which a single probability distribution is dichotomized. Thurstone’s dichotomization leads to the graded response model (Samejima, 1969) in which no sufficient statistics exist. Thurstone’s conception was shown in Figure 1, mine in Figure 3, which is the only one AWW reproduce. The conceptions were again explicitly contrasted for other implications in Andrich (1995).
Bernoulli Random Variables in the Resolved Design and Order of Success Rates
We now introduce the probabilistic element in the decisions at the thresholds. Let
be a Bernoulli random variable for the response of (hypothetical) Judge k to essay n at threshold
where the response
It is stressed that although an empirical, resolved design could be constructed, it is hypothetical and used to motivate the development. The motivation establishes a required threshold order and probabilistic responses at each threshold. In addition, it is noted that AWW provide more or less complicated definitions of order. Here, the definition is embedded in the intended relative difficulty of the thresholds in the resolved design, where the intended order is clear. All other interpretations follow from this definition.
The Case for Rasch’s Simple Logistic Model at the Thresholds
Not only must Equation (3) hold in data, but to approach the analogy to measurement, distances between thresholds should be invariant with respect to differently proficient essays. For both conditions to hold, the responses need to be characterized by the SLM (Rasch, 1961; Wright, 1997). It is emphasized that justifying the application of the SLM in this way, as with the specification of threshold order, is a priori to any collection of data and that data may not follow the SLM. The model is taken as a hypothesis of the data, not merely their description.
The SLM takes the form
where
To explicate connections to the PRM, the two standard psychometrics response functions,
If there were a resolved design, then the SLM could be applied, a test of fit conducted, the estimates
An immediate implication of Figure 2 is that the relative order of response probabilities at the thresholds is specified. It consolidates that at every essay location, the probability of success at
Construction and Analyses of Response Spaces
The resolved design of the first two rows of Figure 1 is now bridged mathematically to the standard design of the last row in the figure.
The Experimentally Independent Response Space Ω
By definition
The joint probability of all response vectors is
and
The space
The space Ω and subspace Ω G for m = 2
Implications of the Required Threshold Order
Consider the intersection of the two dichotomizations of Rows 1 and 2 of Figure 1. The intersection of {Fail} versus {Pass or Distinction} and {Fail or Pass } versus {Distinction} is the decision {Pass}. The response pair is
In addition to Pass, two more pairs of responses in the resolved design can be placed on the continuum of Figure 3. First, if the decisions at
In contrast, the response pair
The Subspace
G
Let
be a vector of responses in which the first x variables, corresponding to the required ordering
Bridging the Resolved Design to the Standard Design With the Guttman Space
Within the response subspace
Then each sum
The Guttman Subspace Ω G and responses in the standard design
Because X
n
= x determines a unique vector in
The inferred category for each total score in the standard design, and compatible with Figure 3, is also shown in Table 3. It is reemphasized that because a Guttman vector could not be chosen without the specification of threshold order, the inference depends on the a priori specification of threshold order in the resolved design, which in turn arises from the specified category order in the standard design.
The Probability Structure of
G
The sum of probabilities of the Guttman vectors within Ω is
Because
ensuring
The factor H
n
does more than ensure that
The Doubly Conditioned Space
Consider a second-level subspace
Then,
where by definition, from Equation (5),
We return to Equations (5) and (11) shortly, but first note that Equation (11) implies a dichotomous response. Let
Then for
Just as a threshold on a continuum was specified for each dichotomous response in Ω and characterized the probability of a successful response with the SLM, the same can be specified in the spaces
where now
From the construction of the Guttman space, the variables
Returning to Equation (11), rather elegantly it is identical to Equation (5). In addition, because Equation (12) is an elaboration of Equation (11), it too is identical to Equation (5). That is,
The result in Equation (14) seems remarkable: It is central to the interpretation of reversed thresholds. It shows that for
If the probabilities in Equation (14) are both characterized by the SLM, then Equation (4) is identical to Equation (13), and the threshold values from the two sets of variables must also satisfy the following identities:
It is stressed that the identity in Equation (15) is not merely an interpretation that depends on specific judges making decisions. Instead, it follows from the mathematical structure of the response spaces, including the definition of the Guttman space based on threshold order.
Construction of the PRM Through the Guttman Subspace
Introducing explicitly the SLM of Equation (4) into the probability of a response in
Equation (16) simplifies to
where (a)
is the normalizing factor. Equation (17) is a general form of the PRM for the response of a single person to a single polytomous item with m thresholds and m+ 1 categories.
Again elegantly, the coefficient x of β
n
is exactly the score
Furthermore, because γ
n
includes explicitly all thresholds, the probability of a response in any category x is a function of all thresholds, and not just of the pair
The exponent in the numerator in Equation (17) of the PRM, irrespective of how it is motivated, is deceptive in that it appears to involve only a subset of the parameters governed by the total score. If all the parameters are written in full, then the Guttman structure is immediately present. Thus, let
is a Guttman vector for the score x compatible with a hypothesized threshold order. Then in vector notation, Equation (17) becomes
Though less efficient, but making the structure clear, the PRM may be written as
showing that the score x implies succeeding on the first x thresholds and simultaneously not succeeding on thresholds x+ 1 and above. This structure is a property of the model.
Alternatively, using matrix notation,
where
In the case of two thresholds, the exponent in Equation (21) is
in which the Guttman structure, relative to a hypothesized threshold order, is unmistakable. In addition, the parameters β n , δ1, δ2 are on the same continuum, being the same parameters as those in the resolved design.
Three consequences of the Guttman structure have been established. First, although the responses are in categories, and there is reference to categories, the only parameters in the model are the thresholds. Second, if a threshold is exceeded, then all other thresholds before it in the order are also exceeded, and following the first one failed, all subsequent ones are also failed. Thus, the Guttman structure, as a hypothesis of the threshold order, is integral to the model. However, there is no constraint that forces the threshold values to be ordered in the data. Third, using the PRM, the result in Equation (14) is confirmed. Thus, the conditional probability of a response in one of two adjacent categories in the PRM of Equation (17) simplifies to
which is exactly the SLM of Equation (4) of the experimentally independent space.
AWW imply an algebraic understanding Equation (23) as a consequence of Equation (11), but without realizing its implications. In particular, they seem to think that the result in Equation (23), which is linked to the standard design, is no longer relevant in the original resolved design. The identity of Equation (23) and Equation (4) in the two designs and corresponding response spaces is central to the inferences regarding disordered thresholds, which continues to be elaborated in the article.
Other Parameterizations of the PRM
First, however, and for completeness, we note other parameterizations of the PRM also considered by AWW. Equation (17) for a polytomous response in a standard design characterizes the response to a single item. If, as is usual, there is more than one item
where m j is the maximum score of item j. This form, where the thresholds can take different values for different items, is the so-called partial credit parameterization of the PRM and is the one used in the simulations shown later in the article.
The following reparameterization, with no change to the model, is used in AWW. Let
The
A further specialization in which all items have the same values of the thresholds
Equation (26) is known as the rating scale parameterization and may be relevant if all items have the same response format as in attitude surveys. This equation was the one first specified in Andrich (1978) in the resolution of the scoring functions and category coefficients in Andersen (1977). For the purpose of this article, these parameterizations are details for noting, as they are in AWW. For convenience, and without loss of generality, Equations (17) and (24) are used when considering a single item and multiple items, respectively.
The Latent Dichotomous Responses at the Thresholds
Figure 4 shows the category characteristic curves (CCCs), the probabilities of the responses of each category, for two ordered thresholds as in Figure 2. Thus, persons of low, medium, and high values on the continuum are, respectively, most likely to obtain scores of 0, 1, and 2. Figure 4 also shows in dotted lines the latent TCCs in the standard design from Equation (23). It is stressed that these dotted TCCs in Figure 4 are theoretically identical to the TCCs in Figure 2 characterizing the responses in the resolved design. To reinforce this point, Figure 5 juxtaposes Figures 2 and 4 where, for further clarity, the CCCs in the latter have been suppressed.

CCCs and latent TCCs in dotted lines for two ordered thresholds in the standard design

Simulation Study 1: Two Models of Analysis in Two Response Spaces
The key result above is that the values
Two sets of two independent assessments were simulated, taken as responses from two Expert and two Novice judges, for each person. Clearly, given that the data are simulated, there are no real judges involved; judges are used simply to help explicate the relationships between spaces and implications of threshold order. The two sets of judges are characterized by subscript
To emphasize the structure of the data analyzed, Table 4 shows the first 10 persons of the data matrix in which the left part of the table shows the set of dichotomous responses y nkj of each of the four judges for each person. They are responses from the resolved design. The right part shows the total scores
Excerpt of Responses According to the Bernoulli Responses and the Subset of Guttman Vectors
for the Experts and for the Novices for which the responses conform to the Guttman structure based on the hypothesized order of the thresholds.
In the demonstrations that follow, the Guttman subspace was defined, and therefore the scores constructed, according to the intended ordering of the thresholds. As stressed above, to choose a Guttman subspace, there is no alternative but to hypothesize a threshold order. Thus, for experts, it was based on a compatible ordering; for the Novices on an incompatible ordering. The non-Guttman vectors (which could not be identified with the standard design) were treated as missing data. The number of Guttman vectors in each set is shown in Table 5. As a result of the incompatible ordering, there are many less Guttman vectors for the Novice judges than for the experts, a feature returned to later in the article.
Estimates of Parameters Using the SLM and PRM in Data Simulated According to the SLM
Note. N = 20,000; β~N(0, 1).
Three questions arise from Simulation 1 to which answers will be given. First, will the PRM give correct threshold estimates for the Experts? Second, despite the Guttman vector being incompatible with the a priori threshold order, will the PRM give correct threshold estimates for the Novices? Third, what are some symptoms of reversed thresholds in the PRM? The answers to these three questions provide the understanding of the role of reversed thresholds in the PRM. As will be demonstrated, the answer to the last question is more complex than AWW’s conclusion, noted earlier, that reversed thresholds are simply a function of low frequencies of responses in some categories.
The responses from the two spaces, Ω and Ω G , were analyzed using the SLM of Equation (4) and the PRM of Equation (24), respectively.
We note that in the PRM, as in the SLM, within any set of responses
Table 5 shows the generating thresholds and their estimates from the two analyses. This table also shows that the estimates from the PRM were obtained from a smaller number of persons than from the SLM. As expected, it shows that the generating thresholds were recovered excellently using the SLM with a root mean square deviation (RMSD) of
More distinctively, and although estimated from a smaller number of responses, the RMSD = 0.028 for estimates using the PRM indicates virtually the same accuracy in recovering the thresholds.
Thus, Simulation 1 illustrates that a PRM analysis of responses in Ω G estimates identical threshold values as an SLM analysis of responses in the experimentally independent space Ω, including the values where the thresholds are reversed relative to the a priori order specification.
Despite the theoretical derivation above that proves that the PRM in the subset of responses in Ω G has the same thresholds as the SLM for all responses in the space Ω, it seems remarkable that, despite the incompatible scoring relative to the intended order, the thresholds for the Novices are estimated correctly. This identity in theory and in estimation leads to an understanding of the implications of disordered threshold estimates.
Establishing the Guttman Response Space From the Structure of the PRM
For the inference of this article, the significance of Equation (14), which shows the identity of dichotomous response variables in the experimentally independent space Ω and the corresponding latent dichotomous response variables in the doubly conditioned space
We begin now with the standard design in Row 4 of Figure 1, and bridge it to the resolved design in Rows 1 and 2 of Figure 1. The previous analysis, which proceeded from the resolved design, gives the cue how to proceed. Although for efficiency the same notation as above is used, the derivations are distinct from those in the above sections.
Let
Let the sequence of latent, dichotomous, Bernoulli variables
where
Because of the intended category ordering, x > x− 1, Z nk = 1, and Z nk = 0 are, respectively, the latent successful and unsuccessful dichotomous responses between categories x− 1 and x. For example, conditional on the latent dichotomous response of Pass or Distinction, Distinction is the successful one. Therefore, Equation (27) is a starting point that is consistent with the intended ordering of the categories. This starting point is referenced in AWW.
To infer a response space for X
n
= x, Equation (27) must be expressed as
From Equation (27), let
Theorem
Given the definitions of
where
and where, for convenience,
Equation (29) implies a very clear structure of successes and failures at successive thresholds x for each response
in which a sequence of x 1s is followed by a sequence of (m−x) 0s; moreover, no other vector of responses is compatible with
The vector
Thus, P
nk
defined in Equation (27) with an unknown space
and
Equation (32) is identical to Equation (11). It has a Guttman structure relative to the intended category order and is the structure that was introduced to bridge the resolved design to the standard design. Reintroducing the SLM from Equation (27) into Equation (32) gives the PRM of Equation (17). Therefore, the Theorem, together with the specification of the SLM for the latent dichotomous response at thresholds, establishes that the Guttman structure, relative to the putative ordering of the categories, is an integral part of the PRM.
Inferring an Experimentally Independent Response Space
From
Although the ordering of the categories implies a Guttman sequences of successes and failures at the thresholds, at this point there is no structural implication for the ordering of the values of the thresholds. To infer the required ordering of the threshold values, it is necessary to bridge the standard design to the resolved design using the Guttman subspace Ω G . Specifically, given the space Ω G and its probability structure, we infer the existence of a hypothetical, complete spaceΩof whichΩ G is a subspace.
From Equation (32) we have
from which H n > 0. Therefore,
Now consider the space
implied in Equation (33) is the subspace
Equation (35) confirms that
In summary, the inference is as follows: Given there is a Guttman response space Ω
G
composed of m latent, but dependent, Bernoulli variables with the probability structure of Equation (32),
For completeness, let these dichotomous responses in Ω be defined by Bernoulli random variables
where
Equation (36) is the bridge from the standard design to the resolved design and is identical to Equation (14), which was the bridge from the resolved design to the standard design.
It is stressed that if an experiment were conducted in which a set of judges worked independently according to the resolved design of Table 2, and the same or other judges worked within the standard design of Table 1 for the same essays and the same definition of the categories, then there is nothing that guarantees that data collected in two designs would give identical empirical threshold estimates. However, that is not the point of the above inference. The point is to infer a hypothetical, experimentally independent response space Ω with
Thus, as anticipated early in the article, the inference is as if the data at hand, collected in the standard design with its probabilities of success at the thresholds, were collected in a resolved design with the same probabilities of success at the same thresholds. This is a mathematical inference from a mathematical structure.
The Required Success Rate Order in the Standard Design
As noted above, Equation (32), which showed that the inferred response space was Ω
G
relative to the ordering of the categories, specified nothing about the values of the inferred thresholds. And this lack of specification seems to confuse AWW. However, in the inferred space Ω, the required and intended order of the threshold estimates is readily specified, namely,
Deductions from mathematical models are tautological. However, the deductions can provide genuinely new insights. The derivation that, for categories to demonstrate empirical order, it is necessary that
With the above development, misunderstandings by AWW in the following statements can be exposed. For example, they state, While this statement seems eminently reasonable it raises two questions. Does this statement, and the derivation that corresponds to it, impose an order requirement on the values of the tau parameters? Does the statement clarify what the latent processes are? What Andrich repeatedly states but neither proves nor logically defends is that an ordering of the tau parameters is related to category ordering or is imposed through the Guttman structure. (AWW, 2012, p. 9)
It can now be highlighted from the above development that there is in fact a disjunction between the values and the Guttman structure of the thresholds. Because they assert that I have not shown that the order of the threshold values is a property of the model, it seems that AWW cannot, or will not, make a distinction between the structure of the thresholds in the PRM, which requires a hypothesized order, and values of the thresholds, which may or may not violate the hypothesized order. Indeed, there is nothing in the structure, nor in the estimation process, that forces the values of thresholds estimates to be ordered. If there were such a feature, then the ordering would be a property of the model (as in the graded response model), and not a property of the data, and therefore could not disclose problems with the category ordering. The consistent argument made previously, and repeated here, is for evidence that the thresholds are ordered in the data, as disclosed by the PRM. The PRM can make such a disclosure because the Guttman structure inherent in it characterizes the intended ordering of the thresholds, which in turn reflects the intended ordering of the categories. Then, the freedom of the threshold values in the PRM to reflect the property of the data in Ω G is identical to freedom of the threshold values in the SLM to reflect the property of the data in the full space Ω. This freedom for the values to reflect a property of the data, including a violation of the hypothesized order of the thresholds, was illustrated in Simulation 1.
Category Characteristic Curves and Disordered Thresholds
Figure 4 showed the CCCs and the latent TCCs of the PRM for three categories with ordered thresholds conceived to be from Expert raters in Simulation 1. Figure 6 shows the CCCs for the PRM for the Novice raters with reversed thresholds, together with their latent TCCs. AWW argue that such a figure provides no evidence of anomaly in the operation of the category order. Their argument is examined here.

CCCs and latent TCCs in dotted lines for the two reversed thresholds in the standard design
Consider a person with a proficiency located exactly at
AWW see no inconsistency in this case and write, At the point of equal probability for pass and distinction we are not tossing up if a student should be a pass or distinction: In fact we are more confident that they are a fail (see Figure 5). (AWW, 2012, p. 12)
The inference in the above quote goes beyond anything that can be inferred from the PRM. The model is a static model and has nothing to do with “tossing up” to make a decision, or deciding that in any case “we are more confident that the person is a Fail.”The model only describes the inferred probability distribution of responses of a person at a particular location to the item implied by the data. Although there are no replications of any person to any one item, the PRM estimates the probability distribution of responses to the item as if the person were to respond to the item, independently many times. To expand on this point, Figure 6 shows three vertical lines with dots on the curves, which characterize the distribution of responses for three locations, including one that is at the threshold of Pass and Distinction. At each location on the continuum, there is such a triplet of points characterizing the probability distribution, which sums to 1. And in this distribution, a person who has a 50% chance of a Distinction relative to Pass (and therefore implies substantial proficiency) simultaneously has a 78% chance of a Fail (which therefore implies substantial lack of proficiency). At the other end, a person who has 50% chance of a Fail relative to the Pass has 78% chance of a Distinction. These implications are contradictory. To make the point concrete, it is as if, based on all the evidence, an instrument for measuring height implies that the person on the margin of 180 cm and 181 cm in height (180.5 cm) is simultaneously most likely to be 179 cm in height! It would be understood that the measuring instrument has produced a contradiction and would be examined.
In contrast, there is no such contradiction in Figure 4 with ordered thresholds, where the respective probabilities for
Third, in the space Ω in which the inferred responses are independent at the thresholds, the person with location
One might wonder, to paraphrase AWW in the quote above, how a PhD student would interpret similar remarks to theirs—namely, that although as three examiners they had a consensus that the dissertation had the quality to be at the margin of a Distinction relative to a Pass, because of their understanding of conditional probabilities, they were already “more confident” that the dissertation’s quality was a Fail, and so declared it a Fail!
Frequencies of Responses in Categories and Reversed Threshold Estimates
The article now turns to the second implication in AWW, which concerns the role of frequencies and reversed threshold estimates. As noted earlier, AWW indicate that we examine parameter estimation and show that for any given set of abilities the relative frequency of the number of responses in each category of an item is the only determinant of whether the estimated parameters are ordered in value or not. (AWW, 2012, p. 2)
They subsequently provide an equation, using a joint maximum likelihood method of estimation, from which they write further: For a given set of students, the parameter estimates for δ
i
, τi1, and τi2 depend solely on the number of observations in each category; they are completely independent of the abilities of the students who respond in each category. (AWW, 2012, p. 15)
That, by definition, the estimation of the item parameters in the class of Rasch models (Andersen, 1977; Rasch, 1961) can be made independent of the person parameters is so well known that it hardly needs reinforcing, even though AWW’s joint maximum likelihood equations, which have person parameters on the right side, do not make this self-evident. Therefore, in their emphasis of frequencies in the equation, AWW must also be saying something else. What they are saying can be gleaned from their next mention of frequencies in conjunction with their Figure 8, which is similar to Figure 6 of this article with reversed thresholds. They state, The item characteristic curves in Figure 10 show a low curve for score Category 1, reflecting the low frequency of responses for this score category and resulting in disordered thresholds (3.75 and 0.07). (AWW, 2012, p. 28)
(Their reference to “Item characteristic curves” is what are referred to in this article as CCCs.) The implication, consistent with other comments in their article, is that a very small frequency in a category in the sample as a whole leads to reversed threshold estimates. If this is AWW’s interpretation, then it results from a superficial analysis of the role of frequencies in the PRM. In addition, even if their comments needed no qualification, they are merely descriptive. That is, AWW do not go on to explain why such a low frequency of a score of 1 produces the reversed thresholds. Before returning to explain the reason for reversed thresholds as a violation of the hypothesized Guttman structure and threshold order, simulations that demonstrate that reversed thresholds are not simply a result of low frequencies are presented.
Examples of Low Frequency and Correctly Ordered Threshold Estimates and High Frequency and Reversed Threshold Estimates
Simulations 2 and 3 show that categories with very low frequencies (as low as 2 but in principle even 0) need not produce reversed threshold estimates, and Simulation 4 shows that categories with large frequencies can. To show that where the frequencies might be low in the order is irrelevant to the conclusion, one example with low frequencies in the extreme category, and one with low frequencies in the middle category, are included.
Simulation 2
Simulation 2 had small frequencies in the extreme categories. It had 10 items with four categories (three ordered thresholds). The distance between all pairs of successive thresholds within an item was 3 logits and the sum of all thresholds was 0.0, giving an origin of 0.0 for both the generating parameters and their estimates. Two thousand persons,
All the category frequencies and threshold values and their estimates are shown in Appendix A. Table 6 shows frequencies for Item 10, in which the frequency of the last category is only 2. With conditional estimation, the person parameter distribution played no role. Taking all thresholds together, the correlation between the parameters and their estimates is 0.999. In particular, for Item 10, in which the last category has a frequency of just 2, and in which the threshold location is 6.00, Table 7 shows that its estimate using RUMM2030 is 6.52. The standard error of the estimate for this threshold is 0.802, giving the standardized residual of z = −0.648, a value well within an acceptable range of no significant difference. Thus, despite the very small frequency, there is no reversed threshold estimate.
Frequencies in Response Categories for Simulation 2
Generating Thresholds and Their Estimates in Simulation 2
Consistent with the simulation parameters, Figure 7 shows the person distribution is not well aligned with the distribution of the thresholds. This misalignment results in few responses in the extreme categories of the difficult Item 10. However, because there is no structural problem in the responses, it does not result in reversed thresholds.

Person and threshold distributions for Simulation 2
Simulation 3
Simulation 3 had small frequencies in a middle category. It had 8 items each with 9 categories (8 thresholds) with a threshold mean of 0. Such large numbers of categories are present in some criteria in national writing assessments in Australia. The 1,000 persons were generated from two subpopulations of 500 persons,
Frequencies in Response Categories for Item 5 for Simulation 3
Thresholds Estimates for Item 5 of Simulation 3
In Table 8, the category with score 4 has a low frequency with only 11 (1.1%) responses. Nevertheless, Table 9 shows that its thresholds are ordered correctly. Figure 8 shows the estimated person distribution is estimated to be clearly bimodal, as generated. Again, because of the lack of persons in the middle of the continuum, there are relatively few responses in the middle category of Item 5. However, also again, because this alignment is the cause of a low frequency, and there are no structural problems in the responses, the threshold estimates are not reversed.

Person and threshold distributions for Simulation 3
In summary, Simulations 2 and 3 show that both relatively, and absolutely, low frequencies do not necessarily lead to disordered thresholds.
Simulation 4
Simulation 4 shows that large frequencies do not preclude disordered thresholds. It is based on a real example that concerned ratings of muscle tone in five putatively ordered categories according to a clinician-rated scale and is detailed in Andrich (2011). Those details are not reproduced here. However, all items showed that thresholds three and four were reversed, and on the average, the reversal was by 0.34 logits.
To ensure that the data for this article had no misfit features that might be present in real data, and that disordered thresholds are not explained from any particular judging process, they were simulated according to the PRM. The generating parameters were similar to the estimates from the real data. Simulation 4 has 1,000 persons distributed as
Frequencies in Response Categories for Simulation 4
Thresholds and Their Estimates in Simulation 4

Person and threshold distributions for Simulation 4
Table 10 shows that all categories have substantially large frequencies, and in particular, that the frequencies in the last category are greater for each item than those in the first category. In addition, Items 1 to 4 have more responses in the last category than in the second last one, and Items 6 to 8 have it the other way around. This pattern is as in the real example. Nevertheless, Table 11 shows that the threshold estimates are reversed only for Thresholds 3 and 4, and that they are reversed for all items, consistent with the real example. The average size of the reversal in the estimates is 0.4 logits. Again, small frequencies are not the cause of these reversed thresholds. For completeness, Figure 9 shows that close to half of the people are located approximately above 0.5 logits where the threshold estimates are reversed. A structural response anomaly is present in this case.
To summarize the point of the above simulations, estimates of the thresholds are independent of the distribution of persons in the PRM, but the distribution of frequencies themselves is very much dependent on the distribution of the persons relative to the distribution of the thresholds. And depending on the misalignment of persons and thresholds, there might be large or small frequencies in categories. If misalignment is the only source of low frequencies, threshold estimates are not disordered. Disordered thresholds can also appear when there is no low frequency in any category, in which case there are structural problems in the responses that, in real data, need a substantive explanation.
The Distribution Leading to Disordered Thresholds and the Test of Fit
Having demonstrated that a low frequency in a category of the whole sample is not the source of disordered thresholds, the article now explains the sense in which a relatively low frequency produces reversed thresholds, and in the process, it demonstrates that the test of fit used by AWW to justify disordered thresholds is irrelevant.
Simulation 5
To facilitate the above purposes with sound estimates of person locations, Simulation 5 had 26 dichotomous items and two polytomous items with maximum scores of 2. The dichotomous item difficulties were distributed in equal intervals between −3 and 3. The distribution of 20,000 persons was
Estimates of Parameters of Two Polytomous Items in Simulation 5
Note. N =20,000;
Frequencies (and Proportions) of Responses for Items 27 and 28 in Simulation 5
For Item 28, the frequencies are similar to those of AWW’s example in their Figure 7, except that the frequency of the Score 1, although low, is greater than that of Score 2 and therefore is not the lowest frequency as in their example. In addition, the frequency of Score 0 for Item 27 is lower than those of Scores 1 and 2 of Item 28. Despite these relative frequencies, the threshold estimates of Item 27 are ordered and those of Item 28 are not, confirming that low frequencies in the data as a whole cannot explain disordered threshold estimates. The explanation requires a closer examination of the distribution of the responses of each person, and not the data as a whole.
Bimodal Response Distributions for a Single Person
Although it can be shown algebraically, from Figure 4 with ordered thresholds, it is evident that for a location between the thresholds,
Expressed in terms of parameters, and therefore generalized, the ratio
which is independent of the person parameter. In general, for m > 2,
Equation (37), which connects a distribution of responses with thresholds of an item explicitly, and which is the only one that seems to do so, characterizes a single person. However, it also characterizes the distribution of every person. This has two immediate implications. First, as noted earlier, that every threshold estimate is affected by all frequencies in a complex way, and that the distribution of Equation (37) must be distinguished from a summary distribution and frequency of responses of all persons in a sample of data. In their justification of low frequencies in the data set as a whole as the cause of reversed thresholds, AWW do not make this distinction.
Second, each person has only one response to an item, and the probability distribution in Equation (37) is inferred from responses to all items. Therefore, to make any comparisons between observed proportions of responses (which are counterparts to the probabilities) and the threshold estimates using Equation (37), it is necessary to take persons with homogeneous estimates rather than the whole sample. In the PRM, classifying persons by the total score, their sufficient statistic is ideal. However, classifying persons into class intervals based on total scores is often more practical. This method is used in tests of fit and is used in particular by AWW in the example they present.
Consider now Item 28 whose parameters are shown in Table 12. According to Equation (37),
Equation (38) holds in general for all persons. In particular, it implies that the inferred response distribution of any person between the two thresholds,
Disordered Threshold Estimates and the Test of Fit
The contradictory implications of a bimodal distribution also provide interpretative problems for the test of fit used by AWW. In particular, it is necessary to invoke the caution drummed into students in elementary statistics classes that it is misleading to characterize a bimodal distribution by a single mean. For example, it would be misleading to characterize the bimodal distribution of persons in Figure 8 with the mean of 0.021. And there is no reason to revoke this principle just because it appears in a less obvious guise in a psychometric distribution of a homogenous group of persons’ responses to a single item. To consolidate how the class interval mean is misleading when thresholds are disordered, Table 14 shows the observed means and expected values in 9 class intervals for Item 28. The table also shows the ratio of Equation (37) for both estimated probabilities and the observed proportions in each class interval.
Nine Class Intervals for Item 28 With Disordered Thresholds in Simulation 5
Note. OM = mean of the class interval; EV = expect value,
In addition, Figure 10 shows the EVCs and observed means in the class intervals. In Figure 10, the observed means are close to their expected values, and therefore the responses are taken to fit the model. This is the evidence used by AWW to justify that there is no contradiction in reversed threshold estimates. However, as argued below, this evidence is meaningless for their purpose.

EVC of Item 28 with reversed thresholds of Simulation 5
First, we confirm that the observed mean in a class interval is meaningless in characterizing a distribution of responses. Table 14 for Item 28 shows that the ratio of Equation (37) is less than 1 for all class intervals in which it can be calculated, and therefore implies reversed thresholds. It is less than 1 for both observed proportions and the estimated probabilities. Because Item 28 is difficult, all but one ratio is undefined for the first four class intervals. For persons within the range of the thresholds, the observed distribution is explicitly bimodal. For example, Class Interval 9 in Table 14 has a mean of 2.206 close to the mean of the two thresholds, (2.65 + 1.59)/2 = 2.121 (from Table 13), and the observed proportions, notated “OBS P” and shaded in Table 14, are bimodal. This bimodality with more scores of both 0 and 2 than of 1 renders the observed mean misleading in characterizing the distribution.
Second, we show that the closeness of observed means to their expected values from threshold estimates is irrelevant to threshold order. Whether disordered or not, threshold estimates reflect a property of the data. Then these threshold estimates are used to obtain the distribution from the model and the expected values. Illustratively, for Class Interval 9, the estimated distribution is
Thus, for two reasons, first that a mean with a bimodal distribution of responses is misleading, and second that the thresholds estimates themselves are used to calculate the expected value, the evidence of fit displayed in AWW’s Figure 9 (and Figure 10 in this article) has no bearing on the interpretation of reversed threshold estimates. Bimodal distributions do not appear with Item 27, which has ordered thresholds.
Frequencies and Reversed Thresholds
The above analysis of bimodal distributions for each person implied by reversed thresholds also allows a qualification of AWW’s assertion that low frequencies lead to reversed threshold estimates. For this qualification, consider again Figure 6 of Simulation 1, which has reversed threshold estimates. It can be inferred from this figure that reversed thresholds appear when, based on all the evidence from the data, the very people (locations) who should obtain a particular score, do not do so, but instead obtain scores on either side simultaneously with greater frequency. In summary, a relatively low observed frequency of a particular category, and simultaneously inflated frequencies on either side, from those persons who are expected to have the highest frequency causes disordered thresholds. This qualification of comments made by AWW is essential to understanding the role of frequencies on threshold reversals.
Figure 6 reflected the distribution from the simulated Novices in Simulation 1. The disordered thresholds arose because the incompatibility between the threshold order and the hypothesized Guttman vector meant that those persons who were expected to score 1 most frequently were not selected for analysis in with the PRM. Instead, not only less persons but also those for whom the hypothesized Guttman vector was incompatible with the threshold order were selected, resulting in the disordered threshold estimates.
The AWW Example With Reversed Threshold Estimates
Although the emphasis in this article is on the use of the PRM for a rating format, the scoring structure of the example of a partial credit item provided by AWW in their Figures 6 to 9 is analyzed briefly from the perspective of this article. The options for the scoring are reproduced in Table 15, together with the threshold estimates. Students were instructed to show their work.
Scoring Key in Example in AWW
AWW claim that, despite severely disordered thresholds, there are no problems in the ordering. Nevertheless, consider this example from the perspective of the location of a person who is exactly at the threshold
It is not enough to describe the reversed thresholds to be a result of a low frequency of responses in the score of 1 for the whole sample; instead as argued consistently from Andrich (1979) to Andrich (2011), it is necessary to seek a substantive explanation for such an anomaly. For example, the constructors of the item need to be informed of the anomaly, which they may be able to explain by having discussions with markers, and so on.
However, in the AWW example the potential problem in the scoring seems so blatant that there is no need to go to the markers to form a first hypothesis. According to the marking key, a score of 2 is given for the correct answer with work shown, but only a score of 1 for a correct answer with no work shown. On the assumption that cheating was sufficiently controlled for by invigilation during the assessment, a correct answer, however it was obtained, suggests the same proficiency as having a correct answer with work shown.
In assessments of this kind, students are encouraged to show working, which, in case it is essentially correct but there is, say, a calculation error, they may gain partial credit. In fact, the above scoring key does just that with the option “Correct method but with calculation error,” which also is given a score of 1. Therefore, to give only a mark of 1 for a correct answer when no working is shown is giving only partial credit as a punishment for not following instructions. That is, a second variable, irrelevant to mathematical proficiency, is introduced in the marking key. Therefore, it seems not surprising that someone who has 0.5 probability of answering the question correctly without showing working, in fact, has a 0.95 probability of answering it correctly showing the working. In addition, it seems that “Correct with no work shown” is qualitatively different from the other two options that also obtain the partial credit score of 1, that is “Correct method but with calculation error,” and particularly, “30 with calculation leading to 30,” which is in fact incorrect. This raises a complementary problem with the scoring key—students who provide the wrong answer of 30 showing calculations achieve the same score as students who actually provide the right answer without showing their work. On reflection, this seems unfair on the students who answer correctly but do not show their work. If this were a high-stakes test, say entry to university for highly selective courses in which less than one scaled mark can make a difference, it suggests that this scoring key would lead to appeals and be untenable.
AWW use the example to argue that there is no problem demonstrated by reversed thresholds. It seems that they are so wedded to the evidence of the test of fit, shown to be irrelevant, that none of them has studied the marking key itself. Even without the reversed threshold estimates, a close examination of the key would suggest exactly the problems that are manifested by the PRM analysis. This article shows that their example is excellent in demonstrating the power of the PRM to expose, with disordered thresholds, problems in the scoring structure of items. In this case, at least two problems can be hypothesized: first, that one of the response options given a partial credit score of 1 demonstrates the same proficiency as that given the maximum score of 2, and second, that two responses that clearly do not show the proficiency of a correct response are given the same score as one of the correct responses. Of course, the above explanation would need to be tested by rescoring performances.
Discussion
In the applications of the SLM to dichotomously scored items, the difficulties of the items (their thresholds in the nomenclature of this article) are integral to the substantive understanding of increasing values of the variable. Placing persons on the same continuum strengthens the interpretations. Wilson, in particular, makes much of this possibility and has used the term Wright map to characterize such a display (Wilson, 2005). Unfortunately, the item order is generally not spelled out in advance and is used mostly post hoc to confirm that the ordering makes substantive sense. Of course, even if the confirmation is post hoc, the power of the analysis is that the item location estimates may disconfirm a sound substantive order.
In the case of items with more than two ordered categories, there is immediately an a priori condition of order that should reflect an understanding of increasing values of the variable. The unequivocal a priori ordering of categories, together with possible empirical evidence to the contrary, provides the opportunity to use the PRM in more than a statistically descriptive way. As shown in the article, this possibility rests on being able to make inferences from the PRM as if responses at the thresholds were experimentally independent and analyzed with the SLM.
This article has addressed, with a detailed derivation of the PRM and with simulation studies, the three major contentions of AWW indicated at the beginning of the article: (a) that category ordering is independent of the threshold order and that the Guttman structure with order of the thresholds is not part of the PRM, (b) that low frequency is the source of the reversed thresholds, and (c) that evidence of the kind of test of fit justifies reversed thresholds. The article rebuts the first two contentions and demonstrated that the test of fit AWW use is irrelevant to their case.
In challenging the relevance of order in the PRM, AWW seem to want, as with the graded response model, that the order of the values of the thresholds be a property of the model rather than only a property of the data. They seem to think that just because estimates of threshold values can be reversed, the structure of the model is not compatible with the required order of the thresholds. This article shows that the choice of the Guttman vector rests on having a hypothesized order of the thresholds based on the intended ordering of the categories and that reversed threshold estimates implies a rejection of the hypothesized order.
In printing out the thresholds from the graded response model, as they do in Figure 7 of their partial credit item, rather than relying on the thresholds from the PRM with which they analyze the data, AWW must sense that there is something wrong. Like Masters and Wright (1997), who state that it is a disadvantage of the method when no region of the continuum is defined as in Figure 6 of this article, AWW must see reversals as a problem with the model rather than with the data. They effectively abandon the PRM in favor of the graded response model whose thresholds are never reversed, no matter what problems there might be in the empirical operation of the categories. Switching between these two models in particular, which are structurally incompatible, confirms a purely data-descriptive perspective in the application of models. For example, it has been known since 1966 (Rasch, 1966) that summing probabilities of adjacent categories in the PRM destroys the sufficiency of the model (Andrich, 1978, 1995; Roskam & Jansen, 1989), yet that is exactly what AWW do when they calculate the thresholds for the graded response model from the probabilities estimated from the PRM.
AWW make another summary comment that is irrelevant to the mathematics of the model—they write, We do not dismiss the observation that if certain psychological processes are assumed, then a numerical ordering of the tau (threshold) parameters would be a substantive requirement. (AWW, 2012, p. 24)
They seem to imply that the mathematics of the model, and the derivation that used judges illustratively, depends on psychological processes. It does not—all the examples presented in this article used simulated data to illustrate the mathematical relationships leading to reversed thresholds. Psychological judging processes do not have to be understood to require that a rate of success at Pass be greater than a rate of success at Distinction, conditionally (in the space
Again, recognizing that there is something wrong with reversed thresholds, they write, We do not dismiss the possibility that disordered estimated parameters may well be an indicator of a problem with an item. (AWW, 2012, p. 24)
And they give particular examples, low frequency of a category, local dependence, and differences in discrimination at the thresholds. This article explains the qualification to the low frequency in detail, and Andrich (1983, 1985a) explained the relationship between local dependence and threshold estimates as well as the case of different discriminations at thresholds. These particular cases were pursued because of a recognition that reversed thresholds presented an anomaly, and which AWW apparently accept as such in these cases. However, from the AWW perspective, there would have been no reason to search for an explanation in these cases. Admitting that reversed thresholds might be an anomaly in an item is inconsistent with AWW’s argument that there is no inherent, structural problem with reversed threshold estimates. Thus, there may be problems that have never been thought of, some may permeate the whole data set as in Simulation 4, and some may simply be idiosyncratic to a particular item, as in the one they parade as having no problems but which seems to have major problems with the partial credit scoring key. The argument from the mathematical structure and implications of the PRM for rating, and ordered categorical data in general, in this article and elsewhere, is that to confirm that the categories are working as intended, the threshold estimates need to be ordered and that any item in which they are not ordered needs to be examined to try to understand the source of the disorder.
Finally, AWW write, What we are concerned about, however, is proliferation of the view that disordered parameters are indicative of a set of items that are not working because the categories are not ordered. (AWW, 2012, p. 24)
What I am concerned about is that Adams, Wu, and Wilson, who use Rasch models, do not appreciate the elegance, the coherence, and the power of the PRM to disclose that categories intended to be ordered are not operating as required. I am concerned that they use statistical sophistry, exemplified by the incongruous statement that if one understands conditional probabilities, then one understands why a performance of sufficiently high proficiency to be at the margin of a Pass and Distinction is simultaneously of such poor proficiency that it is more likely to be classified as a Fail, to justify ignoring reversed threshold estimates. Unfortunately, this sophistry potentially hides problems, as in the example they present, in every item with putatively ordered categories with reversed threshold estimates.
Conclusion and Further Exposition
As analyzed in the article, AWW use a test of fit that is irrelevant to the implications of threshold order. It is acknowledged, of course, that a satisfactory test of fit confirms that the model describes the data and that there is a sense in which it is better that the data fit the model than when they do not. In depending so much on a descriptive account of the data based on fit, without giving primacy to the substantive understanding of parameters in a context, AWW adopt a common statistical modeling paradigm set in place by the founder of modern statistics, Karl Pearson (MacKenzie, 1981; Norton, 1978). This tradition, importantly, contrasts with the paradigm of R. A. Fisher in which statistical methods are applied to specifically designed experiments in which the substantive meaning of variables is paramount (Box, 1978; MacKenzie, 1981). Although the argument in this article arises from Fisher’s paradigm, there is not the space to consider further the implications of the differences between these two paradigms in a measurement context. These implications have been broached for the case of ordered response categories in Andrich (2011). They are examined in more detail in Andrich (in press).
Footnotes
Appendix A
Thresholds Estimates in Simulation 3
| Item |
|
|
|
|
|
|
|
|
|---|---|---|---|---|---|---|---|---|
| 1 | −5.42 | −3.77 | −2.41 | −1.22 | −0.07 | 1.17 | 2.62 | 4.41 |
| 2 | −5.12 | −3.64 | −2.29 | −1.02 | 0.23 | 1.49 | 2.84 | 4.31 |
| 3 | −5.09 | −3.56 | −2.14 | −0.79 | 0.52 | 1.81 | 3.10 | 4.43 |
| 4 | −4.82 | −3.30 | −1.99 | −0.79 | 0.40 | 1.68 | 3.14 | 4.88 |
| 5 | −4.83 | −3.10 | −1.64 | −0.37 | 0.81 | 2.01 | 3.33 | 4.86 |
| 6 | −4.78 | −2.90 | −1.42 | −0.19 | 0.94 | 2.13 | 3.54 | 5.31 |
| 7 | −4.48 | −2.93 | −1.52 | −0.20 | 1.07 | 2.36 | 3.69 | 5.13 |
| 8 | −4.19 | −2.83 | −1.51 | −0.21 | 1.09 | 2.40 | 3.74 | 5.12 |
Acknowledgements
Many discussions with Guanzhong Luo and Irene Styles over many years, and Stephen Humphry more recently, have contributed to improving the original demonstration of the threshold structure of the unidimensional, polytomous Rasch model, the special cases of which are known as the rating scale and the partial credit models. Stephen Humphry, Pender Pedler, Tim Dunne, Barry Sheridan, Ida Marais, Irene Styles, and Josh McGrane provided valuable comments on earlier versions of this article.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This research was in part funded by Australian Research Council Grants and by the Australian Curriculum, Assessment and Reporting Authority and the Curriculum Council of Western Australia as Industry Partners and by Pearson.
