Abstract
Extreme response style is the tendency of individuals to prefer the extreme categories of a rating scale irrespective of item content. It has been shown repeatedly that individual response style differences affect the reliability and validity of item responses and should, therefore, be considered carefully. To account for extreme response style (ERS) in ordered categorical item responses, it has been proposed to model responder-specific sets of category thresholds in connection with established polytomous item response models. An elegant approach to achieve this is to introduce a responder-specific scaling factor that modifies intervals between thresholds. By individually expanding or contracting intervals between thresholds, preferences for selecting either the outer or inner response categories can be modeled. However, for a responder-specific scaling factor to appropriately account for ERS, there are two important aspects that have not been considered previously and which, if ignored, will lead to questionable model properties. Specifically, the centering of threshold parameters and the type of category probability logit need to be considered carefully. In the present article, a scaled threshold model is proposed, which accounts for these considerations. Instructions on model fitting are given together with SAS PROC NLMIXED program code, and the model’s application and interpretation is demonstrated using simulation studies and two empirical examples.
1. Introduction
Cronbach (1946) observed that responses to questionnaire items are affected by individual response styles influencing the response patterns independently from item content. One well-known response style, the so-called extreme response style (ERS), denotes the tendency of an individual to prefer the outer response categories, thus using the rating scale as if it were dichotomous (e.g., Greenleaf, 1992). Relevance and impact of ERS regarding the objectivity and validity of questionnaire data have been verified repeatedly (e.g., Baumgartner & Steenkamp, 2001; De Jong, Steenkamp, Fox, & Baumgartner, 2008; Khorramdel & von Davier, 2014; Naemi, Beal, & Payne, 2009; Weijters, Schillewaert, & Geuens, 2008). Thus, if not accounted for, ERS is likely to compromise results of studies analyzing item responses at least to some extent.
In recent years, several latent trait models have been introduced accounting for ERS in ordered categorical responses. Essentially, there are two main modeling approaches, although models within these approaches are rather heterogeneous themselves. The first approach, which will not be considered further in this article, uses a dichotomous conceptualization of ERS. Specifically, a dichotomous indicator variable, indicating whether an extreme response category was selected or not, is used to measure ERS (e.g., Böckenholt, 2012; Khorramdel & von Davier, 2014; Plieninger & Meiser, 2014). The second approach, which is considered in detail in the following, models ERS by modifying the spacing between responses for each responder individually. Individuals having a decreased spacing between responses assign their responses within a limited range, whereas responders with an increased spacing are more likely to select the extreme response categories. In contrast to the first approach, an extreme response is considered as a consequence of an increased spacing rather than the exclusive preference for the outer versus inner categories of a response scale.
Modifying the spacing between responses can be achieved in various ways. One general approach assumes heterogeneous thresholds (e.g., Javaras & Ripley, 2007; Jin & Wang, 2014; Johnson, 2003; Rossi, Gilula, & Allenby, 2001). The term heterogeneous thresholds means that instead of one set of thresholds applying to all responders, an individual set of thresholds is assumed for each responder. In this way, thresholds are adjusted to account for individual ERS.
A parsimonious way to model heterogeneous thresholds is to include an additional latent variable measuring ERS to an established polytomous item response model. This additional variable operates as a responder-specific scaling factor contracting or expanding intervals between thresholds to adjust differences in response spacing for each individual. However, while scaling thresholds is a sensible and useful approach for modeling ERS, this approach cannot be used arbitrarily in connection with any polytomous item response model as previous research on heterogeneous thresholds might have suggested. Instead, it is important to carefully consider the type of category probability logit and also the required constraint on the thresholds if they are modeled in connection with a scaling factor.
The use of a responder-specific scaling factors is not novel and has been used in other forms to model idiosyncratic response behavior. Ferrando (2009) proposed a model for ordered categorical items with responder-specific residual scales, and Lubbe and Schuster (2017) proposed a model that assumes responder-specific factor loadings. These models, although conceptually similar, achieve different purposes when compared to a model with heterogeneous thresholds. Differences between them and the model proposed in this article will be outlined below.
The article is organized as follows. First, we introduce a new model using scaled thresholds. Second, we demonstrate why we consider this new approach of modeling scaled thresholds as superior by contrasting it to other previously proposed model specifications. Third, we compare the present model to other types of scaling factor models, specifically those of Ferrando (2009) and Lubbe and Schuster (2017), and point out structural differences. Fourth, we demonstrate the small sample parameter recovery of our model using Monte Carlo simulation. Fifth, we analyze two empirical samples of questionnaire responses. Finally, results are summarized and discussed.
2. The Scaled Threshold Model (STM)
In the following, a model is considered that uses a responder-specific scaling factor accounting for ERS on the basis of the graded response model (GRM; Samejima, 1969). According to the GRM, the probability for individual i to select for q-categorical item j response category k is
where
For modeling ERS, thresholds parameters
where
In this decomposition of the
The distributional assumptions for the latent variables are
It follows that ω
i
is log-normally distributed having an expected value of 1.0 and a variance of
Responder-specific thresholds (
It is natural to consider extreme responders and midpoint responders as opposite extremes of the same continuum. For defining the extremes of scaling factor (ω i ) accordingly, the positioning of the β parameters is important. Specifically, β j needs to be placed in the location that corresponds to the center of the response scale as it is perceived by the responder. Only then will responses converge toward the middle categories as ω i increases and will be shifted away from the scale middle toward the extreme response categories as ω i decreases.
Positioning β
j
in the response scale’s perceived middle is achieved by fixing the median of threshold deviations
To illustrate the manner in which scaled thresholds account for different response styles, consider a four-categorical scale with labels strongly disagree, disagree, agree, and strongly agree. In this case, β
j
characterizes the middle threshold, that is, the threshold between categories disagree and agree. Figure 1 gives the category characteristic curves (CCCs) of responders with three different scaling factors and the same arbitrary trait (θ) for a four-categorical item (j). Item parameters are

Category characteristic curves of the jth item assuming three responder-specific scaling factors. The plots show how the thresholds are pulled toward the perceived middle
The first scaling factor value of
For a response scale with an odd number of categories, the principle is equivalent. As an example consider a five-categorical response scale with response labels strongly disagree, disagree, neutral, agree, and strongly agree. For this scale, the β parameter is positioned in the middle of the neutral category. Again, responders with small scaling factor values prefer the extreme response categories, whereas responders with large ω values prefer the neutral category.
3. Alternative Model Specifications
In the following, we contrast the current model specifications to potential modifications that might seem natural but will limit the model’s capability to account for ERS. Specifically, we consider modifying the threshold constraint and the type of category probability logit. These modifications are of particular relevance because they have been used previously for scaling factor models.
3.1. Threshold Constraint
The median constraint on threshold deviations (
As was outlined, by positioning β
j
in the response scale’s perceived middle, extreme and midpoint responders are defined as opposite extremes on the ω continuum. More specifically, only in this case will the probability of selecting a middle category consistently increase as ω
i
increases. To ensure that β
j
is always positioned in the response scale’s middle, the median constraint on threshold deviations
To illustrate, let’s consider a set of thresholds pertaining to a five-categorical response scale. Specifically, we assume thresholds of

β parameter positions of a five-categorical scale for median and sum constrained threshold deviations.
If the median constraint is used, β j lies between both thresholds surrounding the neutral category. If, however, the sum constraint is applied, β j lies in the disagree category. While the different threshold deviation constraints are of little consequence for responders with average ω values, it becomes relevant for responders with large ω values, who are considered midpoint responders. Increasing the probability of selecting the middle category can be achieved by expanding the interval between the thresholds surrounding it. In connection with our example, this means shifting the second and third threshold farther away from each other. Obviously, this is achieved best if β j is positioned within the middle category. However, if β j lies outside of the middle category, a different result is obtained. Specifically, for large values of ω, the probability of selecting the disagree category approaches 1.0 for the present example if the sum constraint is used. Clearly, this appears to be an undesirable model property.
3.2. Category Probability Logit
The type of category probability logit used in connection with scaled thresholds needs to be considered carefully. Two commonly used types of logits when modeling ordered categorical data are the adjacent category logit and the cumulative logit (e.g., Agresti, 2003).
The adjacent category logit considers the probabilities of two consecutive categories k and
It is commonly used in connection with Rasch rating models such as the rating scale model (Andrich, 1978), the partial credit model (Masters, 1982), or the generalized partial credit model (Muraki, 1992).
The cumulative logit considers the probability of selecting any category higher than the kth category versus the remaining categories, that is
A well-known model using this logit is the GRM (Samejima, 1969).
In latent variable models, thresholds characterize the locations where the category probability logits are zero on the latent trait continuum, that is, where the logit’s numerator and denominator probabilities are identical. Thus, for models with adjacent category logits, thresholds mark the transition between adjacent categories, which equals the trait for which either of two consecutive categories is selected with the same probability. For models with cumulative logits, thresholds separate regions of a continuum corresponding to the ordered response categories, that is, the point where either of both regions is selected with probabilities of .5 (e.g., Masters, 1982). It follows that if thresholds based on either logit differ, then scaling factors that modify them will operate differently as well.
While our model uses cumulative logits, the model of Jin and Wang (2014) uses adjacent category logits. Their model can be expressed in terms of the logit of the probabilities of two consecutive response categories (see their equation 10), that is
where θ
i
is the trait, μ
j
is an item difficulty parameter, and
Before we compare both logit types for modeling ERS with scaled thresholds, note that, in general, adjacent category logits are by no means inferior to cumulative logits. Indeed, in some cases, they are more flexible because they are not restricted to modeling ordered response categories but can be applied to nominal data also. However, in the present context, their flexibility introduces unnecessary complexity. Indeed, depending on the specific model parameters, scaling factors modeled in connection with adjacent category logits do not yield a meaningful interpretation in terms of measuring ERS. This limitation can be explained best by outlining two problematic properties.
First, consider the ordering of thresholds. As is well known, when using cumulative logits, thresholds have to be ordered by definition, that is,
For scaling factor (ω) to unanimously account for ERS across items, thresholds of all items are required to be ordered. Only then will decreasing ω values yield increased probabilities for selecting the outer response categories for all items and vice versa. However, if the order of thresholds differs between items, their scaling with ω yields different results.
To illustrate, consider 2 three-categorical items. Item parameters of the first item are

Probability of an extreme response of the model of adjacent category logits for two exemplary items and three different ω values:
Although this is only one of many conceivable patterns for disordered thresholds, it clearly shows that effect and interpretation of scaling factor ω depend on the thresholds’ order. Obviously, this is an undesirable model property because ω is intended to measure ERS for all items similarly. Models using cumulative logits, on the other hand, do not have this limitation because their thresholds are ordered and, therefore, they yield a unanimous interpretation of the scaling factor.
To illustrate the second problematic property of using adjacent category logits, Figure 4 shows the CCCs of responders with three different scaling factors for a three-categorical item. To emphasize the distinction between middle and outer response categories, the curve of the middle category is printed in bold, whereas dashed lines are used for the curves of the outer response categories. Item parameters are

Category characteristic curves of scaled threshold models based on adjacent category and cumulative logits for different ω values. Item parameters are
Scaled thresholds (
The first row of Figure 4 gives the CCCs of a responder with a tendency for preferring the middle category (
As the graphs on the left in Figure 4 pertaining to adjacent category logits show, the probability of an extreme response does not approach 1.0 for all trait levels as the scaling factor approaches zero. Specifically, if θ is close to β, there is still a considerable probability for selecting the middle category, even if
Clearly, using adjacent category logits in connection with the outlined assumptions fails to fit response patterns of most persistent extreme responders, that is, responders who exclusively use the outer response categories. Not accounting for such response patterns is problematic because they are frequently observed in empirical data (e.g., Bachman & O’Malley, 1984; Baumgartner & Steenkamp, 2001). Using the STM, this is not an issue because as it’s scaling factor ω approaches zero, the probability of selecting the middle category approaches zero, too.
4. Other Scaling Factor Models
To illustrate the unique features of the STM, we compare it to (a) the GRM for measuring person reliability (PRM) of Ferrando (2009) and (b) the graded response differential discrimination model (GRDDM) of Lubbe and Schuster (2017). Similar to the STM, they are both modified versions of the GRM using an additional scaling factor.
The PRM was proposed to account for individual differences in reliability of responses. This is achieved by including a responder-specific scaling factor in the form of a person discrimination parameter. Although this model was not developed as an ERS model, its parameters also account for differential preferences for extreme categories. Specifically, the PRM assumes the following category probability function:
where, in contrast to the GRM in Equation 1, the item-specific discrimination parameter aj has been replaced with a responder-specific discrimination ai modeling the inverse of the residual standard deviations for each responder. Similar models also including a random residual scale have been proposed by Cleveland, Denby, and Liu (2000) and Hedeker, Mermelstein, Hakan, and Berbaum (2016).
The GRDDM is more closely related to the STM. However, instead of shifting thresholds, it shifts the expected response of an individual toward or away from the scale midpoint using responder-specific discrimination parameter
Similarly to the STM, β
j
is defined by the median constraint as an item-specific midpoint parameter, and
The STM cannot be equated with the other two models. Nevertheless, the respective approaches bear obvious similarities. For both PRM and GRDDM, the probability of an extreme response changes depending on the value of the scaling factor. However, the change in extreme response probability is directly linked to discrimination parameters, which, if modeled responder-specific, are also indicators of the reliability of an individual’s response pattern. Similarly to a high item discrimination, a high person discrimination implies that item response and trait correspond more closely to each other than for an individual with a low person discrimination.
The ai parameter of the PRM may be considered as a responder-specific scale parameter, whereas the
The STM, on the other hand, does not assume any relation between the reliability of a response pattern and the probability of an extreme response. Its scaling factor only affects the thresholds and, therefore, does not alter any parameter associated with response reliability. In this way, the STM exclusively models response spacing by expanding or contracting intervals between response categories.
5. Simulation Study
Consistent parameter estimates for polyomous item response models with scaled thresholds can be obtained by using either Markov chain Monte Carlo methods or a marginal maximum likelihood approach (e.g., Lubbe & Schuster, 2017; Rossi et al., 2001). However, consistent estimation does not necessarily preclude biased parameter estimates in small samples. Thus, in the following, we will investigate the parameter recovery of the STM in small samples using Monte Carlo simulation based on marginal maximum likelihood estimation.
We generated data sets with 8 five-categorical items for samples sizes of
For each of the six data constellations, resulting from the combination of sample size and scaling factor variance, we generated 100 samples. Model fitting was performed using SAS PROC NLMIXED with adaptive Gauss-Hermite quadrature using five quadrature points and quasi-Newton optimization. Exemplary SAS PROC NLMIXED code is provided in the Appendix section. Using the default settings for convergence of the NLMIXED procedure, all models reached convergence.
Table 1 gives the average parameter estimates of the STM for the six simulations together with the true parameter values. Clearly, the deviation of the average estimates from the true parameter values is extremely small. Only in case of a sample size of 150, there is a small tendency for β
j
and
Average STM Estimates Under the Four Simulation Study Conditions
Note. Estimates of
In addition, the reliability of the estimates’ standard errors is of practical relevance. It can be investigated by comparing standard error estimates to the empirical distribution of estimates across data samples. Specifically, if the averaged standard errors and the parameter estimates’ standard deviations across replications are close, this can be considered as evidence for the reliability of the former.
As Table 2 shows, averaged standard errors and the estimates’ standard deviations are almost identical. Only standard errors of
Average Standard Errors and Standard Deviation of Parameter Estimates Under the Four Simulation Study Conditions
Note. The first value in each cell is the average standard error and the second value is the parameter estimates’ standard deviation. Standard errors pertaining to
Overall, the STM (if estimated using marginal maximum likelihood) performs very well in small samples. Already for samples of 250 observations, parameters are virtually unbiased. However, as for any simulation study, it has to be noted that generalizations of results have to be considered with caution.
6. Empirical Examples
In the following, results of two data examples are given. The first example serves to illustrate how to interpret model results using a data set with a small number of variables. In the second example, the impact of altering the type of logit is demonstrated empirically using a representative sample of questionnaire responses.
6.1. Example 1
In the first example, we analyze a sample of responses of 582 undergraduate students to 5 six-categorical items measuring the attitude toward aphorisms. We fitted STM and GRM to the data, using SAS PROC NLMIXED with marginal maximum likelihood estimation with adaptive Gauss-Hermite quadrature and quasi-Newton optimization. The number of quadrature points was fixed to 5 for all analyses.
Table 3 gives the estimates of item discriminations (aj), midpoint parameters (β
j
), and scaling factor variance (
STM and GRM Parameter Estimates for the Aphorism Data
Note. Standard errors are in parentheses. STM = scaled threshold model; GRM = graded response model.
Next, we consider latent variable predictions of STM and GRM. For this purpose, we calculated
Latent Variable Predictions of ω and θ Together With the Item Responses for Five Exemplary Responders
Note. Standard errors are in parentheses.
The
Finally, we consider the relation between

Relation between latent variable predictions of ω and number of extreme responses for the aphorism data.
6.2. Example 2
To investigate the impact of the choice of category probability logit when modeling scaled thresholds, we analyzed questionnaire responses of the normative sample of the German version of the NEO PI-R (Ostendorf & Angleitner, 2004). Specifically, we analyzed the six facets pertaining to neuroticism: anxiety (N1), angry hostility (N2), depression (N3), self-consciousness (N4), impulsiveness (N5), and vulnerability (N6). Each facet consists of 8 five-categorical items, with response categories strongly disagree, disagree, neutral, agree and strongly agree. The sample comprises N = 12,003 cases. Negative items were recoded prior to analysis. Analyses were performed with identical specifications as in the previous example.
First, we consider the overall model fit. Table 5 gives the AIC for all six neuroticism facets for the STM with cumulative logits and adjacent category logits, respectively. Clearly, AIC values of the STM with cumulative logits are distinctly lower than for the model with adjacent category logits.
AIC Values of All Neuroticism Facts for Cumulative and Adjacent Category Logit Models
Note. AIC = Akaike’s information criterion.
Next, we consider the impact of disordered thresholds on the scaling factor more closely. Note that all of the neuroticims facets had at least 1 item with disordered thresholds when analyzed with the STM with adjacent category logits. Specifically, the number of items with disordered thresholds were 2, 6, 3, 8, 1, and 2 for facets N1 to N6, respectively. To keep our presentation concise, we selected the neuroticism facet with the largest number of items with disordered thresholds, that is, N4, and the one with the smallest number, that is, N5, for further analyses.
For all items with disordered thresholds, the order of the second and third threshold was inverted, that is,
We calculated

Relation between
It is obvious that for N5, which has only 1 item with disordered thresholds, the relation between scaling factors of both models is closer than for facet N4, where all 8 items have disordered thresholds. Specifically, the correlation of scaling factors for N5 is .905 and for N4 it is .781.
The heteroscedastic shape of the left scatterplot in Figure 6, pertaining to N4, can also be explained by inverted thresholds. Clearly, the heterogeneity between scaling factors of both models becomes larger as ω values increase, that is, for responders who tend toward the Middle Categories 2, 3, and 4. This specific pattern results because thresholds between these middle categories are inverted for all items of N4. Indeed, for responders who avoided Extreme Categories 1 and 5 entirely (gray dots), the relation between
Finally, we consider the relation between scaling factor values of both facets for the STM with cumulative and adjacent category logit, respectively. Because ERS is considered a stable individual characteristic, scaling factors obtained from different item responses should be significantly correlated. Thus, correlating scaling factors of N4 and N5 within each model may be considered as part of construct validation.
Figure 7 gives the scatterplots of the relation between

Relation between log (ω) values of facets N4 and N5 models based on adjacent category and cumulative logits, respectively.
In summary, the empirical results confirm the limitations of using adjacent category logits in connection with modeling a responder-specific scaling factor. Specifically, the issue of inverted thresholds is clearly illustrated by the inversion of scaling factor values for responders who preferred nonextreme categories for facet N4. The STM, on the other hand, does not reveal any inconsistencies. Model fit and the validity of scaling factor values are superior across all neuroticism facets.
7. Discussion
The present article investigates an approach for modeling ERS in polytomous item responses using scaled thresholds. Specifically, a model is proposed that modifies the distances between thresholds for each responder by means of a latent scaling factor, allowing for the simultaneous and independent measurement of trait and ERS.
While the use of a scaling factor for modeling idiosyncratic response behavior is not new and has proven useful to make a clear conceptual distinction between response behavior and trait, the model we propose has novel features. First, compared to other related models, the STM yields a scaling factor that exclusively measures responder-specific differences in the extremity of the response pattern independently from other responder-specific attributes such as the reliability of responses.
Furthermore, we illustrate that the properties of modeling scaled thresholds are sensitive to choices of model parametrization, such as the type of category probability logit and the constraint imposed on the thresholds deviation parameters for model identification. As has been demonstrated, the proposed model, using cumulative logits and specifying a median constraint on the thresholds, has several desirable properties not shared by other scaled threshold approaches.
Using simulations studies, it has been shown that essentially unbiased parameter estimates and accurate standard errors can be obtained even in small samples. While the investigated settings on which our simulation studies are based may represent realistic cases of empirical applications, their scope is nevertheless restricted due to the computationally demanding marginal maximum likelihood estimation. Despite this limitation, we consider the investigated settings sufficient to demonstrate that the STM may also be considered for analyzing small samples. For future research, on the other hand, it would be desirable to conduct more extensive simulation studies.
Similarly to other scaling factor models, the STM can be extended in different ways. For instance, models with more than one trait can be considered. Moreover, within the nonlinear mixed model framework, it is viable to include covariates. As proposed by Hedeker et al. (2016), one can use this feature to (a) examine whether covariates are related to a specific response style or (b) to investigate differential item functioning by including interactions between covariates and item parameters to the model.
Finally, it has to be emphasized that the disadvantage of using adjacent category logits is specific to the STM context. There are other approaches for modeling ERS that are not negatively affected by the choice of logit, for example, the approach by Tutz and Berger (2016).
Footnotes
Appendix
In the following exemplary, SAS PROC NLMIXED code is provided for analyzing 5 items with five or six response categories, respectively. We include these two examples because the codes for items with odd or even numbers of response categories require different threshold constraints.
The data need to be provided in the so-called long format. Variable y contains the categorical item responses, variable sbjct contains a responder-specific ID, and variable j contains the item number. Further instructions are given in the program code.
As starting values for the model parameters, one might consider using GRM estimates and setting the staring value for ‘vnu’ between 0.01 and 0.15.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
