Abstract
When constructing measurement scales, regular and reversed items are often used (e.g., “I am satisfied with my job”/“I am not satisfied with my job”). Some methodologists recommend excluding reversed items because they are more difficult to understand and therefore engender a second, artificial factor distinct from the regular-item factor. The current study compares two explanations for why a construct’s dimensionality may become distorted: response difficulty and item extremity. Two types of reversed items were created: negation items (“The conditions of my life are not good”) and polar opposites (“The conditions of my life are bad”), with the former type having higher response difficulty. When extreme wording was used (e.g., “excellent/terrible” instead of “good/bad”), negation items did not load on a factor distinct from regular items, but polar opposites did. Results thus support item extremity over response difficulty as an explanation for dimensionality distortion. Given that scale developers seldom check for extremity, it is unsurprising that regular and polar opposite items often load on distinct factors.
Inclusion of reversed items used to be a popular suggestion for scale construction. According to prominent methodologists, reversed items are useful to counter acquiescence response style (Nunnally, 1978), respondents’ tendency to agree with survey items regardless of item content (Benter et al., 1971). In addition, reversed items serve as stop signs that demand participants’ attention in a scale that would otherwise contain monotonous content, which in itself may encourage careless responding (Schmitt & Stults, 1986).
However, the popularity of reversed items is in decline, with recent instruments excluding them (e.g., Ferris et al., 2008; Jonason & Webster, 2010; Lau et al., 2006; Owens et al., 2013; see Weijters & Baumgartner, 2012, for such a statistic in marketing research). Reversed items often have lower internal consistency than regular items. They also contaminate a construct’s dimensionality by loading on a distinct factor—different from regular items—even in a theoretically unidimensional instrument (e.g., each Big Five personality dimension; Biderman et al., 2011; Kam & Meyer, 2015). The incorporation of reversed items may thus weaken the apparent psychometric strength of a scale.
Two major explanations have been advanced to explain the nuisance factor caused by reversed items. The most common explanation is response difficulty: Reversed items are relatively difficult to understand, which in turn may lead to “misresponses” (Gnambs & Schroeders, 2020). If reversed items are more difficult to answer, some respondents may be confused by them when they are mixed in with regular-keyed items. Misresponses are thought to stem from respondents’ inability to answer certain items. According to this perspective, a possible solution is to exclude reversed items. Many researchers have recommended excluding reversed items so that a scale will have a cleaner factorial structure.
A second, more recent explanation is that although respondents are generally logical and capable of answering items, characteristics such as item extremity may induce the apparent two-dimensional structure (Kam & Meyer, 2022; Kam et al., 2021). For instance, when most participants are at the mid-range of a trait (such as human height), extreme items will be disagreed with by most respondents. Therefore, the relationship between regular and reversed items will not be simple and linear. This reasoning is called logical response perspective. An implication is that scale developers should be cognizant of the logical response pattern by testing item characteristics such as item extremity as a possible source of bidimensionality.
The current paper empirically compares the two perspectives (response difficulty vs item extremity) to determine which better accounts for the bidimensionality result. We begin by elaborating on the perspectives and then present empirical studies.
Incapability Perspective: Response Difficulty
Numerous researchers have suggested that reversed items are more difficult to respond to than regular items. Several studies have found that reversed items have lower internal consistency than regular-keyed items. For instance, Benson and Hocevar (1985) showed that reversed items using the negation word “not” have lower reliability than their regular counterparts. In factor analysis, a two-factor model fit the data better than a one-factor model. As the participants in Benson and Hocevar’s study were children, Grades 4 to 6, misresponses were probably the result of comprehension deficiency among the young respondents.
Kam and Fan (2020) developed a strategy for testing construct dimensionality when a subgroup of participants has problems responding to reversed items. Based on the multitrait-multimethod approach originated by Litson et al. (2017), Kam and Fan conducted factor mixture modeling (FMM) to examine heterogeneity in item response. The FMM technique attempts to separate classes of individuals that differ in parameter estimates such as factor loadings. Using job satisfaction as an example, the researchers found a class of respondents (about 20% of participants) who had exceptionally low factor loadings and exceedingly high standard errors on reversed items. Subsequent analysis revealed that this subgroup had more trouble providing consistent responses to reversed items than to regular items. Exclusion of this response-inconsistent class raised the correlation between regular and reversed items from .76 to .87. Kam and Fan thus showed that construct dimensionality can be distorted even when only a minority of respondents find reversed items challenging. Kam (2020) replicated those findings with a dispositional optimism scale, with the correlation between regular and reversed items leaping from .73 to .94 after excluding participants who had trouble with reversed items. As their purpose was to showcase the potential of the FMM research methodology, those studies did not directly examine the causes of response inconsistency.
In a study that examined the impact of verbal deficiency among German students, Gnambs and Schroeders (2020) found that the discrepancies between regular and reversed items in Rosenberg’s Self-Esteem Scale (RSES) decrease as the verbal and reasoning abilities of students increases. This led Gnambs and Schroeders to conclude that “the factor structure of the RSES seems to represent a response style artifact associated with cognitive abilities” (p. 404), and that “respondents that are unable to properly understand and evaluate the content of an item might more frequently resort to acquiescent responding instead of processing the item and elaborating a response and, thus, introduce multidimensionality in an otherwise unidimensional scale” (p. 412). Gnambs and Schroeders’s finding thus echoes the sentiment, common in the literature, that respondents’ errors were a major cause of misresponse.
However, the studies cited above do not differentiate the types of reversed items that may differ in comprehension difficulty. One type attaches a negation word (“not” or “no”) to an otherwise regular item (e.g., “My job is not excellent” or “My job is no good”). A second type reverses the item without, however, using a negation word (“My job is terrible”). To develop reversed items, scale developers usually incorporate both types, implicitly assuming that they have similar properties. For instance, the Affective Occupational Commitment Scale (Meyer et al., 1993) includes both types of reversal (e.g., “I do not identify with the nursing profession”; “I regret having entered the nursing profession”). Brayfield and Rothe’s (1951) job satisfaction index also includes both types (e.g., “. . . my job is no more interesting than others. . .”; “. . .disappointed that I ever took this job”). The statistics were presented by Weijters and Baumgartner (2012) in marketing research: 30% of the scales contain reversed items, and 20% of all items, regular or reversed, contained negation among all items (also reported in Baumgartner et al., 2018). Negated items are thus quite prevalent in survey scales.
As one of the few researchers who analyzed types of reversed items, Baumgartner et al. (2018) analyzed participants’ misresponse data with eye-tracking data. The researchers found that polar opposite items—reversed items that did not use negation words—received more cognitive attention from participants than regular-keyed items, but they also had lower factor loadings as well as larger score differences with regular items. Baumgartner et al. speculated that the large score difference was due to greater response difficulty. Negated items, however, did not differ substantially from regular items in factor loading.
Despite those results, when providing research recommendations, Baumgartner et al. (2018) continued to discourage the use of negated regular items: “Negation can sometimes confuse respondents, especially under peripheral processing” (p. 13). Based on their overall finding that respondents devote closer attention to reversed than regular items, the researchers concluded that “lack of comprehension is an important contributor to misresponse and possibly is a more important determinant than lack of attention.” In sum, Baumgartner et al. (2018) insisted that negation causes comprehension difficulty despite finding that respondents give similar responses to regular and negated items.
Some research suggests that the type of reversed item influences reliability. Schrieshem and Eisenbach (1995) compared the reliabilities of four types of item concerned with leadership (specifically, initiating structure). Two of the four types were negated items. Regular items had the highest reliability (α = .89), followed by negated (regular) items (i.e., regular items with a negation word; α = .84) and polar opposites (α = .82). The worst were negated polar opposites (α = .70). Thus, the lower reliability of the reversed items—the two negated types and the polar opposites—was apparent even among undergraduates in Schrieshem and Eisenbach’s (1995) study. Comprehension difficulty does not exist only in young children as previously mentioned (Benson & Hocevar, 1985) but also in young, educated adults (Schrieshem & Eisenbach, 1995).
Why Do Reversed Items Cause Response Difficulty?
Item response involves four stages (Weijters & Baumgartner, 2012). At the comprehension stage, respondents pay attention to an item and interpret its meaning. At the retrieval stage, respondents search for memory instances that are consistent or inconsistent with the item’s content. Consistent memory retrieval is probably more likely than inconsistent memory retrieval because the former is less effortful, but this positive-test strategy will lead to a biased response (Kunda et al., 1993). At the judgment stage, respondents compare their own beliefs with item content and develop a general evaluation. At the response stage, the respondents translate their general evaluation (e.g., moderately affirmative) to one of the response options (“agree” or “strongly agree”).
Based on the sentence comprehension literature (Weijters & Baumgartner, 2012; Weijters et al., 2013), negation items pose challenges in the comprehension and judgment stages. First, negation items are probably harder to comprehend because respondents need to mentally revert the meaning of the item. For instance, given “My job is not good,” a respondent may translate “not good” into “bad” or “awful.” Such additional processing taxes respondents’ cognitive resources at subsequent stages when they need to retain the reinterpreted items in the memory. During the judgment stage, respondents may need to disagree that their job is not good. Disagreeing with a negation may impose a greater cognitive load, which may lead to a misresponse. Such cognitive work may be especially formidable for respondents with lower verbal ability, reasoning skills, or working memory.
If reversed items pose a response challenge for participants, a natural recommendation would be to stop using them. Indeed, as noted earlier, many researchers have called for excluding reversed items from any scale. Schriesheim et al. (1991) stated that “one cannot help but wonder why item reversals are still advocated as desirable by psychometricians” (p. 77). In a later paper, their position did not change: “The time-worn general recommendation that questionnaire instruments should contain a mixture of regular and reversed items may warrant abandonment” because reversed items “increase error variance and create a distorted map of construct dimensionalities” (Schriesheim & Eisenbach, 1995, p. 1190). Suárez-Álvarez et al. (2018) echoed the sentiment. When reversed items are added, they said, “the measurement precision of the instrument is flawed”; “examinees’ scores differ significantly from those obtained in tests where all of the items are of a similar form” (p. 156).
Logical Response Perspective: Item Extremity
Although response difficulty is currently the dominant explanation for item misresponse (e.g., Baumgartner et al., 2018), Kam and colleagues (Kam et al., 2021; Kam & Meyer, 2022) have challenged this view. They propose that participants’ apparent misresponses can be the result of logical responding behavior. While respondents at the two ends of a given trait may readily agree or disagree with regular and reversed items, respondents in the middle may disagree with both because the item content does not match their standing on the trait. The researchers demonstrated the phenomenon by using a construct that is undeniably unidimensional—human height.
Kam and Meyer (2022) reasoned that responses to regular and polar opposite items can be inconsistent due to average respondents. An average height respondent will disagree with both the regular item “I am tall”and its polar opposite “I am short” (see Table 1 for a full explanation). When polar opposite items are reversed, the score of regular items (“agree”) will differ from the score of the polar opposites, causing regular and polar opposite items to load on separate factors in factor analysis. Thus, in the case of human height, the empirical dimensionality of a construct as shown in factor analysis clearly does not match its theoretical dimensionality.
Responses for Tall Items and Short Items
Note. Response inconsistency after score reflection for polar opposite and negated regular items has been given in bold. Regular and polar opposite items. Tall respondents will agree with a regular and disagree with a polar opposite. Short respondents will disagree with a regular and agree with a polar opposite. Average height respondents will disagree with both the regular and the polar opposite. In this example, the response after reversal for reversed items (i.e., polar opposite and negated regular) will differ between regular and polar opposite items and between negated regular and negated polar opposite items. Negated regular items. For a negated regular item, tall respondents will disagree, average height respondents will agree, and short respondents will agree. After reflection, the answer for the negated regular item will be “agree,”“disagree,” and “disagree,” respectively. These answers are consistent with the response to a regular item (“agree,”“disagree,” and “disagree”). Thus, regular and negated regular items should load on the same factor. Negated polar opposite items. For such items (e.g., “I am not short”), tall respondents will agree, average height respondents will agree, and short respondents will disagree. The answer is consistent with the answer for the reflected polar opposite item (“agree,”“agree,” and “disagree”). Therefore, polar opposites and negated polar opposites would load on the same factor, whereas negated regular items and negated polar opposite items would load on separate factors. R = reversed items.
How does the logical response perspective explain negated items? Contrary to nonnegated items, responses to negated items and its parent items (regular or polar opposite) can be parallel. Let us first discuss negated regular items. An average height respondent will agree with the negated regular item, “I am not tall.” After the answer for negated regular items is reflected, their score (“disagree”) will be consistent with their score on the regular item, “I am tall” (“disagree”), causing both item types to load on the same factor. The same happens for negated polar opposites. An average height respondent will agree with the negated polar opposite item (“I am not short”). The answer for the negated polar opposite (unlike the negated regular) is not reflected. Therefore, the score of negated polar opposite (“agree”) will be consistent with the score of reflected polar opposite “I am short” (“agree”), causing both types to load on the same factor. The logical response perspective integrates the early but often neglected research area of item extremity (McPherson & Mohr, 2005; Spector et al., 1997) with the properties of the negated item, offering a unique view of how the two types of reversed items—polar opposite and negated—differ.
So far this speculation regarding the difference between polar opposite items and negated items has received empirical support only for the construct of human height. Height was chosen as the initial test because it is undeniably unidimensional. Thus, any bidimensionality results can be caused only by a methodological artifact. Kam et al. (2021) showed that regular and polar opposite items load on distinct factors, and that negated items and their parents (i.e., regular with negated regular; polar opposite with negated polar opposite) load on the same factor. The largest response difference between regular and polar opposite items was found among participants of average height; these findings were replicated by Kam and Meyer (2022).
It may be wondered whether the same reasoning applies to constructs other than human height. Kam and Meyer (2022) found that job satisfaction items (regular) and job dissatisfaction items (polar opposite) load on separate factors when the items have extreme wording (e.g., “excellent job” vs. “terrible job”). Thus, these results appear to challenge the claim that the bidimensionality results are simply the outcome of response difficulty (Gnambs & Schroeders, 2020) because the items were all simple to understand in Kam and Meyer (2022). Those researchers, however, have not tested their hypotheses regarding negated items in any psychological scales.
How Does a Logical Responding Pattern Affect Psychological Measures?
To account for psychological measures such as job satisfaction, only a slight modification from the previous height example is needed (see Table 2). When items are extreme (e.g., “My job is excellent” or “My job is terrible”), they will be disagreed with by most respondents, including of course individuals with average satisfaction. After reflecting the score of polar opposite items (“disagree”→“agree”), regular (“disagree”) and polar opposite items (“agree”) will have discrepant responses, causing them to load on separate factors. Similarly, when negated items are extreme (e.g., “My job is not excellent” or “My job is not terrible”), they will both be endorsed by average satisfaction individuals. When the negated regular items are reflected (“agree”→“disagree”), their responses will differ from that of negated polar opposites (“agree”), again causing a two-factor solution. Nevertheless, regular items (“disagree”) and reflected negated regular items (“disagree”) will have the same response and thus load on the same factor. Similarly, polar opposites (“agree”) and reflected negated polar opposites (“agree”) will have the same response and load on the same factor.
Responses for job Satisfaction Items and Job Dissatisfaction Items
Note. Extreme regular and polar opposite items. An extreme regular item will be agreed with by satisfied respondents but disagreed with by average satisfied and dissatisfied respondents. An extreme polar opposite item will be disagreed with by satisfied and average satisfied respondents, but agreed with by dissatisfied respondents. After score reversal for the polar opposite, the answer for averaged satisfied respondents would not match (given in bold in the table for reference). Extreme negated items. Similarly, a negated regular item will be disagreed with by satisfied respondents but agreed with by average satisfied and dissatisfied respondents. A negated polar opposite item will be agreed with by satisfied and average satisfied respondents, but disagreed with by dissatisfied respondents. After score reversal for the negated regular items, the answer for averaged satisfied respondents would not match (given in bold for reference). The response inconsistency mainly comes from average satisfied respondents. Moderate items. Average respondents will agree with a regular and negated polar opposite item, but disagree with a polar opposite item and a negated regular item. After score reversal for the polar opposite and the negated regular, all items will match in the answer (i.e., “Agree”), resulting in no response inconsistency. R = reversed items.
Does bidimensionality occur for scales with moderate wording? Probably not. Meta-analysis shows that job satisfaction average is well above the midpoint of a scale (Wilkin, 2013), meaning that average individuals are quite happy with their work. When wording is not extreme, an average satisfaction respondent will agree with a regular item (“My job is good”) and disagree with a polar opposite (“My job is bad”), rather than disagreeing with both types (as in the case of height and extreme job satisfaction), and reflected polar opposites (“disagree”→“agree”) will have the same response as regular items (“agree”). Similarly, an average respondent will disagree with negated regular items (“My job is not good”) and agree with negated polar opposites (“My job is not bad”). Thus, reflected negated regular items (“disagree”→“agree”) and negated polar opposites (“agree”) will have the same response as regular items and polar opposites (both “agree”). Regular, polar opposite, and their negated counterparts will all load on the same factor. The same reasoning and factor loading pattern can be extended to other constructs such as life satisfaction.
The Current Research
To our knowledge, no research directly compares the two major explanations of the bidimensionality phenomenon: response difficulty versus item extremity. True, other potential explanations exist, namely, acquiescence response style, careless responding, and content noncorrespondence between regular and reversed items (see Kam et al., 2021, for a brief review). However, acquiescence is unlikely, as it accounts for only a small percentage of the variance. 1 On the contrary, the impact of careless responding varies wildly because it is sample-specific (from about 5% in Weijters et al., 2013, to 40% or more in Oppenheimer et al., 2009). Thus, the influence of careless responding from one study probably does not generalize to another. Often, it is challenging to revert a regular item’s content because it has an unclear or nonequivalent opposite. As pointed out by a reviewer, the polar opposite of the item “I am clever”—namely, “I am unintelligent” or “I am stupid”—may elicit a reactive response from participants due to the item’s extremely low desirability (Kam, 2018), which may in turn distort the scale’s dimensionality and result in content noncorrespondence between regular and reversed items, but in this preliminary investigation we only employ items that are easy to reverse their meaning. For these reasons, we will focus on the two major perspectives—response difficulty and item extremity—and assess their ability to explain the bidimensionality phenomenon.
If response difficulty is more important, negation should play a large role: Negated regular items should load on a factor distinct from regular items, and negated polar opposites should load on a factor distinct from polar opposites. This follows from the assumption that negation causes trouble for comprehension. However, provided items are written clearly (e.g., “My job is excellent,”“My job is terrible”), polar opposites should not cause more difficulty than regular items.
If, on the other hand, item extremity is more important, then for extreme items, regular and polar opposites should load on separate factors, and negated items should load on or follow their parents (i.e., negated regular items with regular items; negated polar opposites with polar opposites). For moderate items, all four item types should load on the same factor.
Two samples will be employed to examine the research question. The first involves the reanalysis of previously published job satisfaction data. Research on job satisfaction–dissatisfaction dimensionality is timely. At least one study has shown that satisfaction and dissatisfaction load on separate factors (Credé et al., 2009), prompting theorists to call for reconceptualization of the construct’s dimensionality (Judge et al., 2017). In our data set, the results for regular and polar opposite items were previously shown to load on separate factors when items are extreme (Kam & Meyer, 2022), but the data for negated items were never published. Suppose negated items load on a factor independent from their parents regardless of item difficulty. This result would support the response difficulty perspective because negation complicates item comprehension and response judgment. On the contrary, if negated items and their parents load on the same factor, the result would support the item extremity perspective.
The second sample involves life satisfaction data that were not previously published. Although life satisfaction is conceptualized as a cognitive, global evaluation of one’s life quality (Pavot & Diener, 1993), research has found that life satisfaction and dissatisfaction have distinct nomological networks, with income having a stronger impact on dissatisfaction than satisfaction (Boes et al., 2010a, 2010b) and with dissatisfaction providing a stronger driving force of behaviors than satisfaction (Kageyama & Sato, 2021). An examination of its dimensionality is thus justified.
The construct of satisfaction was used for both studies because it is relatively easy to write items that provide a fair test of the response difficulty versus item extremity perspectives.
Method
Sample 1: Job Satisfaction
Part of the data has been published in Kam and Meyer (2022). In their analysis, Kam and Meyer included only regular-keyed and polar opposite items for job satisfaction. The current data set contains both negated regular items and negated polar opposite items.
Participants
Participants were 497 adults aged 18 to 60 years who reported themselves working full-time in the United States. They completed an online survey through Amazon Mechanical Turk (MTurk) in exchange for a small remuneration (US$1). All of the MTurk participants were previously qualified as high-quality respondents by CloudResearch. They reported themselves native speakers of English who have completed at least a high school education. Only participants who successfully passed three interspersed attention check items (e.g., “I occasionally eat cement”) were included in the analysis, leaving the final sample of 475 (214 females, 260 males, and one other gender; Mage = 39.99, SD = 9.53). In terms of ethnic background, 67.37% were Caucasian, 9.26% African, 7.58% Asian, 4.84% Hispanic, and 10.95% other ethnicities or unidentified.
Measure
The measure was constructed from two popular measures: the Job in General Scale (Smith et al., 1985) and the Illinois Job Satisfaction Index (Credé et al., 2009). The resulting measure has 16 items with moderate wording and 16 items with extreme wording. For each wording type, four are regular items, four are polar opposite items, four are negated regular items, and four are negated polar opposite items. Item order was randomized by the survey software Qualtrics; thus, each participant received a different order. The items were answered on a 5-point Likert-type scale (1 = strongly disagree; 5 = strongly agree).
Sample 2: Life Satisfaction
Participants
Participants were 508 adults aged 18 to 60 years who completed an online survey through Amazon Mechanical Turk in exchange for a small remuneration (US$1). They reported themselves native speakers of English who have completed at least a high school education. All were previously qualified as high-quality respondents by CloudResearch. Only participants who successfully passed four interspersed attention check items (e.g., “I have never used a computer”) were included, leaving the final sample of 466 (220 males, 244 females, one other gender, and one unidentified; Mage = 38.66, SDage = 19.96). In terms of ethnic background, 74.89% were Caucasian, 7.51% African, 4.29% Asian, 4.94% Hispanic, and 8.37% other ethnicities or unidentified.
Measures
The measures were adapted from two existing life satisfaction measures: the Satisfaction with Life Scale and the Riverside Life Satisfaction Scale. For the extreme version, there were a total of 12 items, of which regular, polar opposite, negated regular, and negated polar opposite had three items each. For the moderate version, there were a total of 16 items, of which regular, polar opposite, negated regular, and negated polar opposite had four items each. Item order was again randomized by survey software. While all participants completed the regular and polar opposite items, they were randomized to complete the two negated item types (i.e., negated regular and negated polar opposite) of either the extreme version (n = 252) or the moderate version (n = 214). All items were answered on a 5-point Likert-type scale (1 = strongly disagree; 5 = strongly agree). The current study analyzes the two groups separately.
Analysis Strategies
To examine construct correlation, all four types of items will be subjected to parallel analysis using principal components (PA-PCA). Simulation research suggests using principal components rather than principal axis factors for parallel analysis because the former has been found better at recovering the true number of factors in simulation investigations (Crawford et al., 2010; Lim & Jahng, 2019). In addition to parallel analysis, we employed exploratory structural equation modeling (ESEM) with robust maximum likelihood estimator (MLR) to compare the one-factor and two-factor solutions, while modeling and controlling for acquiescence response style by setting a random intercept latent factor that loaded at 1’s on all items (Maydue-Olivares & Coffman, 2005). Geomin rotation, which has the advantage of allowing the data to reveal an optimal factor loading pattern, is employed for ESEM. ESEM was conducted using Mplus 8.8 (Muthén & Muthén, 1998–2007). In a good model, the magnitude of a factor loading should be sizable (e.g., λs > .40) and model fit should be good. While some methodologists do not believe in model fit index cutoffs (e.g., Hayduk et al., 2007), generally a model with a higher Tucker–Lewis index (TLI) and comparative fit index (CFI) and lower root mean square error of approximation (RMSEA), standardized root mean square residual (SRMR), Akaike information criterion (AIC), and Bayesian information criterion (BIC) should be preferred. We will then examine the correlation among the four types of items using confirmatory factor analysis (CFA) with MLR to account for non-normality. Two types of items with a strong correlation (e.g., r > .90) may belong to a single factor. Separate analyses of extreme and moderate items will be carried out.
Finally, we will compare participants’ response time among the four item types, assuming that response difficulty correlates positively with response time. In response to a reviewer’s suggestion to trim outliers, response times over three standard deviations above the mean are excluded for each item.
Actual items, data, and analytic outputs can be found at https://osf.io/a8p9h/.
Results
Sample 1: Job Satisfaction
Extreme Items
Parallel analysis suggested a clear, two-factor solution for the extreme version (Figure 1) because the eigenvalues of the first two components were higher than their corresponding resampled data eigenvalues. The follow-up ESEM revealed that the two-factor model (χ2 = 199.23, df = 88, TLI = .96, CFI = .97, RMSEA [95% confidence interval (CI)] = .05 [.04, .06], SRMR = .02, AIC = 15,689.33, BIC = 15,955.79) obviously fit the data better than the one-factor model (χ2 = 969.48, df = 103, TLI = .74, CFI = .78, RMSEA [95% CI] = .13 [.13, .14], SRMR = .11, AIC = 16,823.85, BIC = 17,027.85), with sizable factor loadings in the two-factor model (all λs > .40). The correlation between extreme regular and extreme polar opposite items was estimated to be −.65, 95% CI = [−.56, −.73], a magnitude that is far from a perfect correlation. Using CFA with a latent factor on extreme regular items and another on extreme polar opposites, 2 the correlation was statistically different from perfect negative unity (r = −.73, 95% CI = [−.68, −.79]). The correlation using CFA is higher than that using ESEM probably because CFA does not allow cross-loadings (i.e., the loading of dissatisfaction items on satisfaction latent factor is constrained to be null). The correlations between the four types of items are shown in Table 6.

Parallel analysis using principal component analysis (PA-PCA). The cross line represents data eigenvalues and the broken line represents resampled data eigenvalues: (A) extreme job satisfaction; (B) moderate job satisfaction; (C) all job satisfaction items; (D) extreme life satisfaction; (E) moderate life satisfaction; (F) regular and polar opposite life satisfaction items (no participants complete all negated regular and negated polar opposite items).
Moderate Items
Parallel analysis suggested a clear one-factor solution for the moderate version because only the first eigenvalue is higher than its corresponding resampled data eigenvalue. The follow-up ESEM revealed that the two-factor model (χ2 = 184.95, df = 88, TLI = .97, CFI = .98, RMSEA [95% CI] = .05 [.04, .06], SRMR = .02, AIC = 12,918.94, BIC = 13,185.39) fit the data better than the one-factor model (χ2 = 261.94, df = 103, TLI = .96, CFI = .97, RMSEA [95% CI] = .06 [.05, .07], SRMR = .02, AIC = 13,034.41, BIC = 13,238.41), but the loadings for the second factor in the two-factor model are all weak (most λs < .24). Thus, the factor loading pattern favored the one-factor model. The CFA model showed that the correlation between regular and polar opposite latent factors was strong (r = −.96, 95% CI = [−.94, −.98]), further corroborating that moderate regular and moderate polar opposite items form unidimensionality. The correlations between the four types of items are shown in Table 6. They are close to perfect unity (1 or −1).
Combined Results
Parallel analysis suggested a clear two-factor solution for the combined version (extreme and moderate items together) because the first two eigenvalues are higher than their corresponding resampled data eigenvalue. The two-factor solution (χ2 = 890.31, df = 432, TLI = .95, CFI = .96, RMSEA [95% CI] = .05 [.04, .05], SRMR = .02, AIC = 27,495.54, BIC = 28,028.45) fits the data better than the one-factor solution (χ2 = 2131.12, df = 463, TLI = .85, CFI = .86, RMSEA [95% CI] = .09 [.08, .00], SRMR = .06, AIC = 29,075.30, BIC = 29,479.14) in the ESEM results. In the two-factor ESEM solution, one factor is loaded mainly by extreme satisfaction items and their negated counterparts, and another factor is loaded mainly by extreme dissatisfaction and their negated counterparts. Moderate items tend to load on both factors, with a weak tendency to prefer the dissatisfaction factor. The cross-loadings of the moderate items on both factors imply that they correlate reasonably well with both factors. The correlation between the two factors is far from −1 (r = −.75, 95% CI = [−.69, −.81]). A summary of the factor loadings is reported in Table 4.
Omega Reliability and Factor Loadings
Table 6 shows the omega reliabilities of the four item types. The consistent findings across extreme and moderate items are that (a) regular items tend to have slightly higher reliabilities than polar opposite items and negated regular items; and (b) negated polar opposites tend to have the lowest reliability (ωs = .81 for extreme items and .86 for moderate items), with 95% CIs that do not overlap other types of items. The low reliabilities for the negated polar opposites are probably caused by double negation in some of those items (e.g., “not dissatisfied with my job,”“not an unpleasant job”). As shown in Table 3, those double negation items have a slightly lower factor loadings than other negated polar opposite items in the one- and two-factor solutions.
Standardized Item Loadings in Exploratory Structural Equation Models With Geomin Rotations on Extreme and Moderate Job Satisfaction Items
Note. The word “NOT” is capitalized in the survey to minimize careless responses. Items can be found on the following website: (Blind for review purpose). Loadings > .40 are in bold. Only item stems are shown. Items can be found at https://osf.io/a8p9h/.
Response Time Analysis
Table 6 shows the response time for each condition. For extreme items, the main effect of keying direction was statistically significant, F(1, 474) = 5.09, p = .024, with regular items having slightly longer response times than polar opposite items, negated or not, Ms = 3.30 vs. 3.15, Cohen’s d = 0.08. The significant main effect of negation indicates that negated items take longer to respond to than nonnegated items, F(1, 474) = 296.21, p < .001, Ms = 3.82 vs. 2.64, Cohen’s d = 0.70. The interaction between keying direction and negation was not statistically significant, F(1, 474) = 0.86, p = .355. For moderate items, the main effect of keying direction was statistically significant, F(1, 474) = 108.26, p < .001, with polar opposite items having longer response times than regular items, negated or not, Ms = 3.16 vs. 2.65, Cohen’s d = 0.34. The significant main effect of negation indicates that negated items take longer to respond to than nonnegated items, F(1, 474) = 392.02, p < .001, Ms = 3.45 vs. 2.37, Cohen’s d = 0.75. The interaction between keying direction and negation was statistically significant, F(1, 474) = 10.36, p < .001, with the response time increase for negated items higher for polar opposite items than for regular items (difference = 1.24 vs. 0.92). Thus, the consistent finding over moderate and extreme items is longer response times for negated items, but polar opposite items do not always have a longer response time than regular items (because the opposite trend was found for extreme items).
Sample 2: Life Satisfaction
Extreme Items
Similar to the results of job satisfaction, for life satisfaction parallel analysis suggested a two-factor solution for the extreme items (Figure 1). The first two eigenvalues of the principal components were higher than those from the resampled data set. The ESEM (see Table 5) showed that the two-factor model fit the data well (χ2 = 41.85, df = 42, TLI = 1.00, CFI = 1.00, RMSEA [95% CI] < .001 [<.001, .04], SRMR = .02, AIC = 6,923.51, BIC = 7,092.93) with sizable factor loadings in the two-factor model (all λs > .40), whereas the one-factor model features a nonpositive definite latent variable covariance matrix. Using confirmatory factor analysis 2 (CFA) with a latent factor on extreme regular items and another on extreme polar opposites, the correlation was statistically different from perfect negative unity (r = −.78, 95% CI = [−.72, −.85]). The correlations between the four types of items are shown in Table 6.
Moderate Items
As with job satisfaction, for life satisfaction parallel analysis suggested that the four types of moderate items form one factor. Although the two-factor ESEM solution (χ2 = 121.27, df = 88, TLI = .98, CFI = .99, RMSEA [95% CI] = .04 [.02, .06], SRMR = .02, AIC = 6818.95, BIC = 7034.37) fits the data better than the one-factor ESEM solution (χ2 = 259.30, df = 103, TLI = .93, CFI = .94, RMSEA [95% CI] = .08 [.07, .10], SRMR = .04, AIC = 6969.46, BIC = 7134.40), the factor analytic results show that most regular and polar opposite items, and their negated counterparts, tend to load on a single factor in the two-factor ESEM solution. Using CFA with a latent factor on extreme regular items and another on extreme polar opposites, the correlation was not statistically different from perfect negative unity (r = −.99, 95% CI = [−.98, −1.00]). The correlations among the four types of items are shown in Table 6.
Combined Results
Participants completed the two negated item types (i.e., negated regular items and negated polar opposites) of either the extreme version or the moderate version; no one completed both versions. Therefore, we included only regular items and polar opposite items in the current analysis. Parallel analysis suggested a one-factor solution for the combined version (extreme and moderate items together) because only the first eigenvalue is higher than its corresponding resampled data eigenvalue. The two-factor solution (χ2 = 208.11, df = 63, TLI = .95, CFI = .97, RMSEA [95% CI] = .07 [.06, .08], SRMR = .02, AIC = 13,307.21, BIC = 13,539.29) fits the data better than the one-factor solution (χ2 = 317.61, df = 76, TLI = .94, CFI = .95, RMSEA [95% CI] = .08 [.07, .09], SRMR = .03, AIC = 13,452.34, BIC = 13,630.54) in the ESEM results. In the two-factor ESEM solution, one factor is mainly loaded by extreme satisfaction items and the other factor is mainly loaded by extreme dissatisfaction items. Moderate items load on both factors, with a weak tendency to prefer the satisfaction factor. The cross-loadings of the moderate items on both factors imply that they correlate reasonably well with both factors. The correlation between the two factors is far from −1 (r = −.82, 95% CI = [−.74, −.90]). A summary of the factor loadings is reported in Table 4.
Standardized Item Loadings in Exploratory Structural Equation Models With Geomin Rotations on All Items
Note. Loadings > .40 are in bold. No participants completed both negated extreme items and negated moderate items for life satisfaction.
Omega Reliability and Factor Loadings
Table 6 shows the omega reliabilities. Extreme items have lower reliability than moderate items, perhaps because there are fewer of them. Consistent with the job satisfaction results, in both the extreme and moderate versions, negated polar opposites have the lowest reliabilities. In the SEM solution, one of the double negation items has a low factor loading (“NOT dislike my life”; λ = .62; see Table 5). The other two double negation items (“NOT discontent with my life” and “NOT dissatisfied with where I am in life”) have lower factor loadings than their nonnegated, regular counterparts (“content”; “satisfied”) in the one-factor ESEM solution. Double negation is apparently not the only problem, however, because another item also has poor loading (“Other people NOT living better lives than my own”) in the one-factor ESEM solution. This item may be especially demanding because it requires comparative judgment on top of negation.
Standardized Item Loadings in Exploratory Structural Equation Models With Geomin Rotations on Extreme and Moderate Job Satisfaction Items
Note. The word “NOT” is capitalized in the survey to minimize careless responses. Items can be found on the following website: (Blind for review purpose). Loadings > .40 are in bold. The one-factor solution for extreme items fails to achieve successful convergence. Only item stems are shown. Items can be found in the following website (https://osf.io/a8p9h/).
Correlations (Lower Diagonal) and Omega Reliability (Diagonal) [95% Confidence Interval] Among Item Types
Note. Scale unreliabilities were corrected using confirmatory factor analysis. Bonferroni comparisons were conducted to compare four response times for each wording type (i.e., one comparison for extreme wording and another comparison for moderate wording). Different subscripts represent statistically significant difference. For correlations, all ps < .001. Omega reliabilities are shaded in gray. Correlations over .90 are in bold.
Response Time Analysis
Table 6 shows mean response times for each condition. For extreme items, the main effect of keying direction was not statistically significant, F(1, 251) = 0.01, p = .913, but the main effect of negation was significant, F(1, 251) = 105.19, p < .001, with negated items having longer response times than nonnegated items, Ms = 4.10 vs. 3.11, Cohen’s d = 0.32. The interaction between keying direction and negation was statistically significant, F(1, 251) = 11.32, p < .001, with the response time increase for negated items greater for polar opposites than for regular items (difference = 1.31 vs. 0.70).
For moderate items, the main effect of keying direction was not statistically significant, F(1, 213) = 1.81, p = .180, but the main effect of negation was significant, indicating that negated items take longer to respond to than nonnegated items, F(1, 213) = 31.04, p < .001, Ms = 4.72 vs. 3.51, Cohen’s d = 0.21. The interaction between keying direction and negation was also statistically significant, F(1, 213) = 7.33, p = .007. Again, the response time increase for negated items was greater for polar opposites than for regular items (difference = 1.73 vs. 0.68). Thus, the consistent finding over moderate and extreme items is longer response times for negated items, especially when the item is also a polar opposite.
Discussion
The current study explores a long-debated question: Why do regular and reversed items load on separate factors? We investigated two major explanations—response difficulty and item extremity. Although the current study finds support for the existence of response difficulty, it favors the extremity position.
Response Difficulty
The consistent finding across both constructs (job satisfaction and life satisfaction) is that negated items take longer to respond to than nonnegated (regular and polar opposite) items, lending support to the claim that the addition of a negation term such as “not” increases response difficulty (e.g., Baumgartner et al., 2018). Interestingly, polar opposite items do not always take longer to respond to than regular items but interact with item negation to cause longer response times (with the only exception for extreme job satisfaction items in the current study).
In support of the response difficulty perspective, items with the word “not” may have lower factor loadings, thus damaging the scale’s reliabilities, when they are combined with polar opposites. Negated polar opposites often involve double negation (e.g., “not dissatisfied”): explicit negation with the words “no” or “not,” together with implicit negation provided by a prefix (e.g., the “dis” in “dissatisfied”). Double negation may be confusing. When an item involves only one negation (e.g., negated regular items), the damage to reliability is not apparent in the current study. The damage could, however, be higher among younger respondents (not in the current study) whose verbal reasoning skills are not yet fully developed. For example, Benson and Hocevar (1985) found that Grade 4 to 6 children had more trouble with negated regular than regular items, and Gnambs and Schroeders (2020) found that the discrepancy between regular and reversed items was larger among younger respondents low in verbal reasoning ability.
It is noteworthy, however, that double negation is not the only source of low reliabilities and higher response times. The moderate item “Other people NOT living better lives than my own” (actual items are shown on the OSF website) features single negation, but it is rather long and it requires a comparative judgment; the combination may account for its very low factor loading (λ = .48) relative to other life dissatisfaction items in the one-factor ESEM solution (Table 5). Speculatively, a comparative judgment item may be more taxing on participants’ cognitive resources than other types of judgment. The original item (“Other people living better lives than my own”), taken directly from the Riverside Life Satisfaction Scale (Margolis et al., 2019), also had a lower factor loading (λ = −.72) than other life dissatisfaction items (λs = −.87 to −.93). Furthermore, longer items, in general, may be more difficult to understand.
Item Extremity
And yet response difficulty is unlikely the only cause of bidimensionality. Response difficulty fails to explain why the bidimensionality results appear in the extreme but not in the moderate item condition. Furthermore, even when polar opposites involve double negation (which exacerbates response difficulty), they may still load on the same factor as “normal” polar opposites, implying that response difficulty alone is not sufficient to cause items to load on separate factors. The message here is that reversed items may form a distinct factor from regular-keyed items even in situations without much comprehension difficulty (e.g., between regular and polar opposite items). We are thus led to conclude that bidimensionality owes more to item extremity than to response difficulty, at least among adults with reasonable reading ability, as in the current study. The result is consistent with the major argument, discussed earlier, of the logical response perspective: Respondents who are capable of answering survey questions may still show an inconsistent response pattern commonly but erroneously described as “illogical” or “misresponse.” Response difficulty plays a greater role for participants still developing their comprehension skills (Benson & Hocevar, 1985), whereas item extremity effects are likely prevalent across most age and ability groups.
Implications for Construct Dimensionality Decisions
Item extremity may explain why regular and reversed items load on separate factors. Specifically, polar opposite items, but not negated items, are likely to load on a distinct factor when item content is extreme. Traditionally, some scales such as personality measures typically use reversed items to counter acquiescent response style (e.g., Soto & John, 2017). Many reversed items are polar opposites—not negated items—possibly due to the desire to make comprehensibility as high as possible (Baumgartner et al., 2018). Scale developers, however, tend to overlook item extremity. Thus, polar opposite and regular items may not load on the same factor, causing researchers to question the presumptive unidimensionality of a construct. The situation is not without irony: By avoiding negated items, and by not attending to item extremity when developing polar opposite items, scale developers may actually be inadvertently contributing to dimensionality confusion—the exact opposite of their intentions of minimizing response style artifacts.
Factor analysis is a common tool for discovering the dimensionality of a construct. Scale developers create a pool of items, which they then subject to factor analysis (Lambert & Newman, 2022); items that do not load on the same factor are excluded. This procedure, however, results in scales that are highly similar in item characteristics (e.g., keying direction). Scales with an excellent profile of unidimensional psychometric indices tend to be admired, although they may have insufficient domain sampling, resulting in circumscribed validity (Clifton, 2020).
The lesson of the current study is that empirical techniques such as factor analysis should not be decisive in determining a construct’s dimensionality. Rather, the determination should be the result of solid theoretical consideration.
Consider job satisfaction. The construct has been defined as one’s global evaluation of their job (Weiss, 2002). Researchers who advocate bidimensionality between satisfaction and dissatisfaction (Credé et al., 2009), however, have suggested that people may form a positive (satisfaction) attitude and a negative (dissatisfaction) attitude about an object, each of which may have unique antecedents and consequences. A job may fulfill my competence need (“it gives me a strong sense of agency”) but have overly long working hours. A worker may have a satisfying relationship with coworkers, but the job requires a time-consuming daily commute. One’s job may evoke both a positive feeling and a negative feeling because every job has some positive and some negative aspects. Therefore, the question about the dimensionality of job satisfaction is whether individuals formulate a cohesive global evaluation of their job. If an employee synthesizes a unified evaluative judgment based on their feelings, then surely the construct should be considered unidimensional. The current study’s empirical results strongly support this unidimensional conclusion, as the correlation between (moderate) job satisfaction and dissatisfaction items did not significantly differ from −1.
Implications for Inclusion of Reversed Items
If a construct is theoretically unidimensional, the next important question is this: Should a researcher include reversed items? Many researchers have advised excluding reversed items because switching the keying direction may confuse respondents. However, the current study did not find much lower reliabilities or factor loadings for negated regular items and polar opposites than for regular items. A participant who fails to notice keying direction may also fail to notice subtle nuances in content, which may also lead to inaccurate answers. Surveys deliberately constructed out of homogeneous item content (only with regular items, for example) may inadvertently inflate item correlations. Some careless response detection methods rely on the inclusion of reversed items (e.g., Ulitzsch et al., 2022), although inclusion of those items will inevitably lengthen a scale and cause respondents’ fatigue.
Despite our recommendation to include reversed items, certain types are probably best avoided. Compared with polar opposites (“My job is bad”), negation items (“My job is not good”) may not vastly widen the domain sampling of the original item (“My job is good”). Similarly, double negation items—which pose even greater challenges—were found here to have lower factor loadings; thus, they should be excluded. Other cognitively taxing items, such as those requiring a comparative judgment, should not be combined with negation. Finally, various methods of detecting careless responding should be used to exclude invalid responses before data analysis (Kam & Meyer, 2015).
Some researchers have suggested that another reason for avoiding reversed items is that reversed items have their own method-specific variance, which is hard to manage. According to this view, excluding reversed items will eliminate the method-specific variance related to reversed items. Our counterargument is that regular items still create their own method-specific variance. Excluding reversed items leaves the contamination of method-specific variance from regular items unchecked. By excluding reversed items, researchers will have no method to estimate or exclude method variance specific to regular items, which may cause regular items (and their constructs) to have inflated correlations with each other (see Kam & Meyer, 2015, for a detailed discussion).
In sum, we encourage researchers to use both regular and polar opposite items for construct measurement. However, the use of negation words such as “not” or “no” should be avoided because they make comprehension more difficult without broadening a construct’s domain sampling.
Limitations
The current study compares two main explanations for the bidimensionality of regular items and reversed items in factor analysis. Some limitations should be noted. First, we capitalized the word “NOT” to minimize the risk of participants missing the keyword. We did so to isolate the effect of item negation from careless responding and to compare negated and nonnegated items. The results suggest that negated and nonnegated items have a similar nature and thus load on the same factor. Without isolating the effect, any bidimensionality between negated and nonnegated items can result from careless responding or comprehension difficulty, and we cannot isolate the actual cause. Future research should include but not capitalize the negation word to examine the incremental influence of careless responding on negation items.
The second limitation is that we used adult native speakers of English who were born in the United States, and thus they are supposed to be competent in answering survey questions. Misresponses are likely to increase if respondents are not fully developed speakers of the language, such as young children and adolescents.
Finally, we have tested our reasoning only with satisfaction-related constructs (job satisfaction and life satisfaction), owing to the current debate about their dimensionality (Boes et al., 2010a, 2010b; Credé et al., 2009). Although there is no apparent reason to question the generalizability of the findings beyond those two constructs, future research may still want to examine the replicability of our results with other constructs, such as optimism.
Conclusion
The current study compared two explanations for the cause of a factorial distinction (bidimensionality) between regular and reversed items. Results favored the item extremity over the response difficulty explanation: Item characteristics are the most likely reason regular and reversed items load on separate factors. Findings emphasize the importance of robust theoretical analysis, rather than pure empirical results, in determining the dimensionality of a construct.
Footnotes
Declaration of Conflicting Interests
The author declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: The current research is financially supported by Multi-Year Research Grant (MYRG2022-00073-FED) from the University of Macau to the author.
