Abstract
We propose the use of balanced item parcels to account for method effects caused by acquiescent responding. The use of balanced parcels avoids the need to model method effects explicitly and results in a parsimonious specification of measurement and full structural equation models in the presence of unwanted method effects, particularly when a scale consists of a relatively large number of items. Balanced item parcels are sums or averages of individual items consisting of an equal number of regular and reversed items measuring the same construct. When regular and reversed items are combined into parcels, method effects cancel out (assuming that the method effects affecting the regular and reversed items in a parcel are equal in magnitude), and model fit and parameter estimates will no longer be negatively affected by acquiescent responding. We discuss why balanced item parceling works and when it is likely to prove beneficial, and we present a step-by-step procedure explaining how to use balanced item parceling in practice. We also report a brief hypothetical example to illustrate the proposed approach.
Keywords
Many constructs in organizational research are measured by means of scales consisting of multiple items presented in a Likert format, in which respondents have to indicate the strength of their (dis)agreement with each item on a rating scale ranging from, say, strongly disagree to strongly agree. In principle, the observed responses can be used directly as manifest indicators of latent constructs in a confirmatory factor analysis or structural equation model. However, if measurement scales contain many items, it may be infeasible or at least unwieldy to use individual items as indicators, and researchers sometimes measure constructs based on parcels of items, that is, indicators created by summing or averaging subsets of individual items within scales or subscales (Holt, 2004).
The use of parceling in empirical research has been controversial. Several shortcomings of this technique have been identified (see Bandalos, 2002; Bandalos & Finney, 2001; Little et al., 2002; Marsh et al., 2013; Rhemtulla, 2016). First and most importantly, parceling may hide deficiencies in individual scale items and misrepresent the factorial structure of the measurement scale. For example, the poor convergent validity of individual items may be concealed, or a factor solution based on parceled indicators may appear to be unidimensional when the individual items are in fact multidimensional. Second, in some studies, parceling has been found to lead to biased estimates of (structural) model parameters, although in other studies, parceling was shown to yield better estimates than models in which individual items were used as indicators. Based on these findings, the usual recommendation has been that parceling should not be used when the factor structure underlying a set of items is not well understood. In particular, when researchers want to develop a new measurement instrument or seek to assess the construct validity of an existing scale, individual items rather than parcels of items should be used in the analysis.
However, under the appropriate circumstances, aggregating items into parcels can have several advantages (Bandalos, 2002; Bandalos & Finney, 2001; Holt, 2004; Little et al., 2002). First, parceled indicators tend to have better distributional properties than individual items (i.e., they are more continuous and more normal). Second, models based on parceled items are simpler, require the estimation of fewer parameters (which may be advantageous when the sample size is small), and often fit the data better. Third, parceled items tend to be more reliable and yield more stable parameter estimates.
In this research report, we identify another advantage of item parceling that has not been examined in prior research. Item parceling is also useful for dealing with one of the main drawbacks of using reversed items in measurement scales. Reversed items are items whose keying direction is opposite to the polarity of the construct being measured (e.g., “tends to be quiet” vs. “is talkative” in an extraversion scale). On the one hand, reversed items can be beneficial because they broaden the retrieval strategies that respondents use when probing their memories for relevant content, which should enhance content validity (Tourangeau et al., 2000; Weijters & Baumgartner, 2012); they serve to keep respondents alert, preventing them from mindlessly checking the same response to a series of related questions (Podsakoff et al., 2003); and, most importantly for the present purposes, they correct for bias due to acquiescent responding (Baumgartner & Steenkamp, 2001; Mirowsky & Ross, 1991; Paulhus, 1991). On the other hand, these advantages may be offset by the tendency of reversed items to introduce method variance due to acquiescent responding, which, when unmodeled, can lead to poor fit of models to data, reduced internal consistency of the items, and complex and possibly misleading factor structures (e.g., a unidimensional construct may appear to be multidimensional).
Item parceling allows researchers to account for method effects caused by reversed items without having to model acquiescent responding explicitly. However, in contrast to the usual practice of combining items into parcels randomly, the parceling must be done strategically by combining an equal number of regular and reversed items within each parcel. We call this balanced item parceling. As long as the underlying construct is unidimensional and regular and reversed items within a parcel are affected equally strongly by method effects, balanced item parceling takes care of method effects implicitly (by counterbalancing the keying direction of the items within each parcel) and results in a parsimonious specification of measurement and full structural equation models because there is no need to introduce method factors or correlated uniquenesses to model method effects explicitly.
In the remainder of this research report, we discuss why balanced parceling works, when it can be expected to have beneficial effects, and how it should be used in practice. We also present a brief hypothetical example to illustrate the use of balanced parceling.
Why Balanced Parceling Works
The primary advantage of balanced parceling is that it is an effective strategy for controlling certain types of method effects. Although Likert-type agree-disagree response scales are the most commonly used scale type, they are susceptible to (dis)acquiescent responding. That is, people’s responses may reflect not only the construct of interest but also individual differences in respondents’ tendency to agree or disagree with items regardless of content (Billiet & Davidov, 2008; Billiet & McClendon, 2000; Cheung & Rensvold, 2000; Danner et al., 2015; Kam, 2016; Kam & Meyer, 2015; McClendon, 1991; Watson, 1992; Weijters et al., 2010). If a scale contains no reversed items, it is impossible to distinguish between substantive and stylistic responding (i.e., a person scoring high on an extraversion scale could be an extravert or an acquiescent responder) unless special scales are available to measure acquiescent responding directly and control for it by including acquiescence as a covariate. In contrast, if a scale contains an equal number of regular and reversed items (i.e., the scale is balanced), the biasing effect of acquiescence on overall scale scores can be avoided because the upward bias for regular items is neutralized by the downward bias for reversed items (Baumgartner & Steenkamp, 2001; Paulhus, 1991). Nonetheless, although overall scores are purged of method effects, a measurement model for the individual items in a balanced scale will still have a poor fit to the data unless method effects are modeled explicitly.
Parceling provides an alternative that also corrects for acquiescence but has several advantages over other approaches. First, in contrast to combining regular and reversed items into a single overall score (which corrects for measurement error only incompletely), multiple items are retained, and measurement error in the parceled indicators can be taken into account explicitly. Second, in contrast to using individual items as indicators, the specification of method factors or correlated uniquenesses can be avoided (because method effects are controlled for at the parcel level), and a congeneric measurement model containing only substantive factors should yield an acceptable fit to the data (assuming the model is otherwise specified correctly). As the number of factors and the number of items per factor get larger, it becomes increasingly more likely that a measurement model based on the individual items will fit the data poorly, and the explicit modeling of method factors gets more difficult (in terms of sample size requirements and estimation problems). The use of balanced item parceling becomes an attractive modeling strategy in this situation.
Appendix A, available in the online version of the article, demonstrates analytically why parceling is effective provided certain conditions are satisfied. Specifically, the parceling must be done such that an equal number of regular and reversed items is combined into each parcel. We call this balanced parceling because it requires a balanced scale in which half the items are regular items and half the items reversed items. Intuitively, parceling eliminates method effects because acquiescence has countervailing effects on regular and reversed items, and when the two effects are equal, they cancel each other out, and there is no need to model method effects explicitly. For example, if individual differences in extraversion–introversion are measured with two items, “I am someone who is talkative” and “I am someone who is outgoing, sociable,” as in the Big Five Inventory (John & Srivastava, 1999), someone who strongly agrees with both items could be an extravert or an acquiescent responder. In contrast, an acquiescent responder would agree with both “I am someone who is talkative” and “I am someone who tends to be quiet,” and once the second item is recoded and the two items are combined, acquiescent responding will not affect the aggregated score (assuming that acquiescence affects both items to the same extent).
If only two items (one regular, one reversed) are combined into a parcel, the method effects for the two items have to be equal (otherwise the upward bias caused by acquiescence on the regular item will not be neutralized by the downward bias on the reversed item) and the validity of this assumption has to be tested. How this can be done (without specifying a measurement model for the individual items) will be discussed in the following. If multiple regular and reversed items are combined into a parcel, the requirement is that, on average, the method loadings of regular items equal the method loadings of reversed items. Again, the validity of this assumption can be tested based on the fit of the model with parceled indicators, as described in the following.
In summary, balanced parceling eliminates method effects due to acquiescence from the parceled indicators because indicators that neutralize acquiescence are strategically combined. As a result, systematic errors caused by acquiescent responding do not have to be modeled explicitly, and more parsimonious measurement models can fit the data well.
When Balanced Parceling Is Likely to Prove Beneficial
For balanced parcels to be free of method variance due to acquiescence, the method loadings of regular and reversed items have to be equal. Appendix B, available in the online version of the article, demonstrates analytically why this assumption must be satisfied and what happens if it is not. Intuitively, the upward (downward) bias caused by (dis)acquiescence on regular items has to be exactly offset by the downward (upward) bias caused by (dis)acquiescence on reversed items. As mentioned previously, it is advantageous to have multiple regular and reversed items within each parcel because in that case, the assumption that the method loadings of regular and reversed items be equal is more likely to be satisfied. The reason is that as long as the average of the method loadings for the regular items is equal to the average of the method loadings for the reversed items, the upward and downward biases caused by acquiescence for regular and reversed items, respectively, will offset each other. Therefore, an increasing number of items per factor will not only simplify the measurement model when parceled indicators are used (relative to the situation in which a measurement model is specified for the individual items) but also make the assumption of equal method loadings more tenable (because multiple regular and reversed items can be combined into a parcel).
Balanced parceling is most straightforward for balanced scales (i.e., scales that consist of an equal number of regular and reversed items). In this case, one simply combines each regular item with a reversed item; alternatively, several regular and reversed items can be aggregated into a parcel (if the number of scale items is a multiple of 4). When a scale contains no or very few reversed items, balanced item parceling cannot be used. However, when a scale is nearly balanced, a variant of balanced item parceling is possible. Assume that an instrument consists of six regular and five reversed items. In this case, one could form two parcels consisting of two regular and two reversed items each and one parcel consisting of two regular items and one reversed item. However, to avoid overweighting of the regular items in the third parcel, the two regular items should be averaged before combining the (averaged) regular items with the reversed item (see Appendix B in the online version of the article).
Although not all measurement scales contain reversed items and some contain only a few, key constructs in organizational research are measured by means of balanced, or at least approximately balanced, scales. Examples include dispositional optimism (Li et al., 2019; Scheier et al., 1994), self-esteem (Liu et al., 2019; Rosenberg, 1965), core self-evaluations (Judge et al., 2003), and socially desirable responding (Paulhus, 1991). Job satisfaction has been measured with balanced scales (Kam & Fan, 2020; Kam & Meyer, 2015), and some scales are approximately balanced; for instance, the Job Satisfaction Survey has 36 items with 19 reversed items, and although not all subdimension scales are balanced, balanced parceling could be used for the overall measure (Spector, 1997). In the personality area, the HEXACO personality factors (Bourdage et al., 2015; Hershfield et al., 2012; Lee & Ashton, 2004, 2007, 2018) are measured with approximately balanced scales, with each of the six personality dimensions containing 10 items, four to six of which are reversed.
Balanced parceling is most relevant for longer scales, and ideally, at least eight items, four of which are reversed, should be available. Although balanced parceling can be used for a four-item scale consisting of two regular and two reversed items, a one-factor model with two (parceled) indicators is not identified (without additional restrictions), and another factor has to be included in the model to achieve identification. Even a model with three parceled indicators (based on a six-item scale with three regular and three reversed items) has zero degrees of freedom, which makes model testing impossible for a single-factor model. If a scale contains at least four regular and four reversed items, a stand-alone factor model will be overidentified, and the fit of the model can be tested.
In summary, as implied by the name, balanced parceling requires balanced or nearly balanced scales, and it will probably be an attractive modeling strategy and perform better for longer scales (consisting of eight or more items). The most critical prerequisite for the balanced parceling approach to work well is that the (average of the) method loadings for the regular items equal the (average of the) method loadings for the reversed items (within a given parcel). This assumption can and should be tested, but the assumption of equal method loadings is less stringent when parcels are formed based on multiple regular and reversed items.
How to Use Balanced Parceling in Practice
If a researcher decides to use parceling and the aforementioned conditions are met, the step-by-step procedure explained in this section can be used. As a first step, we recommend creating two different parcel allocations, which we refer to as balanced parceling and isolated parceling. As already explained, in the balanced parceling approach, each parcel consists of an equal number of regular and reversed items measuring the same construct. In the isolated parceling approach, each parcel consists of either regular or reversed items (but not both).
For concreteness, consider a scale consisting of eight items, four regular items (p1, p2, p3, p4) and four reversed items (n1, n2, n3, n4). Assume that the reverse-keyed items have not been recoded so that greater agreement on these items indicates a lower score on the underlying construct. With balanced parceling, the four parcels consist of averages of one regular and one reversed item (with the reversed item subtracted from the regular item). With isolated parceling, two parcels consist of the mean of two regular items, and two parcels consist of the mean of two reversed items. Table 1 illustrates this example.
Illustrative Parcel Allocation Scheme for Eight Items.
Note: This table shows how one can allocate eight items (four regular and four reversed items) to four parcels, according to two alternative parcel allocation strategies. In the polarity column, p versus n indicates whether the item is a regular (positive polarity) or reversed (negative polarity) item. Pa1 to Pa4 denote Parcels 1 through 4. For situations with more than eight items, additional items need to be allocated to the four parcels by extending the scheme (and reweighting items to ensure equal weights for regular and reversed items within each parcel in the balanced allocation). For instance, if the scale consists of 10 items, both two regular items and two reversed items can be averaged to get a scale with four regular and four reversed items.
In the next step, researchers should fit a congeneric factor model (with freely estimated factor loadings and unit factor variance) to the parceled indictors for both the balanced and the isolated parcel data (as illustrated in Figure 1) and compare the results. If the individual items are contaminated by acquiescent responding, model fit will be poor for the isolated parcel data and better for the balanced parcel data. In fact, if the regular and reversed items that are combined within a given parcel are equally affected by acquiescent responding and a unidimensional specification holds for the parceled indicators, the model using balanced parcels should fit the data well. In addition, one should inspect the modification indices (MIs) and expected parameter changes (EPCs) for the residual covariance terms (see Appendix B available in the online version of the article). If method variance due to acquiescent responding is present, one should observe the following: (a) In the isolated parcel data, the MIs for the covariances of the unique factors will be large (i.e., the covariances should be significantly different from zero as long as the statistical test has sufficient power) and the EPCs will indicate positive covariances between the residuals (assuming that all method loadings are positive, which should be the case). The reason is that the unique factors contain method variance due to acquiescent responding, which leads to correlated residuals across pairs of parcels. (b) In the balanced parcel data, acquiescence has been canceled out at the parcel level (provided that the method loadings within each parcel are equal), so none of the residual covariances should significantly deviate from zero. Note that because reversed items (which have not been recoded) are subtracted from regular items when forming balanced parcels, all substantive loadings are positive. However, in the model for isolated parcels, the substantive loadings of parcels based on reversed items will be negative.

Factor model with balanced versus isolated parcels.
In the third and final step, we recommend that a researcher proceed with the model based on the balanced parceling approach, which corrects for acquiescent responding, provided that the model estimated using isolated parcels fits the data poorly (which signals that method variance due to acquiescent responding is present) and the model estimated using balanced parcels fits the data well (which indicates that balanced parceling neutralized method variance due to acquiescent responding). If both parceling approaches result in satisfactory model fit, the researcher can proceed with either of the two models (or items can be allocated to parcels randomly) because there seems to be no need to control for acquiescent responding. If neither approach results in an acceptable fit, there are other measurement problems besides (or in addition to) the presence of acquiescent responding, and the researcher should critically evaluate the items in the scale, possibly dropping some of the items from the analysis (if there are clear reasons for doing so) and collecting new data to validate the modified scale.
It should be noted that although balanced parceling helps to account for method variance due to acquiescent responding, it is no substitute for actions taken to detect insufficient effort responding (Huang et al., 2012, 2015) or careless responding (Kam & Chan, 2018). Respondents who do not read items attentively and thus misrespond to instructed response checks (e.g., for this item, please select “strongly disagree”) are unlikely to provide meaningful answers, and their responses will not contribute substantive variance (even after controlling for acquiescence). Therefore, researchers should try to identify careless respondents based on instructed response items and other means (Kam & Chan, 2018) before using balanced parceling.
Empirical Illustration
Appendix C, available in the online version of the article, provides the R code to generate and analyze synthetic data to illustrate the balanced parceling approach described in the previous section, and Appendix D, available in the online version of the article, presents the corresponding code to run the simulation in Mplus. The example assumes that there are four regular and four reversed items, with the substantive factor accounting for 81% of the total variance in the individual items and the method and unique factors for 3% and 16%, respectively. We purposely chose a rather modest amount of method variance due to acquiescent responding to demonstrate that a model that does not account for method effects will fit poorly even when the method effects are small. From the individual-item data, we constructed two parceled data sets with four parcels each (first step), where each parcel consists of either one regular and one reversed item (balanced parcels) or two regular or two reversed items (isolated parcels), as shown in Figure 1. We assumed a sample size of 200 and generated 1,000 replications for each parcel allocation.
For the balanced item parcels, the average χ2 value across the 1,000 replications was 1.99 with two degrees of freedom; the average values of root mean square error of approximation (RMSEA), standardized root mean square residual (SRMR), Comparative Fit Index (CFI), and Tucker-Lewis Index (TLI) were 0.023, 0.003, 0.999, and 1.00, respectively. Thus, a congeneric factor model fits the data very well, as expected, because parceling eliminates the method effects contaminating the individual items when the method effects of items within a parcel are equal. In contrast, the fit of the congeneric model for isolated parcels was poor. The average χ2 value was 60.3, and the average values of RMSEA, SRMR, CFI, and TLI were 0.379, 0.025, 0.943, and 0.828, respectively. All six modification indices for the covariances between the unique factors were significant (range = 15.7–64.5), and the expected parameter changes suggested a positive covariance between the unique factors.
The averages of the estimated substantive loadings ranged from 0.895 to 0.896 for the model based on balanced parcels and from 0.889 to 0.890 (in absolute value) for the model based on isolated parcels. Although the estimated loadings were close to the true value of 0.90 even in the model for the isolated parcels, the bias and root mean square error were consistently lower in the model for the balanced parcels (for details, see Appendix C, available in the online version of the article). In the example, the method effects were assumed to be fairly small; if acquiescence has a stronger influence on people’s responses, the estimates of the loadings in the model for the isolated parcels will be less accurate.
Discussion
When a scale contains a relatively large number of items so that a measurement model using individual items as indicators is infeasible or impractical and if around half the items in the scale are reversed (i.e., the scale is approximately balanced), balanced parceling is an attractive modeling strategy because (a) multiple indicators are retained and measurement error can be taken into account explicitly (in contrast to the situation in which one overall composite of the available items is formed) and (b) method variance due to acquiescent responding, while present at the level of individual items, is neutralized at the parcel level and therefore does not have to be modeled explicitly (in contrast to a measurement model for the individual items). Balanced parceling is therefore an effective and efficient strategy for dealing with the problem of method variance caused by acquiescent responding. In this research report, we discussed the requirements that must be met for balanced parceling to be beneficial, and we offered a three-step approach on how to use balanced parceling in practice.
Previous research has offered modeling solutions to counter acquiescence bias when a researcher uses item-level, rather than parceled, data. In particular, several authors have proposed that the presence of acquiescence variance in item responses for balanced scales can be modeled by means of a method factor (Billiet & Davidov, 2008; Maydeu-Olivares & Coffman, 2006; Weijters et al., 2013). Thus, modeling solutions for acquiescence are available, and we do not want to suggest that parceling is the only, or generally preferred, approach for dealing with the problem. Rather, we propose that balanced parceling offers important advantages under specific circumstances, such as when a scale contains a relatively large number of items and approximately half of the items are reversed.
There is some evidence that the use of reversed items in measurement scales has declined in recent years, possibly because reversed items frequently create problems such as poorly fitting models when congeneric measurement models are specified and method variance due to acquiescent responding has not been taken into account. We want to emphasize that dropping reversed items does not solve the problem of acquiescent responding; it simply makes method variance due to acquiescence undetectable. A preferred approach might be to form balanced item parcels, which eliminates the need to consider method factors and thus results in a parsimonious specification for measurement models.
For longer scales, a decision may have to be made about the number of parcels and the size of each parcel. For example, in a scale consisting of eight regular and eight reversed items, one could form four parcels of four items or eight parcels of two items. We currently have no empirical evidence about which approach is preferable, but the decision basically involves a trade-off between having a sufficient number of parcels per factor (more indicators are generally better than fewer) and satisfying the assumption of equal method loadings (combining multiple regular and reversed items into a parcel makes this assumption less restrictive). If the method loadings of individual regular and reversed items within a parcel are approximately equal, more parcels are probably better than fewer parcels.
When fewer than eight items are available, the proposed approach to compare balanced and isolated parcel allocations does not work for six scale items because it is impossible to form three isolated parcels from three regular and three reversed items. For four scale items, one can construct two balanced and two isolated parcels, but at least two constructs have to be included in a measurement model for a two-factor model to be identified (unless other restrictions are imposed on the model). However, these complications are not a serious limitation because parceling is most commonly used and most useful for longer scales.
Finally, one important issue associated with parceling in general is that there are often many different ways in which individual items can be allocated to parcels, and each of these allocations will lead to somewhat different results. This is called the problem of parcel allocation variability in the literature (Sterba, 2011; Sterba & Pek, 2012; Sterba & Rights, 2017). To account for this parcel allocation variability and to ascertain how representative a given allocation is of all possible allocations, one can compute the average goodness of fit and the average parameter estimates across multiple parcel allocations. Syntax in SAS and R is available, which enables researchers to automatically create multiple data sets with different parcel allocations to empirically quantify the extent of parcel allocation variability; the reader is referred to the relevant literature for details (Sterba, 2011; Sterba & Pek, 2012; Sterba & Rights, 2017).
Supplemental Material
Supplemental Material, sj-pdf-1-orm-10.1177_1094428121991909 - On the Use of Balanced Item Parceling to Counter Acquiescence Bias in Structural Equation Models
Supplemental Material, sj-pdf-1-orm-10.1177_1094428121991909 for On the Use of Balanced Item Parceling to Counter Acquiescence Bias in Structural Equation Models by Bert Weijters and Hans Baumgartner in Organizational Research Methods
Footnotes
Acknowledgments
The second author gratefully acknowledges support from the Smeal Chair endowment.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
Supplemental Material
Supplemental material for this article is available online.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
