Abstract
This study aims to examine different scale usage correction procedures that are meant to enhance the cross-cultural comparability of Likert scale data. Specifically, we examined a priori study design (i.e., anchoring vignettes and overclaiming) and post hoc statistical procedures (i.e., ipsatization and extreme response style correction) in data from the 2012 Programme for International Student Assessment across 64 countries. We analyzed both original item responses and corrected item scores from two targeted scales in an integrative fashion by using multilevel confirmatory factor analysis and multilevel regressions. Results indicate that mean levels and structural relations varied across the correction procedures, although the psychological meaning of the constructs examined did not change. Furthermore, scores were least affected by these procedures for females who did not repeat a grade and students with higher math achievement. We discuss the implications of our findings and offer recommendations for researchers who are considering scale usage correction procedures.
Keywords
Making valid comparative inferences is contingent on obtaining comparable data across groups (van de Vijver & Leung, 1997). Data from Likert-scale measures often lack the level of comparability that allows for mean comparisons (i.e., scalar invariance) in large-scale educational, psychological, and sociological surveys (e.g., Byrne & van de Vijver, 2010; Zercher et al., 2015). Various procedures have been proposed to enhance the comparability of data from Likert-scale measures, either by modifying study designs or applying post hoc statistical corrections. These scale usage correction procedures aim to reduce the impact of response styles (i.e., systematic tendencies to respond in ways that are unrelated to the target constructs) and reference group effects (i.e., differences in the standards one uses to self-assess attitudes, conditions, and traits).
In large-scale studies, design-based, preventive measures are used much less frequently than post hoc statistical corrections. However, to date, there is no silver bullet for making data from different countries perfectly comparable and valid. This is likely due to researchers’ limited understanding of the nature of implementing such changes to the psychometric properties and external validity of Likert scale data. For instance, one study compared the effectiveness of four scale usage correction procedures on personality and value data for university students across 16 countries, but the use of such procedures did not result in scalar invariance of the target scales (He, van de Vijver, et al., 2017). Therefore, results yielded different conclusions about the relations among the target constructs, and did not provide a clear assessment of which procedure showed the highest validity. Similar conclusions were obtained from studies targeting design-based correction procedures applied in the 2012 Programme for International Student Assessment (PISA) data. In these studies, design-based corrections did not always result in scalar invariance (e.g., Marksteiner et al., 2019), and the relations with the criterion (i.e., achievement) differed between uncorrected and corrected scores (e.g., Vonkova, Papajoanu, et al., 2018). One limitation of this work is that each correction procedure was independently assessed, with the assumption that such corrections operate on data from all respondents to the same extent.
To our knowledge, the 2012 PISA is the only large-scale study that implemented multiple design features—including anchoring vignettes and overclaiming—to account for various scale usage differences (OECD, 2013). Therefore, this dataset is uniquely suited for the simultaneous examination of both design-based and post hoc corrections for scale usage differences in nationally representative data from students in 64 countries. In the current study, we aim to shed light on the changes that may result from various scale usage correction procedures with a novel integration of different procedures across multiple levels of analysis using the 2012 PISA data.
Specifically, we examine the commonalities and differences of different scale usage correction procedures in their impact on Likert scale data at the intraindividual, individual, and country levels. Using student data from the 2012 PISA, we investigate whether and how these corrections targeting scale usage differences change the psychological meaning (i.e., the extent to which the same construct is being measured with raw and corrected scores) and metrics of target constructs, and what individual and cultural characteristics are associated with the extent to which one’s scores are affected by these procedures. Here, we use the term psychological meaning to refer not only to configural invariance (i.e., do the same indicators load on a factor”) but also to the extent that raw and corrected scores reflect the same construct. Understanding what changes and for whom it changes when these procedures are applied can help researchers decide on their application more critically and contribute to studies that yield more valid cross-cultural conclusions.
Below, we introduce the sources of data incomparability and levels of data comparability. Then, we describe four scale usage correction procedures intended to enhance data comparability for which the 2012 PISA data are uniquely suited. We review the effects of these procedures across cross-cultural contexts, and describe the setup of our study.
Correction Procedures for Enhancing Data Comparability
Bias and Equivalence
Issues of bias and equivalence form the cornerstone of cross-cultural research, and provide a framework on sources of and remedies for data incomparability. Bias refers to systematic errors that threaten the validity of measures administered in different cultures (van de Vijver & Leung, 1997). Three sources of bias can be distinguished. Construct bias indicates that the target construct has a different meaning in different cultures. Method bias consists of sources of incomparability due to sampling, instruments, and administration. One main source of method bias in Likert-scale data, the focus of this study, is scale usage preferences. These scale usage preferences can stem from individual and cultural response styles (e.g., extreme response style) and reference-group effects resulting from different standards used by respondents to evaluate themselves and their behaviors. Item bias refers to differences in the item meaning across cultures, and occurs when there is a different probability of endorsing an item given the same trait level of individuals from different cultures.
Bias can negatively impact the level of comparability of scores across cultures. Three main levels of equivalence (i.e., measurement invariance) can be distinguished and statistically tested: Configural invariance means the construct is understood in the same way across cultures. In statistical terms, it means that items measuring a construct exhibit the same configuration of salient and non-salient factor loadings. Metric invariance indicates that items of the construct have the same factor loadings across cultures. With metric invariance, associations among variables can be compared across cultures (e.g., correlations between motivation and efficacy can be compared across cultures, if data from both scales reach this level of invariance). Scalar invariance is achieved when items have the same intercepts (i.e., point of origin) and factor loadings across cultures. Only with scalar invariance can mean scores of scales be validly compared across cultures. These three levels of invariance are not only relevant in cross-cultural comparisons, but also across time for longitudinal data (Urbán et al., 2014), and across levels of analysis in multilevel data (e.g., Jak et al., 2014).
The link between score corrections and equivalence is based on a tacit assumption that differences in the ways respondents use scales (among other sources) pose a challenge for score comparability. Furthermore, it is assumed that the use of corrections for scale usage differences can “remove” method and item bias while not changing the psychological meaning of the measurement, and that comparability should be enhanced by these correction procedures. It is important to note that although some of these correction procedures are widely used, there is not much evidence to support the assumptions behind their usage, nor is there evidence to disconfirm them.
Procedures to Correct for Scale Usage Differences
Several procedures are presumably capable of capturing and correcting for scale usage differences due to different mechanisms and they have been applied in cross-cultural settings. These include a priori design features such as anchoring vignettes and direct assessment of overclaiming (as implemented in the 2012 PISA student questionnaire), and statistical corrections (without the collection of additional data) such as ipsatization and post hoc construction and correction of response styles.
Anchoring vignettes
Anchoring vignettes involve item batteries to correct for reference group effects (Hopkins & King, 2010; King et al., 2004). With this technique, respondents are asked to rate themselves on the target construct, and, on the same scale, to rate hypothetical persons with differing levels of the target construct described in short vignettes. Whereas the systematic differences in ratings on the same vignette supposedly reflect mainly scale usage differences, ratings on the self-assessment are a combination of such distortion and the true trait level. Therefore, the measurement bias due to reference group differences from the self-assessment can be removed by rescaling the self-assessment according to its position relative to the vignette rating (via a parametric or nonparametric application of rescaling).
There are two assumptions underlying anchoring vignettes: response consistency refers to the plausibility that respondents rate themselves and the hypothetical persons described in vignettes the same way, and vignette equivalence requires that the vignettes are understood by all respondents in the same way (King et al., 2004). These assumptions are usually taken for granted and rarely tested empirically. A few empirical studies on the tenability of these assumptions reported rather mixed results (Grol-Prokopczyk et al., 2015; He, van de Vijver, et al., 2017; Jürges & Winter, 2013; Kapteyn et al., 2011), indicating that anchoring vignettes may not serve as the gold standard to correct for individual differences in scale usage without introducing other measurement bias to the data.
Overclaiming
The assessment of overclaiming aims to capture respondents’ self-enhancement tendency independent of one’s ability (Paulhus et al., 2003). Self-enhancement refers to a strategy to maintain a positive image, which is moderated by cultural values. Heine et al. (1999) argued that people from Western cultures tend to self-enhance and show a better-than-average bias in the ability domain, whereas people from East Asian cultures have a stronger self-criticism focus and thus show a modesty bias in their self-reports, resulting in different scale usages and poor cross-cultural comparability. Overclaiming is also found to be related to country-level prevalence of rule violations (Fell et al., 2019). The overclaiming technique was developed to detect and correct for such differences. Respondents are asked to rate their knowledge or familiarity on various concepts, persons, events, and products, some of which are nonexistent; providing ratings of familiarity on nonexistent items (i.e., overclaiming) indicates self-enhancement bias. Subsequently, this overclaiming indicator can be used to correct for scale usage differences due to self-enhancement.
One limitation of this approach is that it is restricted to the knowledge domain, and it may be confounded with one’s memory (Paulhus et al., 2003). Secondly, whether the overclaiming technique indeed measures self-enhancement bias is questioned, as it is argued that the endorsement of nonexistent concepts is related to meta-cognitive processes such as sense-making and misattribution of fluency (Müller & Moshagen, 2018). These potential pitfalls can cast doubt on the validity of its correction.
Ipsatization
Ipsatization is one approach to standardize scores. Score standardization involves adjustment procedures using the mean and (less frequently) standard deviation of individuals (i.e., within-subject) or of each cultural group (i.e., within-group) to partial out scale usage differences (Fischer, 2004). Ipsatized item scores are produced by subtracting the mean of a range of items for each individual from the raw response on each item.
Ipsatization has several potential problems. Subtracting individuals’ mean rating from a set of items may remove genuine, meaningful individual differences when it is unrealistic to assume that the sum of all item ratings is supposed to be identical across individuals. As the sum of all ipsatized scores for each respondent is zero, each ipsatized item score is dependent on scores of other items; the average inter-item correlation among these scores tends to be negative, creating a data dependency that can complicate the analysis of the variance—covariance data.
Post hoc response style construction and correction
Respondents may differ in their response style preferences. The most frequently studied styles include acquiescent response style (ARS; the tendency to agree with an item regardless of its content), extreme response style (ERS; the tendency to use the end points of a response scale), midpoint response style (MRS; the tendency to endorse the middle response categories), and socially desirable responding (SDR; the tendency to present oneself in a positive light). These specific response styles are interrelated (e.g., He & van de Vijver, 2013; Smith & Fischer, 2008) and can be mapped as indicators on a continuum that ranges from response moderation (e.g., ARS and MRS) to response amplification (e.g., ERS and SDR). ARS, ERS, and MRS can be measured indirectly (i.e., using post hoc analyses), through a simple procedure that counts the number of specific responses (e.g., extreme responses) of heterogeneous content and Likert-scale responses (De Beuckelaer et al., 2010). The idea behind this procedure is to recode the Likert-scale responses from a set of items measuring different constructs for the presence and absence of response styles (e.g., recoding original responses on the two endpoints as the presence of ERS and recoding other responses as the absence of ERS) and then summing the recoded scores as an indicator of ERS. This indicator then can be used as a covariate in following analysis to control for the response style. It should be noted that this operationalization of post hoc construction of response styles depends on the availability of suitable data. Ideally, all specific response styles should be extracted with non-overlapping sets of items to ensure their independent assessment, whereas it is often the case with limited data that one can extract one response style and only infer the associations with other response styles. We focus on ERS in the current study because ERS extracted from Likert-scale data of various targeted constructs, in comparison to MRS and ARS, shows higher reliability and consistency (e.g., He & van de Vijver, 2013).
Despite the straightforward operationalization and application of these response styles, the validity of response bias indicators and their use for correcting scores is a continually contested topic in assessment (e.g., Rohling et al., 2011). On the one hand, these response styles may be nuisances that need to be statistically accounted for, but on the other hand, they may also represent valid cultural and individual differences in personality and values.
Impact of these Procedures on Data Comparability in Large-Scale Surveys
Applications of the aforementioned procedures are common, and their effects can result in changes in psychometric properties of target constructs, criterion and predictive validity of scales (i.e., structural relations) and mean comparisons across groups (e.g., Bartram, 1996; Vonkova, Papajoanu, et al., 2018). We argue that correction effects on structural relations and mean comparisons should be built on sound psychometric properties of corrected scores. Therefore, we first focus on the impact of scale usage correction procedures on the psychometric properties of scales, and then highlight potential changes on structural relations and mean comparisons in cross-cultural settings.
A key criterion in improving the psychometric properties involves the enhancement of measurement invariance. We assume that if any of the abovementioned procedures enhances the comparability of cross-cultural Likert-scale data, corrected scores should show a higher level of measurement invariance than raw scores. Two applications of anchoring vignettes in the PISA data from 2012 showed enhanced measurement invariance of scales across dozens of countries (He, Buchholz, et al., 2017; Marksteiner et al., 2019). However, anchoring vignettes did not result in scalar invariance (e.g., the mean scores of rescaled data cannot be compared validly across cultures) in all of the target scales in these studies. The usefulness of applying ipsatization and response style correction in social psychological studies of values and personality traits was shown via a clearer factor structure and more comparable mean scores. For instance, ipsatized value ratings with the Schwartz Value Survey in multidimensional scaling showed a rather stable structure of the 10 basic human value dimensions across countries, although strict scalar invariance was rarely found (Schwartz et al., 2001). When ARS was controlled for, a better recovery of the Big Five personality factor structure was found in individualistic countries, but much less so in collectivistic countries (Rammstedt et al., 2013). Furthermore, none of these procedures helped the data reach scalar invariance (and in some cases, metric invariance) in a 16-country study on personality and values (He, van de Vijver, et al., 2017). There is little research directly assessing the impact of overclaiming on psychometric properties of target constructs in cross-cultural contexts.
Many studies have reported a substantial impact of scale usage correction procedures on structural relations and mean comparisons in cross-cultural surveys. Yet, due to the lack of validity evidence using measures other than self-reports, it is difficult to determine if corrected scores are unequivocally more valid. For instance, when overclaiming in math familiarity was corrected for in the PISA data, the weak and nonsignificant correlation between math familiarity and math achievement became strong and significant in 64 countries (Vonkova, Papajoanu, et al., 2018). Additionally, after performing an anchoring vignette correction on the Classroom Management scale in the PISA data, it was revealed that there may be different implicit standards of self-assessment across countries, and this resulted in substantial change in correlations with students test scores and public expenditure per pupil (Vonkova, Zamarro, et al., 2018). The famous motivation-achievement paradox (i.e., a positive correlation between Likert-scale data assessing motivational factors and achievement at the individual level, but negative correlation when scores were aggregated at the country level) was partially explained or alleviated when anchoring vignettes, overclaiming, and ERS corrections were applied to the Likert-scale data (e.g., He & van de Vijver, 2015b; Kyllonen & Bertling, 2014). These studies seem to suggest higher validity for structural relations among corrected scores. Additionally, using three large surveys of middle-aged and older adults, Mojtabai (2015) found that although American respondents rated themselves as more depressed in the vignettes than European respondents, scores adjusted by anchoring vignettes revealed that American respondents were actually less depressed than most of their European counterparts, with the exception for those from two European countries. Without additional evidence, it is difficult to assess the extent to which this approach resulted in higher validity.
Alternatively, documented changes may not point to substantial changes in the structures of the measure, its relations with other validity measures, and mean comparisons. For instance, an investigation of ipsatization of the Schwartz Value Survey with different methods (i.e., treating data as interval or not) revealed the lack of impact on relationships between the Schwartz value types and other constructs. In the Teaching and Learning International Survey, self-reported teaching beliefs, attitudes, and practices among a nationally representative sample of secondary school teachers in 38 countries were examined. Results indicated that there were very limited changes in country ranking on most self-reported Likert scales and in the sizes of cross-cultural differences with and without corrections for response styles (He & van de Vijver, 2015a). Another study focusing on political interests also reported minimal change of country ranking with adjusted scores based on anchoring vignettes in 12 countries (Lee et al., 2016).
The Present Study
In sum, previous research suggests that a growing body of literature focuses on the effects of applying a single scale usage correction procedure in cross-cultural research. This is likely due to a focus on relating corrected and uncorrected scores to other constructs (e.g., for predictive validity) than examining their impact on the psychometric properties of target scales. Simultaneous applications and comparisons of different scale usage correction procedures are rare, and these studies tend to focus on the relative merits and limits of each procedure (He, van de Vijver, et al., 2017; Kyllonen & Bertling, 2014). Therefore, we know that the scale usage correction procedures reviewed above can mitigate some data comparability issues, but it remains unclear which psychometric properties are changed and the extent to which these corrections could differentially affect individual- and country-level responses. The present study aims to explore this gap in knowledge by comparing and contrasting different scale usage correction procedures using the same data.
The Effect of Scale Usage Correction Procedures on Changes in the Psychological Meaning and Metrics of Target Scales
A first step in evaluating the effectiveness of different scale usage correction procedures is to examine the extent to which they change the psychological meaning and metrics of target scales. It is usually assumed that correction procedures adjust individual responses to a common, comparable scale, without changing the psychological meaning of the target construct. However, some correction procedures have an influence on interitem correlations. For example, anchoring vignette corrections will increase interitem correlations and hence, the internal consistency of the scale (e.g., von Davier et al., 2018), whereas ipsatization introduces a slight negative correlation between the items. Therefore, it seems that for statistical reasons, it is unrealistic to expect that different scale usage correction procedures will yield entirely comparable results. We examine whether and how these four correction procedures alter the psychological meaning and metrics of the target scales in a multilevel equivalence testing framework. This is done by treating the four sets of corrected scores together with the raw scores for each respondent such that they can be seen as pseudo-repeated measures at the individual level, resulting in five sets of responses that are related to each other.
Consistency of Scale Usage Corrections in their Impact on Scores
Second, it is important to explore the potential drawbacks of using these corrections. Are we removing the same source of measurement bias with each correction? These correction procedures are used to alleviate incomparability due to scale usage differences, but they may also impact the scores differently. We examine the global changes in structural relations and mean comparisons (by linking corrected and uncorrected scores with external measures other than Likert-scale self-reports) and explore the similarities and differences in what is being corrected for.
Characteristics Associated with Individuals’ Susceptibility to Corrections
Thirdly, are score profiles of various respondents similarly susceptible to different correction procedures? Research on the correlates of these correction procedures indicates otherwise. With response styles, it has been found that response amplification (e.g., ERS) is positively related to age and negatively related to education and cognitive performance (e.g., van Vaerenbergh & Thomas, 2013). Immigrants and nonimmigrants also differ in their scale usage preferences, which may be moderated by cultural values and the acculturation processes in the receiving country (e.g., Morren et al., 2013). Thus, it is reasonable to assume larger ERS correction effects for older respondents with a lower educational level and an immigrant background. Extending the evaluation of one procedure to multiple procedures, we can investigate the similarities and differences in correction susceptibility for individuals and cultural groups. In other words, for certain groups of respondents, applying different procedures may change (the ranking of) their scale scores dramatically, whereas for some other groups of respondents, corrected and uncorrected scores are relatively stable, indicating higher reliability of their reporting. Taking into consideration of the four procedures, we aim to predict the extent to which people might vary in their susceptibility to corrections based on individual and cultural characteristics.
Research Questions
In this study, we apply scale usage corrections on self-reported scales in the 2012 PISA data based on design procedures including anchoring vignettes and a measure of overclaiming, and procedures focusing on data transformations prior to analysis (ipsatization and post hoc ERS correction), with the aims to simultaneously model: (1) what psychometric property changes are introduced with various correction procedures, (2) any inconsistencies in correction effects across procedures, and (3) individuals’ susceptibility to correction effects. This exploration with an integration of procedures where raw and corrected scores are nested within each individual, and individuals are nested within countries, can reveal fundamental changes in corrections (despite the fact that these corrections may not result in scalar invariance of data in cross-cultural comparisons), and characteristics of subgroups that are more or less affected by the corrections. It is expected to add incremental value to our understanding on when and for whom scale usage preferences can make a difference on the measurement, structural relations and mean comparisons of target constructs.
Method
Data Source
We based our analyses on PISA data collected in 2012. PISA assessed the competencies of 15-year-olds in reading, mathematics, and science (with a focus on mathematics) in 64 countries and economies in 2012 (OECD, 2013). Students were recruited through a stratified sampling procedure to represent the schools and the 15-year-old student population of each country. Each student took a subset of a cognitive test that took 2 hr to complete, and a context questionnaire that took another 30 min following the cognitive test. In this cycle of assessment, two adapted designs (i.e., anchoring vignettes and overclaiming) were featured in the student context questionnaire, 1 and data were available to perform the two post hoc corrections of item responses. All data, codebook, and manuals for the PISA are available for public use on the OECD website (http://www.oecd.org/pisa/pisaproducts/). The data package for this study (including raw and refined dataset, syntax, and output) was deposited on the Open Science Framework (https://osf.io/g36uh/).
Measures
Target Scales
We focused our analyses on the two student self-report scales that included anchoring vignettes in the PISA data. Teacher Support (TS) was measured using four items (e.g., “The teacher helps students with their learning”), with response options on a Likert scale ranging from 1 (strongly agree) to 4 (strongly disagree). Values of coefficient alpha for this scale ranged from 0.63 (Liechtenstein) to 0.88 (Chinese Taipei) with a mean of 0.76 and median of 0.78 across the 64 countries. Classroom Management (CM) was measured using four-items on the same Likert scale (e.g., “My teacher gets students to listen to him or her”), and values of coefficient alpha ranged from 0.31 (Japan) to 0.79 (France) with a mean of 0.68 and a median of 0.71. The CM scale appears to work poorly in Asian and some Middle Eastern countries, as it had a coefficient alpha below 0.60 in Japan, Qatar, Indonesia, Thailand, Albania, Jordan, Russian Federation, Malaysia and Turkey.
Anchoring Vignettes
Two sets of vignette questions were asked, targeting TS and CM, respectively. Each set had three vignette questions on high, medium, and low levels of traits, respectively; the vignettes were asked immediately prior to the self-assessment questions of the two scales. Students rated the vignettes on the same 4-point Likert scale. Following the nonparametric scoring approach with the anchors package in R (Wand & King, 2007), we rescaled each of the self-assessment item responses on the target scales on the basis of responses of the three ordered vignette questions designed for the specific construct. With ratings on three vignettes, the rescaled scores were on a 7-point scale. The rescaled self-assessment score became 1 if the self-assessment was lower than the rating on the vignette of low trait, 2 if it was equal to the low-trait level vignette, 3 if it was in between the low- and medium-trait level vignettes, 4 if it was equal to the medium-trait level, and so on. In cases of tied or inconsistently ordered vignette responses (e.g., respondents did not distinguish the different trait levels in vignettes, as indicated by providing the same rating on different vignettes or reversing the ratings), the rescaled item responses could take a vector of possible values instead of one scalar value, and the highest possible rating was taken as a proxy of the anchored scores, as it has been suggested to enhance comparability and validity in comparison to the lowest possible rating (Kyllonen & Bertling, 2014).
Overclaiming
An assessment of overclaiming was embedded in the Familiarity with Math scale in the student contextual questionnaire. Specifically, three foil items (items referring to concepts that do not exist) were administered along with items on the familiarity with math concepts in 63 countries (note that these items were not administered in Norway). The response options ranged from 1 (never heard of it) to 5 (know it well, understand the concept), and the internal consistency of the 3-item scale ranged from 0.46 to 0.80, with a mean of 0.64 and a median of 0.63. The mean of ratings from the three items was taken as an overclaiming score. 2 Each target item score was regressed on the overclaiming measure, and the residuals from the regression, representing an item score after accounting for overclaiming, were saved as the overclaiming corrected scores.
Ipsatization
The student questionnaire includes 40 items that have response options range from 1 (strongly agree) to 4 (strongly disagree). These comprise the two target scales of TS and CM, as well as Math Anxiety, Math Concept, Student-Teacher Relationship, Sense of Belonging in School, and Attitude towards School. Following the within-subject standardization procedure in the Schwartz Value Survey (Schwartz, 2009), we first reverse-coded the negatively worded items, computed the mean across all available items for each individual, and then subtracted this mean from the raw item responses for the two target scales to produce the ipsatized scores. 3
ERS Correction
It is acknowledged that specific response styles such as ARS, ERS, and MRS all can impact on score comparability. In this study, ERS was chosen as the response style indicator over ARS and MRS, given that it shows higher consistency and reliability than the others. Items in the student questionnaire that used the same response format, but excluding the two target scales, were recoded for the presence of ERS (i.e., an original response of 1 or 4 on a 4-point scale was recoded with a value of 1) and for the absence of ERS (i.e., an original response of 2 or 3 on a 4-point scale was recoded with a value of 0). The recoded ERS scale with 32 items had values of coefficient alpha ranging from 0.84 (Germany) to .93 (Kazakhstan) with a mean of .89. All target item scores were regressed on the ERS measure, and the residuals were saved as the ERS corrected scores.
Individual and cultural characteristics
Other information was extracted from the same questionnaire, including gender (1 [female] and 0 [male]), grade (most 15-years old student respondents are between Grade 7 and 9, here grade was measured as one’s school level in comparison to the national modal school level) whether one has ever repeated a grade (1 [yes] and 0 [no]), socioeconomic status (a composite indicator of economic, social and cultural status), language use at home (1 [same as test language] and 0 [other language]), and immigration status (1 [immigrant] and 0 [native]). Students’ math achievement scores from the cognitive assessment in PISA were analyzed. Math achievement scores were calibrated and scaled using an Item Response Theory based scaling approach, and the scores were computed using five plausible values. Plausible values are a selection of likely proficiencies that are randomly drawn from the marginal posterior of the latent distribution for each student. Thus, analyses involving math achievement scores need to be performed with each of these plausible values and combined for unbiased estimates (Rutkowski et al., 2010). At the country-level, membership in the Organization for Economic Co-operation and Development is indicated (1 [is a member] and 0 [is not a member]).
Analysis Strategy
In addition to raw scores, after applying the four scale usage correction procedures, each respondent had five sets of scores on the target scales. We then conducted the analyses in three steps. In Step 1, we performed a three-level multilevel confirmatory factor analysis with different sets of corrected and uncorrected item scores (Level 1) nested within each respondent (Level 2) nested within countries (Level 3) in Mplus 7.3 (Muthen & Muthen, 1998–2012). Three models were specified to test configural, metric, and scalar invariance across levels of analysis. In comparison to multigroup confirmatory factor analysis (with respondents nested in countries), multilevel confirmatory factor analysis has several advantages. First, it accounts for the dependency of different sets of scores within individuals and within countries, and it can estimate the models based on these sets of scores simultaneously. In contrast, multigroup confirmatory factor analysis needs to be performed on each set of scores separately, which may be problematic as it ignores their dependency. Then, the relative merits of corrections can be judged by eyeballing model fit indexes between the corrected and uncorrected scores. Secondly, multilevel confirmatory factor analysis provides empirical evidence as to whether the psychological meaning and metrics of the construct remain unchanged across sets of scores (i.e., with acceptable multilevel configural and metric invariance), which is not possible in multigroup confirmatory factor analysis. For example, in a multigroup model, obtaining better model fit in a model with the corrected scores vs. a model with the uncorrected scores does not necessarily guarantee improved measurement properties regarding the target construct but it could also be due to changed meaning of the construct.
We made use of Full Information Maximum Likelihood (FIML) estimation to account for missing data, and used senate weights (i.e., rescaled final student weights so that the population of each country equaled 1,000) to ensure that each country contributed equally in the model. We first estimated a configural model, in which the latent factors of TS and CM were indicated by their respective four items at the within-individual, between-individual, and between-culture level). Good model fit for the configural model indicated that the constructs have the same meaning at the three levels of analysis. We then estimated a metric model by adding equality constraints on the factor loadings across the three levels, which assessed whether the measurement units are comparable across levels of analysis. Lastly, we tested a scalar model by fixing the error variances of all items at the individual and cultural level to be zero, which amounted to intercept equivalence across the three levels of analysis (Jak et al., 2014). Good model fit for the scalar model would point to strong comparability across levels of analysis and for cross-cultural comparisons. Model fit was assessed from the comparative fit indexes including the Comparative Fit Index (CFI; > 0.90 acceptable) and the Root Mean Square Error of Approximation (RMSEA; < 0.08 acceptable), and the absolute fit index Standardized Root Mean Square Residual (SRMR; < 0.08 acceptable). The acceptance of a more restricted model was based on the changes of CFI and RMSEA values: ΔCFI to 0.02 and ΔRMSEA to 0.03 from configural to metric models, and to 0.01 from metric to scalar models for both ΔCFI and ΔRMSEA (Cheung & Rensvold, 2002; Rutkowski & Svetina, 2014). In doing so, we could assess whether these corrections change the psychological meaning of the target scales across aggregation levels.
In Step 2, we estimated factor scores of TS and CM using the multilevel CFA model. This model produced five sets of factor scores (based on the raw scores and the four correction procedures, respectively) at the individual level, and one set of factor scores at the country level. Then, we compared the correlation of raw and corrected scores with math achievement at the individual level. Because the data were standardized within countries in the first analysis, the estimated country level factor scores or aggregated scores from the individual-level estimates were not suitable for country comparisons.” Results from the first two steps provide answers concerning global changes and consistency produced by the corrections.
In Step 3, the variations in the Level 1 factor scores (i.e., the five sets of factor scores for each individual) were modelled in a two-level multilevel regression analysis with predictors at both the individual and country levels in Mplus 7.3. Here, we used multiple imputation (in addition to FIML estimation) to account for the missingness found at the item and individual levels. This analysis addresses individual’s susceptibility to corrections. We report the findings in each analysis step below.
Results
The Effect of Scale Usage Correction Procedures on Changes in the Psychological Meaning and Metrics of Target Scales
In the three-level multilevel confirmatory factor analysis with the TS and CM items measuring the two correlated constructs, we specified models of configural, metric, and scalar invariance across intraindividual, individual, and country levels. To retain the within-country variations and facilitate modelling across countries, these five sets of scores were standardized to z-scores per country in the multilevel model. This model involved 310,737 respondents with 1,335,732 within-person observations across 64 countries. In the configural model, the intraclass correlations for items at Level 2 (individual level) ranged from 0.73 to 0.85, and at Level 3 (country level) from 0.02 to 0.06. Table 1 presents the model fit of different models. Results indicated that the configural model fit the data well, indicating equivalent constructs measured across the three levels, which suggests that the corrections did not change the psychological meaning of target constructs. The drop in the CFI values from the configural to the metric model was above 0.02, indicating a less well-fitting model when all loadings were constrained. Modification indices suggested the largest cross-level variation on one item loading (i.e., “My teacher keeps class orderly” from the CM scale), thus the constraint on this item loading was freed. This partial metric invariance model showed an acceptable fit based on CFI and RMSEA, although the SRMR showed considerable difference between the observed correlation and the predicted correlation at the country level. This speaks largely to comparable metrics across levels for the measurement of target constructs. Unsurprisingly, the scalar invariance model did not fit well, as these correction procedures were meant to modify item intercepts and the latent means of target constructs. These findings indicate that the various score corrections could have an impact on score comparability, but not on construct comparability, suggesting that the psychological meaning of scores remained largely invariant across the different scale usage correction procedures. 4
Model Fit in the Multilevel Confirmatory Factor Analysis.
Note. CFI = comparative fit index; RMSEA = root mean square error of approximation; SRMR = standardized root mean square residual.
p < .01.
Consistency of Scale Usage Corrections in their Impact on Scores
To delineate the specific changes brought about by each of these correction procedures on the structural relations and global mean patterning, we made use of the Level 1 factor scores estimated in the first step of our analysis. Because the partial metric invariance model from the multilevel equivalence testing (Step 1) was the best fitting model for the data, the factor scores of TS and CM were estimated using this model. This yielded five sets of factor scores (i.e., raw, anchored, overclaim corrected, ipsatized, and ERS corrected) at the individual level for each construct. The correlations between the factor scores with math achievement at the individual level are presented in Table 2. For both TS and CM, the correlations with achievement for raw scores were negative. Anchored scores markedly changed this pattern as the correlations became positive. For ipsatization and ERS correction, the negative correlations shown in the raw data were. Correction for overclaiming produced the smallest changes, such the correlations were barely changed from these of the raw scores.
Individual-Level Correlations with Math Achievement.
Note. Sets of scores for teacher support and classroom management were estimated using the partial metric invariant multilevel CFA.
p < .01.
Characteristics Associated with Susceptibility to Corrections
Using the factor scores estimated from the partial metric invariance multilevel model described above, the standard deviations across the five factor scores at Level 1 were specified as the dependent variables in the multilevel regression analysis (using data from 304,330 respondents nested in 64 countries). In the null model, we found very little variation at the country level with intraclass correlations of 0.03 for the variability in TS and 0.02 for the variability in CM, respectively. The two dependent variables were then predicted by gender, grade, ESCS, immigration background, math achievement, and OECD membership. Because math achievement was measured with five plausible values, we used multiple imputations to combine results from each and every set of plausible values. The standardized regression coefficients are presented in Table 3. For both variability in TS and CM, males, individuals in lower grades, and those who repeated a grade tended to be more susceptible to correction effects than females, individuals in higher grades, and those who did not repeat a grade, respectively. Additionally, higher math achievement was the strongest predictor of low susceptible to correction effects. There was no substantial difference in variability in TS nor CM between OECD and non-OECD countries.
Standardized Solution in the Multilevel Analysis on the Variability of Correction Effect for Teacher Support (TS) and Classroom Management (CM).
Note. *p < .05. **p < .01.
Discussion
Data comparability is a fundamental requirement for conducting valid cross-cultural comparisons. This study aimed to shed light on the different design and data analytic procedures meant to correct for scale usage differences and the implications of these procedures. We used self-reported data from students in 64 countries collected in the 2012 PISA assessment to compare and contrast the effectiveness of four different scale usage correction procedures: anchoring vignettes, overclaiming, ipsatization, and response style correction. In general, we found that these different procedures did not change the psychometric meaning of the target constructs. We also found that the structural relations and mean levels varied by correction procedure, and that there were both similarities and differences in the effects. Moreover, we found that respondent scores were not equally affected by different correction procedures, and specifically, scores from females, students without a grade repetition and students with higher math achievement were least affected by these correction procedures. We discuss the implications of these findings below, and provide recommendations for addressing the cross-cultural comparability of Likert-scale data.
Scale Usage Correction Procedures Should Not Change the Psychological Meaning and Metrics of Target Scales
It is important to evaluate whether corrections change the psychological meaning of the target constructs before evaluating the effectiveness of each correction procedure. The support of configural and partial metric invariance in a three-level multilevel confirmatory factor analysis (with raw and corrected scores nested within students and within countries) suggests that the same psychological meaning of the constructs and metrics are achieved with and without corrections.
Obtaining evidence for unchanged psychological meaning of the target construct is fundamental for a few reasons. First, this requirement should be met before evaluating the extent to which these corrections indeed improve comparability by evaluating the model fit of the scalar invariance using different sets of scores separately. If this is the case, the model fit of scalar invariance model should also become better with corrected scores than raw scores for the intended construct (Marksteiner et al., 2019). The findings we obtained from the multilevel equivalence testing provide more support for the conclusion that procedures such as anchoring vignettes can indeed enhance data comparability (e.g., He, Buchholz, et al., 2017; Marksteiner et al., 2019). Second, evidence for unchanged psychological meaning of the target construct is needed to further evaluate the validity (e.g., external validity) of corrected scores for the target constructs, described below. In the current study, we used a multilevel analysis approach to test for changes in psychological meaning and metrics using raw and multiple corrected scores. By treating individual scores as pseudo-repeated measures within individuals, we were able to integrate information from different sets of scores. Additionally, using a multilevel confirmatory factor analysis with different constraints served the aim of unpacking the equivalence of the construct being measured (i.e., multilevel configural invariance), the equivalence of strengths of the associations between items and the construct (i.e., multilevel metric invariance), and the equivalence of the item intercepts (i.e., multilevel scalar invariance). The lack of support for the highest level of invariance is not necessarily disappointing, as it points to significant changes in intercepts and latent mean of the construct produced by correction procedures, which perhaps afford higher comparability and validity. This modelling approach properly handles the structure of the data by accounting for dependencies across sets of individual scores, estimates models in a parsimonious way (by integrating different sets of scores within the same model, and provides empirical evidence on the extent to which the psychological meaning across sets of scores is invariant. Given that our findings are based on four correction procedures in the 2012 PISA data, we recommend that future research on the effectiveness of any correction procedures with other data sources first empirically check that the psychological meaning of the target construct remains unchanged specifically using multilevel confirmatory factor analysis.
Scale Usage Corrections Vary in their Impact on Scores
The similarities and differences of each correction procedure’s impact on structural relations can be inferred from the correlations of raw and corrected scores with math achievement (See Table 2). Our results were remarkably similar for the two target constructs with regard to the effect of each correction procedure on structural relations. Anchoring vignettes exhibited the most substantive change in correlations with math achievement, followed by ipsatization, extreme response style correction, and then overclaiming. The different sizes of changes in correlations might be due to different measurement biases being targeted. Alternatively, they may be due to the same measurement bias being tackled to different degrees, while not excluding the possibility that these procedures may bring in other sources of measurement bias.
Given that we had no theoretically based expectations with regard to the true correlations between the constructs of teacher support and classroom management with student math achievement in large-scale cross-sectional data, it is unclear as to which set of corrected scores provides the greatest validity, and we can only extrapolate based on the observed correlational patterning in the current study. Previous studies that have relied on interventions and longitudinal designs to target a certain monocultural group have identified significant positive effects of teacher support and classroom management on student achievement (e.g., Aldrup et al., 2018; Elias & Haynes, 2008; Freiberg et al., 1995), whereas studies that have used cross-sectional designs have shown weaker or nonsignificant associations between teacher support and achievement (e.g., Chen, 2005; Yıldırım, 2012).
Yet, it is possible for a negative association between teacher support and math achievement to occur at one time point, with more teacher support being directed toward students with low academic performance than those with high academic performance, because the former group may need more teacher support than the latter group. Similarly, a possible explanation for the negative cross-sectional correlation between classroom management and math achievement may be that there is a lack of student-centered teaching and cognitive activation in the classroom, which can be detrimental for student achievement. It is also worth noting that the true association would not be known using simulation-based methods. In general, the very limited changes resulting from the correction of overclaiming seem to suggest that overclaiming, embedded in the knowledge domain to capture self-enhancement bias, does not fully capture individual scale usage differences in self-reported attitudes and opinions for Likert-scale format data (Müller & Moshagen, 2018), or that the individual scale usage difference is minimal. 5
In contrast, ERS correction and ipsatization attenuated the negative correlations of math achievement and both constructs, suggesting that they are better at capturing and controlling for response amplification or moderation among individuals. Moreover, the marked changes that resulted from correcting for anchoring vignettes in the current study is in line with previous studies (e.g., He, van de Vijver, et al., 2017; Kyllonen & Bertling, 2014; Vonkova, Zamarro, et al., 2018), which also speaks to the robustness of different operationalizations and modelling approaches in such corrections. Anchoring vignettes appear to be a promising approach to enhance data comparability, given that the psychological meaning of target constructs was not changed in this study, and that anchoring vignettes rescaled scores tend to show higher measurement invariance in multigroup confirmatory factor analysis of PISA data (e.g., Marksteiner et al., 2019). However, further evidence on the tenability of the working assumption is needed to fully utilize this approach with confidence.
Individual Characteristics are More Strongly Associated with Susceptibility to Corrections than Country Characteristics
Modelling variability with standard deviations in data from a large-scale assessment can bring in new perspectives concerning heterogeneity within groups. We modeled the variability of factor scores based on raw scores and different corrections, and found that, irrespective of cultural characteristics, females, those without grade repetition and those with higher math achievement are less susceptible to corrections in the PISA data. The low ICC at the country level suggests that variability in correction effects does not strongly depend on national contexts and cultural characteristics (e.g., affluence level, cultural values and orientations), but rather that individual characteristics matter to a greater extent. Thus, it is unlikely that lack of comparability in large-scale assessment is mainly attributable to a subgroup of countries. Therefore, although one advocated remedy for lack of comparability is partial equivalence (i.e., freeing the equivalence constraints for certain groups), our findings suggest that partial equivalence may not solve the issue. Instead, clustering countries and their measurement in a simultaneous mixture factor analysis may be more appropriate in such large-scale assessment contexts (De Roover et al., 2017).
Furthermore, the large variation at the individual level and the consistent patterning of variability in correction impacts on subgroups of individuals gives rise to new research directions. For these characteristics showing limited susceptibility to various corrections, with their rather reliable measurement—with or without correction—a promising future direction is to simply compare subgroups of respondents with these characteristics across countries. If measurement invariance testing of target constructs with this subgroup (e.g., females without grade repetition) across countries fails to find support of scalar invariance, this would strongly indicate that there are not only quantitative differences on these target constructs, but also qualitative differences, making mean comparisons based on Likert-scale self-reports of these constructs difficult. In this case, it would be important to explore the qualitative differences in these measures via interviews, cognitive labs, focus groups, and so on, as they may reflect meaningful individual and cultural characteristics that have yet to be adequately quantified (e.g., Benítez & Padilla, 2014; Zumbo, 2007).
Limitations
This study is not without limitation. First, we have focused on a single empirical dataset. The PISA 2012 is the only large-scale assessment that incorporates the two design-based procedures of anchoring vignettes and overclaiming, and thus the generalizability of our conclusions is tentative. Future research should examine the extent to which our findings replicate in other datasets with dedicated and well-matched correction procedures and in simulations. Second, one target scale (i.e., Classroom Management) showed rather different internal consistency across countries, indicating that it is not as reliable in some Asian countries. This finding points to a need to further examine how the measurement properties of scores resulting from scales administered in different countries may interact with the impact of corrections.
Conclusions
In the current study, we examined four common procedures to correct for scale usage differences in Likert-scale data from the 2012 PISA. We found that the procedures do not change the psychological meaning of the target construct, but that they do change the origin of metrics. Furthermore, these correction methods share commonalities and differences in their effects on structural relations. Corrections based on anchoring vignettes showed the largest changes in correlations with our criterion, followed by extreme response style correction and ipsatization, with overclaiming showing the least amount of change. Individual susceptibility to these corrections showed greater associations with individual-level characteristics than with country-level characteristics, such that there were correction effects were largest for males, students in lower grades, and students who repeated a grade.
The insights obtained from the current study regarding what and for whom scores are most changed by scale usage corrections, we recommend that researchers check if corrections change the psychological meaning of the target constructs. Secondly, we recommend the use of advanced statistical models such as multilevel confirmatory factor analysis to better delineate the similarities and differences of the impact of these corrections. Taken together, our findings suggest that there is no correction procedure that is clearly better. Anchoring vignettes show much potential to enhance data comparability and validity. However it is important to have matched vignettes and a target construct which can impose high cognitive demand resulting from administering double or triple the amount of questions, and the untenable assumptions of anchoring vignettes remain problematic to ascertain its effectiveness. Using overclaiming ratings in the knowledge domain for attitudinal expressions showed little correction effects, but the effects may be enhanced by better aligning the response options and content domains. Furthermore, the two post hoc statistical correction methods are convenient if, in the case of ipsatization, there are available data to compute a meaningful individual mean across items, and in the case of extreme response style, response styles can be extracted. Therefore, we believe these procedures should be seen as complementary, and not at odds with each other. We recommend that researchers invest in future collection of data that can allow for these procedures to be used effectively. By using well-matched measures such as incorporating the design features necessary to examine scale usage correction in the target construct, the convergence and robustness of their correction effects can contribute to better cross-cultural comparisons. Lastly, we advise researchers to focus on both the qualitative and quantitative differences in targeted constructs with subgroups of respondents across cultures.
Footnotes
Authors’ Note
Fons van de Vijver was involved in the project from the beginning until his sad passing in June 2019.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by the Marie Sklodowska-Curie Individual Fellowship European program [Grant Number 748788]
