Abstract
The issue of race within the context of psychological assessment is important, but often overlooked. Many self-report measures of psychopathology have been developed and validated using primarily White samples. Research regarding the Penn State Worry Questionnaire (PSWQ) and race has produced mixed results, which in turn may present challenges when comparing scores across racial groups. The current article sought to investigate the measurement invariance of the PSWQ and PSWQ-A (an abbreviated version) across four racial groups (White, Black, Asian, and Hispanic) in a sample of 2,489 undergraduate students. Confirmatory factor analysis of a one-factor structure illustrated poor fit across all racial groups for the full-length PSWQ. Two-factor and one-factor with method effects models of the full-length PSWQ each improved on the previous model fit, although the one-factor method effects model was limited by nonsalient factor loadings. Additionally, a separate confirmatory factor analysis indicated good fit for the PSWQ-A. Further analysis of the PSWQ-A suggested measurement invariance across all racial groups, as well as configural, metric, and scalar invariance. These findings advance the literature on the relationship between worry and race, suggesting that direct comparisons on the PSWQ-A between racial groups is appropriate.
The study of potential racial differences in psychological assessment is an important yet underrepresented research area in the development and validation of symptom questionnaires (Helms et al., 2005; Watkins, 2012). 1 Perhaps as a result, research suggests that some measures perform inconsistently across racial groups (e.g., Thomas et al., 2000; Wheaton et al., 2010).
There are many reasons why achieving accurate assessment across racial groups can be difficult. According to Roysircar (2005), this pursuit can be complicated by issues surrounding diagnostic guidelines, academic and research bias, and sparse and ambiguous research surrounding the relationships among race, assessment, and symptomatology. Helms et al. (2005) added that racial groups often are misused in psychological research as causal constructs that can have a direct impact on a dependent variable. In reality, the function of race within most psychological research is unclear, with some suggesting that the use of race is too dependent on other demographic variables to stand alone (e.g., Steinberg & Fletcher, 1998).
Still other research suggests that racial group differences exist in anxiety and related symptoms. Hunter and Schmidt (2010) proposed a model wherein beliefs and attitudes tied to sociocultural groups have effects on behavioral outcomes and clinical assessment. Their model posits that specific sociocultural factors experienced uniquely by different racial groups interact with social stereotypes and thus have an impact on presentation and interpretation of a variety of psychopathology symptoms. Racism, stigma of mental illness, and salience of physical illness all may be experienced distinctively by separate racial groups, who may differentially view certain behaviors as problematic (Hunter & Schmidt, 2010). Each of these factors—individually or in combination—may lead to overreporting or underreporting certain symptoms. Thus, Hunter and Schmidt opined that consideration of racial differences is critical in the context of psychopathology assessment to identify where these differences may exist.
Another consideration is that reliance on White and middle-class populations for exploratory, validation, and outcome research (Helms et al., 2005) clouds the research on these issues to date. For example, Watkins (2012) found that across 104 assessment and treatment outcome studies, 75% did not report the race of participants. Of the 25% that did, 75% of the participants were White. When only three studies that provided the majority of the non-White participants were excluded, more than 90% of remaining participants were White. This pattern of findings shows a clear overrepresentation of White individuals in this research. According to the U.S. Census Bureau (2018), 60% of individuals in the United States self-identify as White and the percentage of the U.S. population who identify as non-Hispanic White is projected to decrease to approximately 50% by 2045 (Vespa et al., 2020). Findings such as those offered by Watkins (2012), in conjunction with the model proposed by Hunter and Schmidt (2010), suggest the need to investigate the performance of extant instruments that were developed using predominantly White samples, in more representative samples. The current study sought to address the second action point—specifically, to examine the measurement invariance of the Penn State Worry Questionnaire (PSWQ; Meyer et al., 1990) in four racial groups.
Measurement of Worry Severity (PSWQ)
The PSWQ (Meyer et al., 1990) is perhaps the most widely used measure of worry in clinical and research settings. It contains 16 items, including five reverse-coded items. The structure of the item scores initially was examined in a college student sample (n = 337), with the racial identification of participants not reported. Meyer et al. asserted that the PSWQ was comprises one general factor, as they conceptualized worry as measured by the PSWQ to be unidimensional. Some data have not supported a unidimensional model fit, including data from studies that have examined the factor structure in samples of White and African American students (Carter et al., 2005), Korean students (Lim et al., 2008), and Mexican students and community members (Padros-Blazquez et al., 2018).
To address the poor fit of the original factor structure, subsequent research has suggested potential remedies. One such remedy is use of a two-factor solution—with one factor representing the five reverse-coded items—which has provided improved model fit (e.g., Fresco et al., 2002). Although some researchers have characterized two substantive factors (i.e., “Worry Engagement” and “Absence of Worry”), the utility of the second factor has been questioned conceptually and practically. More precisely, the two factors are highly intercorrelated (φ = .87; Brown, 2003) and there is a lack of conceptual reasoning for an “Absence of Worry” factor in relation to the original conceptualization of the construct assessed by the PSWQ (Meyer et al., 1990). Alternatively, a second factor may emerge as a result of a method effect. A second remedy to addressing the poor structural validity of the PSWQ is to account for this method effect within a one-factor model (Brown, 2003; Hazlett-Stevens et al., 2004).
A third approach to improving the structural validity of the PSWQ is to drop the reverse-coded items and develop a short form. Hopko et al. (2003) developed the PSWQ-A—an abbreviated, eight-item version that removed the five reverse-scored items and three other positive-scored items. The PSWQ-A showed a unidimensional structure and retained strong psychometric properties comparable to the full-length PSWQ (Hopko et al., 2003). For example, the PSWQ-A was highly correlated with the original PSWQ (r = .92), showed strong internal consistency (α = .97), showed moderate test–retest reliability (r = .63), and had comparable convergent and discriminant validity estimates to its full-length counterpart. This pattern of findings has been replicated (Crittendon & Hopko, 2006; Kertz, Lee, et al., 2014). The PSWQ-A total score has been found to adequately discriminate between clinical severity status groups (e.g., anxiety diagnostic status group vs. nonclinical control group; Wuthrich et al., 2014). Although similar sensitivity and specificity values were found across the two versions of the PSWQ, Wuthrich et al.’s pattern of results suggest some diagnostic accuracy may be lost when using the PSWQ-A total score relative to the full-length PSWQ total score. It is important to underscore that, as reviewed, the use of a summed total score from the full-length PSWQ item pool is not uniformly supported by existing studies that have cast doubt on the tenability of a one-factor model. Alternatively, as reviewed, structural analyses support deriving a PSWQ-A total score from the respective item scores. Overall, the extant literature supports the PSWQ-A as improving on the structural validity of the PSWQ item pool and that the PSWQ-A has the potential to be a viable alternative to the full-length PSWQ in many settings.
Experience of Worry Across Racial Groups
Some studies have explored differences in worry severity and the specific content of worry across racial groups. For example, in a sample of 502 undergraduate students, Scott et al. (2002) observed differences both broadly and across specific domains. Compared with White and Asian American participants, African American participants exhibited lower levels of overall worry severity (p < .001). In terms of content areas, African Americans reported less worry in relationship stability, self-confidence, work competence, and the future (all ps < .001), whereas Asian American participants generally reported high worry on all domains. All groups in the study reported similar levels of worry with regard to finances. This pattern of findings illustrates racial differences in both the severity and content of worry. There may be a number of influencing factors, including parental expectations: Saw et al. (2013) found that Asian Americans’ perceptions of parental expectations explained significant (p < .05) differences with their White peers on worry related to academic and family concerns.
PSWQ and Race
The PSWQ has been examined with regard to potential differences in worry severity across racial groups. This research has yielded mixed findings. Gillis et al. (1995) administered the PSWQ to a quota sample (n = 267; 85% White) selected to correspond with the demographic characteristics of U.S. adults as measured by the national census (i.e., sex, race, income, and age). No significant (p > .05) mean-level difference was found between White and Black groups on the PSWQ total score. Brenes et al. (2008) examined PSWQ-A responses from 1,111 White and Black individuals ranging from 18 to 94 years of age recruited from hospitals. Similarly, there were no significant (p > .05) mean-level differences observed. In contrast, McLaughlin et al. (2007) examined main effects of race on the PSWQ in a diverse (86.8% non-White) sample of middle school children in the Northeast United States (n = 1,065). Significant main effects of race were found for the PSWQ (p < .05, d = 0.20), with Hispanic individuals scoring the highest across White, Black, and Hispanic racial groups. Whereas McLaughlin et al. used a sample with different characteristics (e.g., age) than Gillis et al. and Brenes et al., the results offer a possible contrast in addition to the inclusion of a third ethnoracial group and therefore contribute to lack of clarity as to whether there are racial differences in worry severity as assessed by the PSWQ.
Critical in the examination of group differences in symptom severity is whether the item scores of the measure used to examine those differences function similarly across groups of interest (Brown, 2015). As noted, the PSWQ is used widely to assess worry severity. To the degree to which invariance is supported across the item scores of the PSWQ, mean-level differences can be considered as reflective of true differences in worry severity rather than as an artifact of differential item scores across groups.
To date, two published studies have examined the structure of the PSWQ or the PSWQ-A across racial groups. Carter et al. (2005) compared the PSWQ factor structure in Black (n = 181) and White (n = 180) college students. An exploratory factor analysis revealed that a two-factor solution was most appropriate among White individuals, with one factor consisting of all reverse-coded items. This finding is consistent with the model supported by the findings of Fresco et al. (2002). Additionally, an exploratory factor analysis found that a three-factor solution best fit the data for Black individuals; in that solution, the reverse-coded items split into two factors. However, the third factor contained only two items and therefore represents a potentially underdetermined factor (Brown, 2015). Acknowledging such a concern, the authors reasoned that a two-factor solution—separating reverse-coded items from nonreverse-coded items consistent with Fresco et al. (2002)—was most practical in both groups. DeLapp et al. (2016) examined the PSWQ-A in Black (n = 100) and White (n = 121) college students. A confirmatory factor analysis (CFA) supported the same one-factor solution for both racial groups. Partial invariance was found for the PSWQ-A item scores in that four items showed higher intercepts in the sample of White individuals compared with the sample of Black individuals. This pattern of findings suggests a need to carefully consider what intercept differences across racial groups indicate, as differential item scores may lead to bias in direct comparison of scores across specific racial groups.
The Current Study
There is a need to extend Carter et al.’s (2005) and DeLapp et al.’s (2016) findings as to the structure of item scores and measurement invariance of the PSWQ and PSWQ-A across racial groups. Only DeLapp et al. formally considered measurement invariance and that study examined only the PSWQ-A. Moreover, both studies focused on only two racial groups, despite a clear need to examine worry across other racial groups (e.g., Lim et al., 2008; Padros-Blazquez et al., 2018). The purpose of the current study was to investigate the invariance of the PSWQ and PSWQ-A across four racial groups by examining prominent PSWQ measurement models identified in the extant literature. More precisely, we tested the original full-length one-factor PSWQ model (Meyer et al., 1990), the two-factor full-length PSWQ model (Fresco et al., 2002), the one-factor method effect full-length PSWQ model (Brown, 2003), and the one-factor PSWQ-A model (Hopko et al., 2003). 2 If none of these four measurement models received support in baseline examinations of model fit across groups that traditionally precede measurement invariance analyses (Brown, 2015), exploratory analyses (e.g., removal of specific items) then were considered to improve model fit. We expected the one-factor PSWQ-A model to receive support in the baseline analyses across groups given research findings regarding the structural validity of those item scores (e.g., Crittendon & Hopko, 2006; DeLapp et al., 2016; Hopko et al., 2003; Kertz, Lee, et al., 2014). No predictions were made as to the expected support for the full-length PSWQ measurement models across the groups given the inconsistent findings of existing studies examining those models (e.g., Brown, 2003; Carter et al., 2005; Fresco et al., 2002).
A baseline measurement model with support across each of the examined groups was considered necessary before then examining the measurement invariance of that model (Brown, 2015). Measurement invariance analyses, in conjunction with analysis of latent mean differences across racial groups on a given measure if invariance is sufficiently supported, can offer insight into whether score differences and reported experiences across racial groups are a function of the measure itself or better attributed to sociocultural differences across racial groups (Hunter & Schmidt, 2010). With common use of the PSWQ, it is important to gain an improved understanding of its properties and functioning in diverse populations to aid interpretation and use, both academically and clinically. Based on available research, we hypothesized that there would be a degree of noninvariance between Black and White participants (Carter et al., 2005; DeLapp et al., 2016). To the degree to which noninvariance was identified, local areas of strain (e.g., differentially performing item scores) across groups within measurement models would be identified. Because previous invariance research has not included other racial groups, the analyses involving Asian and Hispanic participants were considered exploratory.
Method
Demographic data were analyzed using IBM SPSS Statistics 26. Participants of the full study in which the PSWQ was included were 2,804 undergraduate students from introductory psychology courses at a large Midwestern U.S. university. 3 Participants were not included in the analyses if they met any of the following criteria: missing race value or reported “multiracial” (n = 126); failed both of the included instructional validity questions (“I sometimes have a fatal heart attack while watching television,” “Please choose ‘agree’ if you are paying attention right now”; n = 175); any missing PSWQ data points (n = 14). As a result, for the purpose of the analyses reported here, we attained the final usable N = 2,489. Each participated in this study for partial credit of a research exposure requirement. Participants self-identified as Asian or Asian American, Black or African American, Hispanic/Latino, or White/Caucasian. The sample was 50.4% White (n = 1,254), 25.6% Black (n = 637), 7.1% Asian (n = 176), and 17.0% Hispanic (n = 422). The sample was 50.1% male, with an average age of 19.45 (SD = 2.48). Procedures for the current study were reviewed and approved by the local Institutional Review Board. Participants were recruited from introductory-level psychology courses between 2008 and 2019 and completed a battery of self-report questionnaires that included the PSWQ. All participants completed the full-length PSWQ, from which PSWQ-A scores were derived.
Data Analysis Plan
Confirmatory Factor Analysis
A series of CFAs were conducted using Mplus Version 8.3 (Muthén & Muthén, 1998-2019) to examine the underlying structure of the PSWQ and the PSWQ-A across the four racial groups. Due to the ordinal-level nature of data obtained from PSWQ item scores, mean- and variance-adjusted weighted least squares (WLSMV) estimation was used for these analyses. Although the developers of the PSWQ put forth a one-factor solution (Meyer et al., 1990), the appropriateness of this factor structure has been questioned (e.g., Brown, 2003; Fresco et al., 2002). Since no one preferred model has yet been established, as previously noted, the three models receiving the most existing support were analyzed in the current sample: the one-factor solution, the one-factor solution with method effects, and the two-factor solution. As noted, the PSWQ-A one-factor solution previously identified by Hopko et al. (2003) was examined in the current sample as well.
In order to assess model fit, multiple common indices were examined (Brown, 2015; Kline, 2016). The root mean square error of approximation (RMSEA) was used with values between less than .05 indicating good fit, values .05 to .08 indicating reasonable fit, values .08 to .10 indicating mediocre fit, and values above .10 indicating poor fit (Browne & Cudeck, 1993). This represents a more stringent interpretation than some other work, which has suggested a cut off of .06 or .07, with scores below these cut offs indicating good fit (Hu & Bentler, 1999; Steiger, 2007, as cited by Hooper et al. 2008). Additionally, standardized root mean squared residual was used to further assess model fit, with values less than .08 indicating reasonable fit (Hu & Bentler, 1999) and values less than .05 indicating good fit (e.g., Diamantopoulos & Siguaw, 2000, as cited by Hooper et al., 2008). Model fit also was assessed using the Tucker–Lewis index (Tucker & Lewis, 1973) and comparative fit index (CFI; Bentler, 1990). For these indices, values less than or equal to .90 were considered to indicate poor fit, values greater than .90 were considered to indicate adequate fit (Bentler, 1990) and values greater than .95 were considered to indicate good fit (Hu & Bentler, 1999).
Additionally, in order to further assess model fit, Akaike Information Criterion (AIC; Akaike, 1987) and Bayesian Information Criterion (BIC; Schwarz, 1978) statistics were computed using robust maximum likelihood estimation, which was utilized only for these indices. Each of these statistics considers both model fit and complexity, with lower AIC and BIC values indicating a better-fitting model in relation to alternative models (Brown, 2015). These two fit statistics were derived using robust maximum likelihood estimation, as they are unavailable using WLSMV in Mplus Version 8.3 (Muthén & Muthén, 1998-2019).
Measurement Invariance Analysis
The CFA approach described by Brown (2015) was used to assess measurement invariance, in a series of increasingly restrictive tests. According to this approach, if each of the four groups demonstrated adequate model fit in the initial CFA, configural invariance is assessed to determine whether the four groups demonstrated equivalent patterns of factor structure and loadings. If support is found for configural invariance, this test is followed by a test of metric invariance, which examines whether item loadings onto the latent construct were equivalent across groups. Finally, if metric invariance is supported, an analysis of scalar invariance, or the equality of item intercepts, is conducted. If each of the above types of invariance is found, then comparison of the groups on the latent mean is interpretable, and the observed item score, for any given true score on the factor, will not differ depending on group membership.
At each of these steps, several considerations are made in order to determine whether invariance is supported. These include an examination of chi-square differences, change in CFI scores, and whether there is overlap in the confidence intervals of the RMSEA (Putnick & Bornstein, 2016; Reis & Judd, 2000). A CFI change score of .002 has been suggested as an appropriate cut off for the detection of metric and scalar invariance (Meade et al., 2008). It is important to note that the chi-square statistic, which is sensitive to sample size, may fail to differentiate appropriately between models with good and poor fit when sample size is too small (Hooper et al., 2008) and may suggest inappropriate rejection of good-fitting models when sample size is too large (Hooper et al., 2008). The chi-square statistic, therefore, should be interpreted with particular caution in the current study.
Results
Preliminary Analyses
Neither the PSWQ nor the PSWQ-A total score was normally distributed. Standardized values indicated that both distributions were significantly positively skewed (PSWQ = 4.24; PSWQ-A = 4.10) and significantly platykurtotic (PSWQ = −6.24; PSWQ-A = −8.12). At the item level, standardized skew values ranged from −28.37 to 13.33 for the full-length PSWQ, and from −2.75 to 8.96 for the PSWQ-A. The standardized kurtosis values at the item level ranged from −11.20 to 12.02 for the full-length PSWQ, and from −11.20 to −9.19 for the PSWQ-A. The range of observed scores spanned the full possible range for all items. Though standardized skew and kurtosis values indicate nonnormality, CFA parameters were estimated using WLSMV and is appropriate to be used when observed data violate normality assumptions (Brown, 2015).
Confirmatory Factor Analysis
PSWQ
Table 1 provides the complete CFA results for each PSWQ model examined across the four racial groups. Overall, the original one-factor PSWQ model did not receive support in any of the four examined groups. The RMSEA value fell in the poor classification for each of the examined groups, with this model providing a poorer fit to the data relative to its comparison (i.e., two-factor) model via chi-square difference testing. The other indices for the one-factor model generally ranged from mediocre/reasonable to good, although the AIC and BIC values were highest for this model relative to the other two full-length PSWQ models. Lower AIC and BIC values are preferred. The two-factor model and the one-factor method effect model both received support when examining fit indices, as indices all met or exceed thresholds for supporting model fit. The AIC and BIC values were generally, although not exclusively, more favorable (i.e., lower) for the one-factor method effect model relative to the two-factor model. Chi-square difference testing supported the one-factor method effect model relative to the two-factor model across each group. All modification indices were less than 4.0 for each model, which, according to Brown (2015), provides further evidence of goodness of fit for these models. When considered alongside conceptual concerns surrounding the meaning of an “Absence of Worry” factor within the two-factor model (Brown, 2003), the one-factor method effect factor ultimately was retained as the measurement model using the full-length item pool for further consideration. The standardized factor loadings of that model included some nonsignificant loadings among the reverse-keyed items and many of the reverse-keyed items showed loadings at a magnitude typically considered nonsalient (i.e., less than .3-.4; Bandalos, 2018).
Goodness-of-Fit Statistics for the Penn State Worry Questionnaire.
Note. Models computed using mean- and variance-adjusted weighted least squares estimation. Δχ2 (**p < .001) computed using Mplus 8.3 DIFFTEST function. PSWQ = Penn State Worry Questionnaire. PSWQ-A = Penn State Worry Questionnaire–Abbreviated; RMSEA = root mean square error of approximation; CI = confidence interval; LL = lower limit; UL = upper limit; CFI = comparative fit index; TLI = Tucker–Lewis index; SRMR = standardized root mean square residual; AIC = Akaike Information Criterion; BIC = Bayesian Information Criterion; df = degrees of freedom.
Fit statistic computed using robust maximum likelihood estimation.bΔχ2 comparing one-factor model and two-factor model. cΔχ2 comparing one-factor model with method effects and two-factor model.
PSWQ-A
The PSWQ-A showed good model fit across three of the groups. In Asian participants, three of the indices suggested good fit, with the RMSEA suggesting reasonable fit; this model was retained. See Table 1 for complete goodness-of-fit results. All modification indices were less than 4.0, which, according to Brown (2015), provides further evidence of goodness of fit for the model. All items loaded significantly (p < .01) onto the one-factor model across all groups, ranging from .71 to .89. See Table 2 for standardized factor loadings. Internal consistency estimates for the PSWQ-A, derived using coefficient omega, ranged from .93 to .94 across all racial groups. Using a combination of standardized item factor loadings and error variances, omega provides an estimate of internal consistency that reduces bias associated with more traditional estimates of internal consistency (e.g., Cronbach’s alpha; Dunn et al., 2014).
Standardized Factor Loadings From Penn State Worry Questionnaire.
Note. PSWQ = Penn State Worry Questionnaire; PSWQ-A = Penn State Worry Questionnaire–Abbreviated.
Reverse-keyed item.
p ≤ .001 (two-tailed).
Model Selection
To conduct measurement invariance analyses, a good-fitting model first must be demonstrated across each group (Brown, 2015). Given that a good-fitting model for the PSWQ-A was established in the sample, measurement invariance analysis for this model could proceed. A good-fitting model, as indicated by goodness-of-fit and lack-of-fit indices, was also found for the PSWQ, specifically the PSWQ one-factor method model. However, as previously mentioned, many of the reverse-scored items had nonsalient loadings across groups. As a result of this pattern of findings, there was insufficient support for summing the item scores into a composite across groups. Any subsequent measurement invariance analysis of the PSWQ one-factor method model was determined to have little practical value because the presence of nonsalient factor loadings in that model suggested lack of support for deriving a total scale score from item scores across groups. Therefore, given the poor model fit in the original one-factor analyses, low conceptual meaning for the two-factor model, and nonsalient factor loadings on the one-factor method model, it was determined that none of the full PSWQ models were appropriate for invariance analysis and therefore these models were not included in the subsequent analysis. Thus, only the PSWQ-A model was retained for continued measurement invariance analysis.
Measurement Invariance Analysis
PSWQ-A
Support for PSWQ-A configural invariance (equal factor structure) was indicated by reasonable to good fit across indices. Support for PSWQ-A metric invariance (equal factor loadings) was indicated by reasonable to good fit across indices. Support for PSWQ-A scalar invariance (equal item intercepts) was indicated by strong fit across indices. These data support configural, metric, and scalar invariance across the four racial groups. Furthermore, BIC decreased at each step of the PSWQ-A measurement invariance analysis; AIC decreased between the configural and metric invariance analyses but increased at the scalar invariance step.
Chi-square difference testing is strongly influenced by sample size (Cheung & Rensvold, 2002; La Du & Tanaka, 1989). Thus, changes in CFI between the models of invariance were examined to assess measurement invariance for which changes less than or equal to 0.002 are considered acceptable (Meade et al., 2008). The ΔCFI was less than .002 between configural and metric invariance as well as between metric and scalar invariance, which suggests invariant properties that could indicate differences either in how the construct of worry is measured or how it is experienced—or both—between racial groups. Overlapping RMSEA confidence intervals across models similarly supports invariance (Table 3).
Goodness-of-Fit Statistics for Measurement Invariance Analyses.
Note. Models computed using mean- and variance-adjusted weighted least squares estimation. Δχ2 (*p < .05. **p < .001) computed using Mplus 8.3 DIFFTEST function. PSWQ-A = Penn State Worry Questionnaire–Abbreviated; RMSEA = root mean square error of approximation; CI = confidence interval; LL = lower limit; UL = upper limit; CFI = comparative fit index; TLI = Tucker–Lewis index; SRMR = standardized root mean square residual; AIC = Akaike Information Criterion; BIC = Bayesian Information Criterion; df = degrees of freedom.
Fit statistic computed using robust maximum likelihood estimation. bΔχ2 comparing metric and configural. cΔχ2 comparing scalar and metric.
Group Differences
Latent mean differences were examined, using the White respondents as the reference group. Compared with that reference group, Black participants showed a deviation of −0.12, z = 2.75, p = .006; Hispanic participants showed a deviation of 0.06, z = 1.13, p = .192; and Asian participants showed a deviation of 0.07, z = 1.17, p = .241. The only significant difference was between White and Black respondents, where Black participants scored, on average, 0.12 lower than White participants on the latent variable.
Discussion
Although the importance of verifying the structural validity of psychological symptom severity measures in racial minority populations is widely recognized, examinations of this type have tended to be limited for the PSWQ and PSWQ-A. This study is the first to our knowledge that explicitly tested the structure of the PSWQ and PSWQ-A across White, Black, Hispanic, and Asian individuals using measurement invariance analysis. This gap in the literature is especially important to address given the observed differences in worry severity across racial groups (e.g., Saw et al., 2013; Scott et al., 2002; Williams et al., 2012). The present analyses inform how to operationalize item scores (e.g., as total scores) and whether mean comparisons are meaningful across racial groups.
PSWQ
For the PSWQ, the present study examined structural models previously proposed within the literature. A series of CFAs revealed that the original one-factor model showed inadequate fit for each racial group as determined using common goodness-of-fit indices. After testing the three primary models found in the literature, it was found that the one-factor model with method effects had the best fit, reflecting prior findings reported by Brown (2003). However, none of the models of the full-length PSWQ were deemed appropriate for subsequent measurement invariance analysis. The original one-factor model (Meyer et al., 1990) illustrated poor fit across groups on goodness-of-fit and lack-of-fit indices. Though the two-factor model (Fresco et al., 2002) improved on the fit indices compared with the original one-factor model, the two-factor model is conceptually inconsistent with the original intention of the measure as proposed by Meyer et al. (1990). Whereas the one-factor method model (Brown, 2003) indicated the best fit across the three PSWQ models, the standardized factor loadings for some of the negatively worded items were nonsalient, calling into question the soundness of summing the items to create one composite score across groups. These concerns add to the existing literature, which reflect an inconsistent and incongruent factor structure of the full-length PSWQ.
PSWQ-A
The PSWQ-A was developed in part to address noted factorial problems of the PSWQ and eliminate potential method effects (Hopko et al., 2003). In the present study, the one-factor PSWQ-A model displayed good fit across all racial groups. The exception was the confidence interval upper limit for the RMSEA value for the Asian sample. Though the RMSEA value was considered indicative of reasonable fit, the confidence interval fell above .10. Of note, the RMSEA is sensitive to small sample sizes (Kline, 2016). A sample size of 200, which represents an approximation of the median sample size across many studies utilizing such methods, is considered typical (Shah & Goldstein, 2006). The number of Asian participants in the current study was smaller (n = 176), but meets minimum recommendations of the sample size to parameters estimated ratio of 10:1 (Kline, 2016). Overall, these findings show that the PSWQ-A single factor solution is supported for each racial group. One reason for the difference in findings across the PSWQ and PSWQ-A may be that the PSWQ-A eliminates reverse-keyed items.
Results supported configural invariance of the PSWQ-A, indicating that the basic organization of the model (i.e., the pattern of factor loadings) is equivalent in all four racial groups. Full metric invariance across the four groups was supported, indicating that the degree to which each item loads onto the latent factor, worry, is equivalent across groups. This finding indicates that no one item is more closely related to worry in one group than in the others, which contrasts the findings of DeLapp et al. (2016). This contrast with previous research may be related in part to the present study’s large sample size or the inclusion of four racial groups. Scalar invariance also was supported for the PSWQ-A, suggesting that any mean differences between any of the four groups on any PSWQ-A item are related to differences in worry and not to measurement error caused by other (e.g., cultural) factors between the groups.
Results revealed some group differences on the PSWQ-A; participants who identified as Black showed significantly lower scores compared with those who identified as White and Hispanic. However, the size of the effects for these group comparisons was small, suggesting that the large sample sizes may have resulted in a Type I error or the true effects have limited practical significance. This interpretation fits with previous research comparing White and Black individuals that displayed a lack of mean-level differences on the PSWQ-A (Brenes et al., 2008; Shrestha et al., 2019). The lack of apparent meaningful differences on the PSWQ-A between the other groups is consistent with the literature comparing White and Hispanic individuals (Nuevo et al., 2007) and White and Asian individuals (Tang et al., 2016).
Conclusions, Limitations, and Directions for Further Research
Overall, our findings suggest that the PSWQ-A measures worry equivalently across samples of White, Black, Hispanic, and Asian individuals. This study provides substantial support for continued use of the PSWQ-A in each of these populations, as well as comparisons of mean worry scores as measured by the PSWQ-A. The measurement invariance analysis provides some confirmation that differences among the four racial groups tested on the PSWQ-A reflect true differences on the latent construct it measures, worry. This questionnaire often is used as a tracking measure in clinical settings, and these results imply that this use of the measure is appropriate for individuals who identify as White, Black, Hispanic, or Asian.
Our findings converge with other research supporting the structural validity of the PSWQ-A (e.g., Kertz, Lee, et al., 2014). Additionally, strong structural validity was not present in established measurement models of the full-length PSWQ. This finding by itself does not necessarily render the PSWQ-A as the preferred measure over the PSWQ, given that other aspects of reliability and validity of the scales are important to consider. For example, the full-length PSWQ may carry advantages beyond structural validity. Fewer items on a measure tend to result in lower reliability via internal consistency. Notably, however, despite containing fewer items than the full measure, the PSWQ-A has shown comparable levels of internal consistency (e.g., Hopko et al., 2003), with strong values found in the present study. However, this does not ensure that the reduced item set of the PSWQ-A offers the same content coverage as the full-length PSWQ. Additionally, considering that the PSWQ-A is unifactorial, whereas the full-length PSWQ appears to not be unifactorial, overemphasizing internal consistency of the respective total scores may be unproductive. Overall, the evidence for use of the PSWQ-A is compelling, particularly when considering the importance of structural validity in psychological measurement. Future studies using independent samples may consider ways to improve the structural validity of the full-length PSWQ, such as more targeted item removal (e.g., only the reverse-keyed items). This practice may be especially indicated should additional research find that the PSWQ holds advantages over the PSWQ-A, such as regarding issues not examined in the current study (e.g., diagnostic accuracy, improved breadth of key content). 4 Pursuant to that point, as noted, preliminary research suggests the full-length PSWQ total score may have greater diagnostic accuracy relative to the PSWQ-A (Wuthrich et al., 2014). Such analyses are based on the premise that the 16 items comprising the full-length item pool should in fact be summed into a total scale and there is no consensus on the viability of that practice from existing study findings, including the present results. Although Wuthrich et al.’s findings warrant replication, it underscores the point that the PSWQ-A may not be preferred in all settings. Nonetheless, the present results further support existing research suggesting the PSWQ-A may help address shortcomings surrounding the structural validity of the full-length PSWQ item pool.
This study is limited by several factors, including a smaller sample of Asian American participants in comparison with the other racial groups. Because some of the indices used to examine model fit and measurement invariance in this study are vulnerable to sample size, it is uncertain whether the Asian group at times differed from the other groups (e.g., showed slightly poorer unidimensional model fit on the PSWQ-A) due to a smaller sample size or due to true differences. Second, there are many considerations to be made when making inferences based on self-identified racial categories. It is unknown whether the individuals in this study are representative of their racial group on the construct examined. It will be important in future examinations of measurement invariance on the PSWQ/-A and other psychological measures based on race to examine other aspects of cultural identification. The grouping that was done in this study did not allow for examining differences within groups regarding the present research question. A related limitation resides in the way in which the racial categories were coded for this study. Participants who identified as Hispanic were placed in the “Hispanic” group even if they also identified as Black, Asian, or White. Additionally, racial categories were presented as mutually exclusive, and any participant who reported being “multiracial,” “other,” or preferred not to respond (n = 126) were excluded from analyses. This practice introduces variance—including the impact of multiple racial backgrounds on perception and development—that should not be overlooked, given the difficulties in parsing the complexities of racial and ethnic identity. Future research should consider the relationship between various combinations of mixed-race backgrounds and self-reported worry. Finally, the present study focused on existing proposed structural models for the full-length and abbreviated versions of the PSWQ. Future research may consider alternate models to improve the structural fit of either version of the measure.
Although both versions of the PSWQ generally are considered dimensional in nature (Kertz, McHugh, et al., 2014; Olatunji et al., 2010) and unselected samples often are used in their examination, it is important to examine measurement invariance across race in a clinical sample. Although that examination may limit the range of observed scores, with participants reporting more uniformly higher ratings on the item scores, it is important to determine whether the present findings hold among individuals who consistently report high worry severity, on average. Next, the current sample consisted of college students and thus is limited in terms of age range and variability of educational background. Given that some research has shown differences in PSWQ scores between age groups (Gillis et al., 1995), performing the same analyses in samples of more varied age and educational background may be important. In the current sample, there was a greater number of males than females among White participants, and more females than males among Black participants. This statistically significant difference on the variable of sex indicates that the results should be interpreted in this context. However, it should be noted that this may be due in part to a large sample size and that the groups showed similar frequencies despite displaying statistically significant differences (41.1% to 56.9% male). Finally, the PSWQ-A was derived from the PSWQ item pool and future research may seek to specifically examine the PSWQ-A using the eight-item pool independent of the parent item pool (Smith et al., 2000).
Despite its limitations, this study offers evidence that the PSWQ-A can be interpreted equivalently across White, Black, Hispanic, and Asian groups. Although more research is needed to address the gaps left by the limitations of the present study, it provides support that researchers and clinicians are justified in using the measure with the current racial groups.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
Methodological Disclosure
We report how we determined our sample size, all data exclusions, all manipulations, and all measures in the study.
