Abstract
Internalized homophobia (IH) refers to negative attitudes and stereotypes that a lesbian, gay, or bisexual (LGB) person may hold regarding their own sexual identity. Recent sociocultural changes in attitudes and policies affecting LGB people generally reflect broader acceptance of sexual minorities, and may influence the manner in which LGB people experience IH. These experiences should be reflected in the measurement properties of instruments designed to assess IH. This study utilized data from three different samples (N = 3,522) of LGB individuals residing in the United States to examine the invariance of a common self-report IH measure by gender identity (Female, Male) and age cohort (Boomers, Generation X, and Millennials). Multigroup item response theory–differential item functioning analysis using the alignment method revealed that 6 of the 9 Internalized Homophobia Scale items exhibited differential functioning across gender and generation. Latent scores based on the invariant items suggested that Male and Female Boomers exhibited the lowest level of latent IH, relative to the other cohorts.
Keywords
Researchers and clinical practitioners often administer self-report instruments to populations characterized by a wider range of ages and backgrounds than the original validation sample. The same instrument may also be used for decades, which allows broader sociocultural changes to shape the interpretations or meaning of individual items, or the scale concept itself. Constructs specific to lesbians, gay men, and bisexual men and women (lesbian, gay, or bisexual [LGB]) may be particularly sensitive to changes over time because of unfolding political events, laws, and judicial decisions (e.g., Obergefell v Hodges, 2015), and changes in public attitudes toward LGB individuals. Internalized homophobia (IH), which refers to the degree to which an LGB person consciously or unconsciously believes negative attitudes and stereotypes regarding LGB people (e.g., as portrayed by institutions, media, within a family or community; Herek, 2007) is one type of minority stress (Meyer, 2003) at the individual level of the ecological environment. IH may be learned through socialization in a stigmatizing environment and may be sensitive to ongoing sociocultural changes.
Conceptualizations of IH appear in the early coming out models theorized by Cass (1979) and Fassinger (1998; see also Szymanski, Kashubeck-West, & Meyer, 2008; Williamson, 2000, for reviews), and decades of research have established a relationship between IH and individuals’ levels of psychological distress (Newcomb & Mustanski, 2010), global well-being, and close relationship formation and functioning (Frost & Meyer, 2009). Recent calls to better address the impact of different LGB identities across the minority stress literature (e.g., Parra & Hastings, 2018) underscores the importance of understanding the extent to which the measurement properties of IH instruments are invariant across individuals with different gender identities and ages. The present study describes multiple group item response theory (IRT) analysis conducted on responses collected across three samples of adult LGB participants who completed a common measure of IH (Wright, Dye, Jiles, & Marcello, 1999; Wright & Perry, 2006). We begin by briefly reviewing recent studies examining measurement invariance (MI) across different sexual and gender minority stress constructs, followed by a discussion of existing measures of IH. After introducing the measure used in the present study, we discuss the rationale for using gender/sex and generation as grouping variables in the ensuing MI analysis. Our initial analysis featured the standard alignment approach described by Asparouhov and Muthén (2014) with six combinations of gender identity (Male, Female) and generation (Boomer, Generation X, and Millennials) as the grouping variables. Cross-group differences in item response properties were further probed using a Bayesian Alignment-within-CFA (AwC) analysis described by Marsh et al. (2018).
Measurement Invariance in Sexual Minority Stress
Although research examining the psychometric properties of IH measures across gender/sex and age is limited, a number of prior studies have evaluated MI of self-report measures related to sexual minority stress. Most of these studies find support for metric (i.e., factor loading) or strong factorial (i.e., factor loading and intercept) invariance as a function of sex or gender identity, which suggests that these measurement instruments measure essentially the same construct(s) across these groups. For example, Frost and Meyer (2012) developed a measure of lesbian, gay, bisexual, and transgender (LGBT) community connectedness, and found evidence of metric invariance across gender/sex (men vs. women) and race or ethnicity (White vs. people of color; Frost & Meyer, 2012). A more recent study examined young adults’ retrospective evaluations of the coping strategies used in response to sexual orientation-related stress during adolescence (Toomey, Ryan, Diaz, & Russell, 2018). Analysis found evidence of metric invariance as a function of self-reported sex (male vs. female), and strong factorial invariance as a function of ethnic group identification (Latino/a vs. non-Latino/a White).
Lehavot, King, and Simoni (2011) explored gender expression among lesbian and bisexual women through the gender expression measure among sexual minority women, and found evidence of partial factorial invariance across butch, femme, androgynous, and none-identifying gender groups. The noninvariance present in the Lehavot study was isolated to a handful of items across two or fewer groups, and follow-up analyses were consistent with latent factor mean equivalence across all groups. In the process of validating a short form for the Anti-Bisexual Experiences Scale, Dyar, Feinstein, and Davila (2019) selected a final revised item pool that exhibited strong factorial invariance across cisgender and gender minorities, as well as across bisexual and nonmonosexual identified individuals. Bauerband, Teti, and Velicer (2019) describe a comprehensive series of MI analyses examining differences in the Everyday Discrimination Scale and the Discrimination-Related Vigilance Scale of the across cisgender and transgender sexual minority individuals. Analyses found evidence in support of a partial metric invariance on the Everyday Discrimination Scale across cisgender and transgender, and within different transgender identities, whereas metric invariance was observed for the Discrimination-Related Vigilance Scale across all items and groups.
MI analysis provides direct insight into how different identities shape minority stress experiences as reflected in constructs created to measure these experiences. Specifically, when applied for this purpose, MI testing directly acknowledges and evaluates potential differences across a group by examining how individuals with different characteristics (e.g., a 60-year-old lesbian) may respond differently to some aspects of a minority stress construct, relative to others with overlapping or different identities (e.g., a 25-year-old lesbian or a 62-year-old gay man).
Measuring Internalized Homophobia
IH, also referred to as internalized homonegativity or internalized heterosexism in the literature, represents a proximal component of sexual minority stress (Brooks, 1981; Mayfield, 2001; Meyer, 2003). In a research context, IH has been utilized as a predictor (Frost & Meyer, 2009; Puckett et al., 2017; Williamson, 2000), mediator (Jellison & McConnell, 2003), moderator (Szymanski, 2006), outcome variable (Rowen & Malcom, 2002), and a symptom indicator in a therapeutic context (Shidlo, 1994). There are numerous published measures of IH (e.g., Herek, Cogan, Gillis, & Glunt, 1998; Herek, Gillis, Cogan, 2009; Ross & Rosser, 1996; see also Berg, Munthe-Kaas, & Ross, 2016, for a recent systematic mapping review of internalized homonegativity), with some focused on measuring IH for a specific group (e.g., the Lesbian Internalized Homophobia Scale; Szymanski & Chung, 2001) and others for multiple groups (e.g., Measure of internalized sexual stigma for lesbians and gay men; Lingiardi, Baiocco, & Nardelli, 2012). Though the exact wording, number of items, and target populations may differ, there are theoretical and linguistic similarities across conceptualizations. Most measures contain item sets (or complete subscales) pertaining to external perceptions of LGB individuals and culture, as well as items describing the respondents’ negative evaluation of their own sexual identity (Herek et al., 1998; Martin & Dean, 1992; Puckett et al., 2017; Ross & Rosser, 1996; Szymanski & Chung, 2001), including a desire to “be heterosexual” and regret over homosexual or same-sex attraction or behavior.
The Wright Internalized Homophobia Scale
The current study focuses on a measure of IH developed by Wright et al. (1999; see also Wright & Perry, 2006). Although this measure was initially developed as part of a project examining minority stress and health behaviors among young adults in the U.S. Midwest, the item content is representative of the larger body of IH measures in the published literature. Participants respond to the items using a Likert-type scale ranging from 1 (strongly agree) to 5 (strongly disagree). Four of the items are worded such that stronger endorsement of the statement reflects positive feelings about LGB identity (e.g., “I am proud that I am gay/lesbian/bisexual”), and five of the items represent negative feelings about LGB identity (e.g., “I often feel ashamed that I am gay/lesbian/bisexual”). For analyses using this scale, the four “positive” items are reverse scored and combined with the remaining items to create an index of IH or homonegativity, such that higher scores represent stronger negative feelings and attitudes about the respondent’s LGB identity.
In the original Internalized Homophobia Scale (IHS) development sample (Wright et al., 1999), interitem consistency reliability was strong (α = .87), indicating that the scale serves as an internally consistent measure of IH. Wright and Perry (2006) also reported the results of an exploratory factor analysis performed on a seven-item version of the measure (omitting Items 7 and 9 from the 1999 version), and found strong evidence supporting the unidimensionality (first eigenvector accounts for 54% of variance) and reliability (standardized factor loadings: range = .519-.816, average = .726, α = .83) of the measure. Wright et al. (1999) also reported that the scale exhibited adequate convergent validity with measures of distress and discriminant validity with measures of self-esteem. The individual items comprising the Wright et al. IH scale are similar to items found across other scales. For example, in Martin and Dean’s (1992) IHP scale (adapted for female respondents), the item “I wish I weren’t lesbian/bisexual” is similar to the generally worded item “I wish that I weren’t attracted to the same sex.” The item, “Sometimes I feel ashamed of my sexual orientation,” developed for use with young men who have sex with men (Puckett et al., 2017) is similar to the Wright et al. IH scale item, “I often feel ashamed that I am gay/lesbian/bisexual.”
Generation, Gender, and Internalized Homophobia
The presence of MI suggests that individuals identifying with the invariant groups may share a similar conceptualization of the underlying construct. Moreover, where noninvariance in item parameters (e.g., factor loadings, intercepts/thresholds) emerges, it may provide insight into how individuals with different identities or characteristics may experience or relate to the corresponding phenomenon differently. In contemporary models of coming out, the majority of individuals who identify as lesbian or gay begin the process of establishing their sexual identity during adolescence and young adulthood, although the average age of coming out and disclosure may be related to the era in which an individual is maturing. For example, Grov, Bimbi, Nanin, and Parsons (2006) in a sample of 2,733 LGB individuals in New York and Los Angeles found that younger cohorts came out earlier than older cohorts of men and women, and women came out later than men, but with no differences by racial identity. A Pew Research Center survey of 1,197 LGBT adults found that “younger gay men and lesbians are more likely to have disclosed their sexual orientation somewhat earlier in life than have their older counterparts” (Pew Research Center, 2013). This difference may reflect changing and more accepting social norms, and that LGB people may come out at any point in life (individuals that will identify as LGB in the future are not included in the sample).
Coming of age in different time periods will reflect broader cultural attitudes during the time period and shape the manner in which individuals internalize negative thoughts and beliefs regarding same-sex attraction. For instance, a recent Pew Research Center (2017) poll found Americans’ endorsement of same-sex marriage rights to be at a historical high across a broad range of demographic categories, and although LGB individuals still face discrimination and harassment in many contexts, the sociopolitical context of the daily lives of LGB persons in 2019 differs significantly from that of 1969.
These changes impact the experiences of LGB individuals depending on when they were socialized and when they came out (Ramirez-Valles, 2016), suggesting that LGB individuals of different generations may experience IH differently. For example, Baby Boomers experienced the civil rights movement and lived through a time in which the Diagnostic and Statistical Manual pathologized homosexuality; Generation X experienced the Stonewall riots and gay (and lesbian) liberation movement; and Millennials experienced Bowers v. Hardwick (1986) and Don’t Ask, Don’t Tell (American Psychiatric Association, 1952; Baunach, 2011; Herek, 2007). Furthermore, the LGB community has collectively experienced an increase in acceptance by society over time (Baunach, 2011). The lived experience of these historical events may also shape older LGB individuals’ perception of more recent advances and attitudinal shifts, further contributing to potential differences in the meaning of IH for different generations of LGB individuals.
As discussed above, gender may also shape individuals’ experiences of socialization and IH (see McCarn & Fassinger, 1996; Szymanski & Chung, 2001), which further underscores the need to explore the validity of measures of IH for different generations as well as by gender (Szymanski et al., 2008). For example, IH is associated with risk for poor mental health outcomes (see Meyer, 2003, for a review) and poor relationship quality (Frost & Meyer, 2009). Some research finds no difference between gay men and lesbians in the role of IH on outcomes (e.g., Frost & Meyer, 2009), while other research finds differences by gender (e.g., McLaren, 2015). For example, Szymanski (2006) found that while psychological distress resulting from heterosexist experiences can be moderated by IH for gay men, IH did not moderate the psychological effects of experiencing heterosexist events for lesbians. These results suggest that an individual’s gender can affect their interpretation and experience of IH, which further underscores the need to evaluate the gender/sex invariance of IH measures.
The Present Study
The recent increase in number of studies examining MI in sexual minority stress-related constructs as a function of gender, ethnicity, race, and sexual identity may reflect more general growth in research on sexual minority stress. However, it may also be indicative of efforts to acknowledge the role of different characteristics and identities in how stressors are experienced. As a minority stressor, IH may be responsive to shifts in broader cultural attitudes and context, and these changes may be reflected by differences in the measurement properties of scales such as the IHS. For instance, over the past few decades, there have been dramatic changes support for LGBT rights and acceptance of LGBT people in American society. The increased prevalence and visibility of openly LGB figures (e.g., in politics, entertainment, and media), and changes in the societal lexicon used to describe the LGB community and its members, are reflective of such a shift.
Given these changes, and the lack of published studies examining MI in IH constructs as a function of respondent age and gender (e.g., Berg, Weatherburn, Ross, and Schmidt’s [2015] review did not include systematic examination of differences by age or gender), and the ongoing use of these constructs in theoretical and exploratory testing of associations between IH and mental health, there is an acute need to understand the invariance qualities of IH measures. Additionally, the aforementioned studies appeared to apply traditional confirmatory factor analysis (CFA) approaches to evaluating MI, which treat the item-level Likert-type responses as continuous indicators. Although many of these studies used robust (e.g., Sattorra–Bentler corrected) estimators to account for excessive skew and kurtosis associated with Likert-type indicators, these approaches provide less detail than the ordinal response model used in the present study. We did not formulate hypotheses regarding the differential functioning of specific items; based on the literature, we did expect to find partial MI for the IH scale as a whole, such that some questions would possess equivalent measurement properties (i.e., factor loadings and threshold parameters) across groups, whereas others would differ meaningfully.
Method
Participants and Procedure
Sample 1
Sexual minority individuals were recruited through e-mails to listservs of LGB groups in the United States to participate in an online survey in 2002 (n = 315; Riggle, Rostosky, Prather, & Hamrin, 2005). The mean age of the sample was 51.6 years (SD = 9.69), with 48.89% of participants being members in the Baby Boomer generation, 48.57% in Generation X, and 2.54% in the Millennial generation. Participants identified as female (53.97%) or male (46.03%). The majority of participants identified as White (87.62%), and the second largest group identified as multiracial (4.13%). Nearly all participants in this sample (98.41%) identified as gay/lesbian, none identified as bisexual, and the remainder chose not to answer or provided another response.
Sample 2
Participants were recruited by e-mails to listservs of LGB groups in the United States and online announcements through LGB focused webpages to complete an online survey conducted in June 2006 (n = 1,450; Rostosky, Riggle, Horne & Miller, 2009). The mean age of the sample was 39.54 years (SD = 12.22), with 16.14% of the sample in the Baby Boomer generation, 42.83% in Generation X, and 41.03% in the Millennial generation. The sample included 60% female participants and 40% male participants. The majority of the sample was White (87.31%), and the second most frequent response was Other (2.56%). The majority of participants (81.49%) identified as gay or lesbian, more than one tenth (10.64%) identified as bisexual, and the remainder chose not to answer or provided another response.
Sample 3
Individuals who identified as LGB were recruited by e-mails to LGB listservs in the United States and announcements posted on LGB focused websites to complete an online survey in November 2006 (n = 1,759; Rostosky et al., 2009). The mean age of the sample was 38.77 years (SD = 12.54), and 15.63% of participants were members of the Baby Boomer generation, 41.16% were members of Generation X, and 43.21% were members of the Millennial generation. Participants identified their gender/sex as female (59.91%) or male (43.09%). The majority of participants identified as White (87.17%), and the second most frequent response was Hispanic/Latino (2.79%). The majority of participants (81.18%) identified as gay or lesbian, more than one tenth (11.20%) identified as bisexual, and the remainder chose not to answer or provided another response.
Measures
Demographic Questions
Participants answered items regarding their age, gender identity, ethnicity, and education level on the self-report survey and were then categorized into generational breakdowns corresponding to the year in which they were born. Birth year was calculated by subtracting their reported age from the year at the time of data collection. Generational classifications were developed based on ranges of years as indicated in the literature and informed by the U.S. Census Bureau. The Baby Boomer generation consists of individuals born between 1945 and 1964 (Colby & Ortman, 2014), Generation X includes those born between 1965 and 1981 (Borges, Manuel, Elam, & Jones, 2006; Bristow, Amyx, Castleberry, & Cochran, 2011), and that the Millennial generation consists of individuals born between 1982 and the mid-2000s (Benfer & Shanahan, 2013; Howe & Strauss, 2007). Participants were presented with the following question: “Which of the following best describes your primary gender identification? (choose from the drop-down menu),” and were provided with the following answer choices: female, male, transgender, intersexed, other, prefer not to answer. Demographic characteristics by sample are provided in Table 1.
Age and Frequencies by Sample.
Note. Column percentages are in parentheses.
Internalized Homophobia Scale
The IHS (Wright et al., 1999; Table 2) is a nine-item self-report scale. The Likert-type items includes responses ranging from 1 (strongly agree) to 5 (strongly disagree). The items address positive and negative attitudes toward the respondent’s self-identification as an LGB individual and their sexual attraction to individuals of the same sex.
The Wright Internalized Homphobia Scale.
Note. GLB = gay, lesbian, or bisexual. Response Scale: 1 = strongly agree, 5 = strongly disagree. (R) identifies reverse-scored items.
Data Preparation and Analysis Plan
More than 99.9% of respondents (N = 3,522) selected the female or male responses to the gender question, and nine individuals selected transgender, intersexed, other, or prefer not to answer. As a result, we were only able to include female and male identifying individuals in the present study. Participants were classified into gender and generational groups based on their responses to demographic questions, resulting in group sizes ranging from 327 (Female Boomers) to 874 (Female Millennials). Participants responded to the IH using a 5-point Likert-type response scale, but examination of group-specific response category frequencies revealed low response rates in some groups for one or more response categories, which often leads to estimation problems in ordinal response models. For example, only two individuals in the Female Boomer group provided a response in the highest category on Item 1, which means the corresponding (fourth) threshold in our ordinal response model would be poorly estimated for this combination of item and group, and any corresponding differential item functioning (DIF) tests would be effectively meaningless.
In an effort to maximize the number of response categories modeled for each item, while ensuring the stability and reliability of group-specific parameters and the corresponding DIF tests, a minimum group-specific item response category frequency of n = 10 was established. More extreme response categories below the minimum frequency benchmark were collapsed into the less extreme adjacent category until this minimum frequency was obtained. Continuing with the previous example, it was necessary to collapse across the fifth (n = 2), fourth (n = 3), and third (n = 24) response categories for Item 1 in the Female Boomer group to obtain the minimum response frequency. For this item, the recoded highest response category now represents the number of individuals who endorsed 3 or higher, whereas the lowest and middle categories still correspond to the first and second response options. Applying this decision rule resulted in collapsed response categories for all items, except Item 9, and the number of threshold parameters listed in Table 3 identifies the number of collapsed categories for each item. 1 Additionally, items expressing a negative sentiment regarding the respondent’s sexual identity (i.e., 2, 3, 5, 7, 8) were then reverse-coded so that endorsement of all items was consistent with negative beliefs about one’s LGB identity (i.e., homophobia).
Aligned Factor Loadings (λ), Threshold (τ), Factor Mean (α), and Variance (ψ) Parameters.
Note. λ = factor loading; τ = threshold; α = latent factor mean; ψ = latent factor variance. Items 2, 3, 5, 7, and 8 were rekeyed so that endorsement reflected higher levels of internalized homophobia. Estimates based on initial alignment analysis using (robust) maximum likelihood using a probit link (normal-ogive item response theory model). Parameters with at least one noninvariant group denoted by *. Fixed parameters in italics. Parameters significantly different from the pooled estimate (p < .05) in bold. Factor means (α) with different superscripts are significantly different from one another.
The analyzed data were composed of individuals from six discrete groups (Female Boomers, Male Boomers, Female Xers, Male Xers, Female Millennials, Male Millennials) who provided responses to nine unidimensional items measured by a 2 to 5 category ordinal response format, which were subjected to a multigroup normal-ogive IRT analysis using Mplus 8.3 (L. Muthén & Muthén, 1998-2018).
DIF Analysis Using the Alignment Method
Traditional approaches for evaluating DIF or MI testing rely on stepwise testing approaches (see Stark, Chernyshenko, & Drasgow, 2006), which can become cumbersome and error-prone as the number of items and groups increases. In contrast, the alignment method for DIF detection overcomes the limitations of prior approaches and allows the analyst to identify noninvariant or differentially functioning items, as well as examine latent mean differences across groups in the presence of partial invariance (Asparouhov & Muthén, 2014; Flake & McCoach, 2018; B. O. Muthén & Asparouhov, 2014).
Alignment analysis begins with a minimally constrained baseline model, identified by standardizing the scale of the latent factor in all groups, and freely estimating all item factor loadings and thresholds. Next, a component loss function is optimized so that group-specific parameters for the latent factor mean (α) and variance (ψ) are rescaled to reflect cross-group differences in the level of the trait, while simultaneously minimizing cross-group differences in factor loadings (λ) and item thresholds (τ). This aspect of the alignment approach is conceptually similar to the process of rotation in exploratory factor analysis, and is designed to produce a solution characterized by “few large noninvariant measurement parameters and many approximately invariant measurement parameters rather than many medium-sized noninvariant measurement parameters” (Asparouhov & Muthén, 2104, p. 497). The scale of the aligned solution is identified by fixing the factor mean (α) to 0 and factor variance (ψ) to 1 in a designated reference group, and estimating α and ψ for each of the remaining groups. 1 A multiple-comparison corrected algorithm is used to identify significant differences in latent factor means, item factor loadings, and threshold parameters across groups.
Bayesian AwC Model
Marsh et al.’s (2018) AwC procedure allows the aforementioned alignment solution to be recast as a standard multigroup IRT model characterized by the same number of estimated parameters and model likelihood. When combined with Bayesian estimation, this extension of the standard alignment analysis provides a number of opportunities to refine the initial solution by examining assumptions that could otherwise not be tested. For instance, the joint maximum likelihood (ML) estimator used in standard IRT models (including the initial aligned model) assumes item responses to be independent after accounting for the latent trait (Embretson & Reise, 2000). Traditionally, this “local independence” assumption has been difficult to evaluate because global model fit indices were not available for IRT models estimated using ML. However, recent advances in Bayesian structural equation modeling (BSEM; Muthén & Asparouhov, 2012) and the development Markov Chain Monte Carlo (MCMC) estimation algorithms allows the analyst to both test, and account for, violation of the local independence assumption. More specifically, BSEM provides insight into model adequacy using posterior predictive checks (PPC), which are based on the distribution of likelihood ratio statistics reflecting the discrepancy between the likelihood of the observed data, and data simulated from the posterior distributions of parameter estimates (Gelman, Carlin, Stern, & Rubin, 2004).
Generally, an estimated model is considered to provide a reasonable fit for the data when the 95% confidence interval (CI) of PPC values contains 0, and an optimal-fitting model is obtained when the median of this distribution equals 0 (i.e., posterior predictive p = .5). Additionally, the BSEM framework allows for a respecification of the aligned solution incorporating all (k * k − 1)/2 possible residual covariance parameters into the model, which will account for any violations of the local independence assumption. Although this model would not be identified using frequentist estimators (e.g., ML, weighted least squares), extensive simulation work by B. O. Muthén and Asparouhov (2012) demonstrates that these parameters are estimable using MCMC through the specification of small-variance priors for the residual parameters.
Results are presented in three phases. First, the results from a standard ML multigroup normal-ogive IRT model using the alignment method are presented. Next, the corresponding Bayesian AwC (BAwC; Marsh et al., 2018) model will be estimated in order evaluate the fit of the initial alignment model, and ensure that cross-group differences in latent factor means and item parameters remain credible after accounting for any violations of the local independence assumption. Specifically, if the initial BAwC model provides a less than optimal fit for the data, a follow-up BAwC will be estimated using small-variance prior residual covariances estimated for each group (B. O. Muthén & Asparouhov, 2012), which should result in an optimally fitting model. Finally, the robustness of the initial DIF analysis will be evaluated by simulating the posterior distributions representing the discrepancy between group-specific and pooled-invariant parameters.
Results
The overall sample was composed of N = 3,522 eligible cases (2,041 female), was predominately White (87.2%) and highly educated (74.47% attended at least some college or technical school). Average age in the overall sample was 40.25 years (SD = 12.68), resulting in a generational distribution composed of 18.82% Baby Boomers (“Boomers”) 42.53% Generation X (“Xers”), and 38.65% Millennials. The large overall sample size and parity of the gender/sex and generation distributions resulted in substantive subgroups that were sufficiently large to detect meaningful invariance effects (Flake & McCoach, 2018). The primary invariance tests were based on 327 Female Boomers, 336 Male Boomers, 840 Female Xers, 658 Male Xers, 874 Female Millennials, and 487 Male Millennials.
Prior to performing the primary analysis, an exploratory factor analysis was conducted using Mplus 8.3 to verify the unidimensionality of the IHS in the present sample. The initial eigenvalues based on the polychoric correlation matrix were consistent with a single factor model: 4.966, 0.912, 0.714, 0.602, 0.547, 0.341, 0.253, and 0.250. Additionally, the rotated two-factor solution was characterized by a number of substantial cross-loadings, and a nontrivial factor correlation (r = −.699), which suggest an overextracted solution. Based on these findings, the aforementioned unidimensional multigroup IRT model was examined.
Initial Alignment Results
Results from the aligned multigroup normal-ogive IRT model are provided in Table 3. Pooled estimates across the invariant groups (weighted by group size) are provided in the rightmost column, and group-specific estimates are provided in the remaining columns. Analyses revealed that 8 of the 9 factor loadings (λ) did not differ across groups, which suggests that for these items the strength of the relationship between the latent trait and observed responses for each item were consistent across gender/sex and generation. Closer examination of the pooled coefficients indicates that Items 1, 4, and 6, exhibited the strongest loadings, whereas Items 2, 5, and 9 were the least reliable. The factor loading for Item 7 (“I feel that being gay/lesbian/bisexual is a sin”) was identified as differentially functioning for the Female Millennial group based on a multiple-comparison corrected procedure, and the loading for this group appeared to be significantly weaker than the pooled estimate, suggesting that this item was less strongly related to the underlying IH trait. Item trace plots and information curves are provided in the supplementary materials available online.
Turning to the threshold parameters, DIF was observed in one or more groups for Items 2, 3, 5, 6, 7, and 9, whereas the remaining items exhibited no meaningful threshold differences across groups. For Item 2 (“I feel uneasy around people who are very open in public about being gay/lesbian/bisexual”), the Female Millennial group exhibited significantly higher values across all three thresholds. This pattern of findings suggests that after accounting for cross-group differences in latent IH as well as the reverse-coding of this item, were more likely to select the middle and highest categories (i.e., mixed feelings, disagree, strongly disagree), over the two lowest categories (i.e., agree, strongly agree). A similar pattern was observed for Female Xers, but only the first threshold (τ1) emerged as significant, indicating that this group was more likely to strongly disagree with this item. Turning to Item 3 (“I often feel ashamed that I am gay/lesbian/bisexual”), Female Millennial participants were slightly more likely to strongly disagree, relative to members of the other gender/sex-generation cohorts. For Item 5 (“I worry a lot about what others think about my being gay/lesbian/bisexual”), the Male and Female Millennial groups exhibited significantly lower values for all threshold parameters, which is consistent with a higher likelihood of selecting the agree or strongly agree response option. Male Millennial respondents also exhibited a significantly higher first threshold for Item 6 (“I feel proud that I am GLB” [R], suggesting that members of this group were more likely to strongly agree with this item. In addition to the factor loading DIF for Item 7, the Female Millennial group showed a significantly lower threshold, which is consistent with a higher probability of selecting strongly disagree. Finally, Male members of the Boomer and Xer cohorts exhibited significantly lower second threshold (τ2) parameters for Item 9 (“I feel that being gay/lesbian/bisexual is a gift” [R]) relative to other groups, and because of the reverse-coding for this item, older men were less likely to endorse the category of this item. In contrast, Female Xer and Millennial participants exhibited lower values for Item 9 τ1. This pattern of threshold differences suggests that Female respondents were more likely to agree over strongly agree, whereas Male respondents had a higher likelihood of selecting agree over mixed feelings.
The alignment analysis also provides tests inflation-protected tests for latent factor mean differences across groups, which provide insight into the degree to which the average level of latent IH varies across gender/sex and generational cohorts, while accounting for DIF in the aforementioned items and groups. In the present analysis, the Male Millennial group was designated as the reference category (αMale Mil.@0; ψMale Mil.@1), so the latent factor means of the remaining groups are estimated relative to this group. Analysis of the aligned solution revealed that both members of the Boomer cohort, and Female Xers exhibited significantly lower latent factor means (αMaleBoom. = −0.468, 95% CI [−0.668, −0.283], d = .47; αFemaleBoom. = −0.450, 95% CI [−0.656, −0.278], d = .45; αFemaleXer. = −0.352, 95% CI [−0.508, −0.222], d = .35), relative to all other groups. Additionally, individuals in the Male Xer and Female Millennial groups (αMaleXer. = −0.316, 95% CI [−0.476, −0.182], d = .32; αFemaleMil. = −0.165, 95% CI [−0.314, −0.039], d = .17) had lower latent means than the older cohort and Female Xers, but were significantly higher in latent IH than Male Millennials.
Bayesian Alignment Within CFA Results
Consistent with Marsh et al.’s (2018) recommendations, the BAwC model was identified by constraining the factor loading and thresholds of an invariant indicator (Item 1) to the estimated value from the robust ML alignment analysis model. Following some initial tuning, 2 posterior distributions for the baseline BAwC model were simulated from 20,000 draws from two MCMC chains running a Gibbs sampler using 10% thinning. As expected, median values for all posterior distributions were identical (within 2 decimal places) to the corresponding point estimates obtained in the initial ML alignment model. In contrast to the initial analysis, which provided no feedback regarding the appropriateness of the specified model, the PPC statistics generated by the BAwC revealed that the baseline model did a poor job of replicating the observed sample data, PPC 95% CI [167.241, 335.733], p < .001. Given that the BAwC imposes no cross-group constraints on model parameters, the model misfit may arise through a violation of the conditional local independence assumption imposed by the standard IRT specification. In other words, after accounting for the endorsement likelihood attributable to the influence of the latent trait, endorsement of some item pairs may be more likely than others. The resulting unmodeled dependence may be one reason that the initial aligned solution failed to adequately replicate the sample data.
Following the procedure outlined by B. O. Muthén and Asparouhov (2012), 36 covariances among conditional item responses were specified for each group (216 total). Priors for the covariance parameters were specified to follow an Inverse Wishart distribution with (IW[0,15]). Posterior parameter distributions for this conditional BAwC model were simulated from 20,000 draws from two MCMC chains using 10% thinning. The location and form of the posteriors were very similar to the baseline BAwC model, but unlike baseline, PPC statistics suggest that this follow-up model did an excellent job of replicating the observed sample data: PPC 95% CI [−73.650, 67.118], p = .525 (complete results provided in the online supplemental materials).3,4
Table 4 provides formal comparisons for each of the parameters and groups identified as invariant in the initial alignment analysis. Estimates listed in Table 4 describe the posterior median and 95% credible intervals for the differentially functioning group, item and parameter (DIF parameter), along with the pooled estimate of the same parameter among the groups identified as invariant (Pooled-invariant parameter), and the difference between the DIF and pooled parameters (Difference). As expected, the posterior median values described in the first two columns are very similar to the corresponding ML point estimates provided in Table 3. Estimates listed in the final column of Table 4 provide a formal evaluation of the DIF suggested by the initial alignment analysis, after accounting for the violation of the local independence assumption. None of the credible intervals for the Difference posteriors contained 0, suggesting that all items and parameters identified as differentially functioning by the initial alignment analysis can be considered reliably different.
Posterior Median and 95% Credible Intervals for DIF, Pooled Invariant, and Difference Parameters.
Note. DIF = differential item functioning; DIF group estimates = posterior median for group, item, and parameter identified as differentially functioning; pooled-invariant parameter = posterior median for invariant group, item, and parameter.
Discussion
The present study examined differences in the relative centrality and endorsement of aspects of IH as a function of respondents’ generation and gender identity. This form of psychometric analysis is important because it can identify differences in how participants construe the items used to measure a construct, which may also provide insight into how individuals of different generations and gender identity may conceptualize the corresponding experience. Although we did not offer any formal a priori hypotheses, prior literature suggested that some aspects of IH would differ in meaning for LGB individuals who belong to different gender/sex and generational groups, due at least in part to changing attitudes regarding LGB culture over time and different experiences and treatment of lesbians and gay men in U.S. culture (e.g., Beals & Peplau, 2005; Brooks, 1981). Following the procedures outlined by Flake and McCoach (2018), as well as Marsh et al. (2018), IRT–DIF analysis found evidence of credible cross-group differences for one item factor loading, as well as threshold DIF for six items from the Wright and collegues IH scale (Wright et al., 1999; Wright & Perry, 2006). The remaining three items exhibited invariance of both factor loading and threshold parameters, and the alignment analysis provided estimates of group-specific latent factor means adjusted to account for the observed DIF.
The presence of significant latent mean differences in IH suggests the levels of the invariant components of this construct are lowest among members of the oldest generation (Boomers) and highest among the youngest generation (Millennials). Our attempt to explain this finding should be tempered by the absence of a priori hypotheses regarding any specific pattern of latent mean differences, as well as the cross sectional design of the study. However, this pattern of findings is consistent with the idea that older individuals may evaluate and respond to day-to-day stressors associated with their sexual minority status differently than younger individuals, which may explain differences in IH. For example, older LGB individuals may be more likely to have fully integrated their sexual identity into their self-concept, and in turn, may have developed better coping strategies to counter ongoing threats to this aspect of their identity (e.g., Berg et al., 2015). In contrast, the emerging adults in this study may possess a less cohesive sense of self, and they have had less time to integrate their sexual identity with the other aspects of their self-concept. As a result, they may be more vulnerable to ongoing stressors associated with their identity, which may lead them to internalize disparaging or threating signals from their external environment (e.g., Cox, Dewaele, van Houtte, & Vincke, 2011).
Although the observed latent mean differences are inherently compelling, these findings only speak to the component of this construct that was invariant across groups. Much of the utility in the present work comes from examining the content of the items that were found to be noninvariant, or differentially functioning, across gender/sex and generation. These findings provide insight into how members of different groups relate to specific aspects of the construct, accounting for any potential differences at the latent level. For example, younger female respondents were more likely to disagree with statements referring to “shame” (Item 3) and “sin” (Item 7) regarding their sexual identity. Additionally, both Xer and Millennial female respondents were more likely to disagree with a statement describing unease about others being open about their identity in public (Item 2), whereas younger male participants were more likely to agree with a statement expressing pride over their sexual identity (Item 6). All Millennial respondents were more likely to agree with a statement expressing concern over others’ evaluation of their sexual identity (Item 5).
Broadly, these results indicate that female and male identifying individuals with similar levels of latent IH endorse items on the IH scale differently, relative to both the centrality of the statement to their IH and to the strength at which they endorse the item. This is consistent with suggestions that lesbians experience oppression as women and as nonheterosexual (i.e., sexual minority), whereas gay men may experience privilege as men even while experiencing stigma for nonheterosexual identity. Furthermore, sex differences in socialization may affect the way females and males experience IH. The importance of the social, cultural, and political environments on male and female socialization may also lend support to the findings that variability in item endorsement by gender/sex can be influenced by generational cohort.
Implications for Clinical Research and Practice
The presence of partial (rather than full) scalar and metric invariance for the items of the IHS means that the observed composite scores typically used in applied research are a less than perfect reflection of latent differences in levels of the trait across groups. More specifically, finding that threshold parameters for six of the nine items exhibited differential functioning across one or more groups, suggests that components of this construct are more readily endorsed by members of different groups who have the same level of the underlying latent trait (Embretson & Reise, 2000). As a result, researchers should interpret any observed differences based on the nine-item composite, across the groups examined here, because any differences may be partially attributable to the impact of the aforementioned threshold noninvariance rather than differences in the underlying latent trait. Nonetheless, the present analysis did find evidence of significant latent mean differences across generations, suggesting that latent factor means (as defined by the invariant item set) were significantly higher among Xers and Millennials, relative to Boomers. Researchers interested in maximizing the validity of comparisons across these groups are advised to either, utilize the AwC framework illustrated here to construct measurement-aligned multigroup SEMs that account for observed DIF, or to rely on composite scores based on three invariant items, though the latter solution is limited by a decreases in composite reliability that may arise from shortened measures.
The presence of differentially functioning (noninvariant) items has interesting implications for clinical practitioners serving LGB clients experiencing IH by providing insight into the way in which members of different genders/sexes and generations experience IH. For example, the tendency for Male and Female Millennial participants to more readily endorse statements expressing concern over what others think about their sexual identity, suggests that their internalized stigma may be more heavily derived from individual-level interactions, rather than structural sources of oppression. Clinical practitioners may find this type of insight to be informative in their ongoing assessment of the client’s progress and treatment plan.
Limitations and Future Directions
In the present study gender/sex was operationalized as a binary variable, which excluded gender nonconforming, gender fluid, gender queer, and transgender persons. Although a recent study (Bauerband et al., 2019) has begun to address this gap, future work should incorporate broader and more inclusive conceptualizations of gender to provide a more nuanced and complete understanding of the intersection between gender and IH. The present study included exclusively same-sex attracted and bisexual identifying individuals, which is appropriate given that all items refer to an individual’s “gay/lesbian/bisexual” identity. Although a substantial percentage of individuals in our sample identified as bisexual (13%), there were not a sufficient number of bisexual individuals in the male and older cohorts to investigate this attribute as a potential grouping factor. However, recent work examining internalized stigma associated with bisexuality (Balsam & Mohr, 2007; Brewster & Moradi, 2010; Dyar et al., 2019; Dyar, Feinstein, & London, 2015) underscores the importance of acknowledging this dimension of diversity within the larger umbrella of gender and sexual minority research.
The generational groupings applied in the present work were based on designations provided by the U.S. Census Bureau, but the nominal classification scheme results in a quasiarbitrary allocation for individuals born near the cutoffs (i.e., those born in 1964 were Boomers, 1965 were Xers). At the present time, MI analysis cannot be performed using continuous moderator variable, but forthcoming advances in statistical modeling may allow of this form of MI analysis in the future. Although our overall sample size was substantial, the findings are still somewhat limited by the unbalanced group allocation. It should be acknowledged that no DIF was detected for the gender/sex-generation group with the smallest sample size (Female Boomers, n = 327), whereas the group with the largest sample size (Female Millennials, n = 874) exhibited DIF on four of nine items. However, it is also worth noting that DIF was detected for one item in both the second-smallest group (Male Boomers, n = 336), and second-largest groups (Female Xers, n = 840).The application of Marsh et al. (2018) AwC model to include a fully saturated residual covariance structure into our follow-up analysis provides additional support for the validity of our findings. However, it is worth noting that the inadequate fit of the initial BAwC model (based on the 95% PPC interval not containing 0), may have arisen from (unmodeled) nonlinear factor loading relationships, and future studies should examine this possibility.
The cross-sectional nature of our design resulted in a confounding of age and chronological time. More specifically, the differences attributed to generation in the present study likely reflect the combined influence of generational and stage-of-life factors. Future studies could overcome this limitation by collecting responses from the same individuals across the life span and performing mixed (between and within) MI analysis. White individuals were overrepresented in the present sample, which limits generalizability. In the future, researchers should collect more racially/ethnically diverse samples of sexual minorities, which will allow for a more precise examination of how intersecting minority group identities might uniquely impact the experience of IH.
Supplemental Material
R1_Supplemental_Materials – Supplemental material for Gender and Generational Differences in the Internalized Homophobia Questionnaire: An Alignment IRT Analysis
Supplemental material, R1_Supplemental_Materials for Gender and Generational Differences in the Internalized Homophobia Questionnaire: An Alignment IRT Analysis by Robert E. Wickham, Renee Gutierrez, Brenna L. Giordano, Sharon S. Rostosky and Ellen D. B. Riggle in Assessment
Footnotes
Acknowledgements
We thank Robert Wickham, who was responsible for the development of the research question, introduction and discussion, and conducted the statistical analyses. A special thanks to Brenna Giordano and Renee Gutierrez, who were responsible for the data cleaning, preparation and reporting, and contributed to the literature review. We also thank Sharon Rostosky and Ellen Riggle, who designed and collected the studies that generated these data, and contributed to the literature review and discussion.
Authors’ Note
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: The data collected for Samples 2 and 3 was supported in part by a grant from the American Psychological Foundations’s Wayne F. Placek Award and the University of Kentucky’s Center for Drug and Alcohol Research.
Supplemental Material
Supplemental material for this article is available online.
Notes
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
