Abstract
The purpose of the current study was to assess the gender invariance of an a priori four-factor solution of behavioral consequences of drinking. Results evidenced strong partial measurement invariance, with marginal structural invariance, which signals that the underlying constructs possessed the same theoretical structure for both men and women.
Alcohol is one of the first drugs of choice for young adults, and alcohol use disorder constitutes one of the greatest health care issues within the United States (Karoll, 2002). Student consumption of alcohol is an ongoing concern for many within higher education. Researchers have sought to understand not only the academic ramifications of alcohol use but also the second-hand sociological effects of alcohol misuse, as well as factors that affect students’ binge drinking. Despite the breadth of such studies, they are not without limitations. For instance, although a preponderance of student drinking research focuses on behavioral consequences, data analytic approaches to understanding measurement in this area tend to be purely descriptive and do not address more precise factors that make up the larger constructs of behavioral consequences of alcohol use.
General consequences of student drinking are both serious and broad. Potential immediate consequences of binge drinking include premature death (e.g., from cirrhosis of the liver), suicide, increased risk for HIV infection, antisocial behaviors, and school-related difficulties. Over time, some long-term consequences of sustained and consistent binge drinking could include unmet developmental tasks, chronic unemployment/underemployment, failed interpersonal relationships, and dysfunctional developmental transitions (Kuo et al., 2002; Sullivan & Risler, 2002; Wechsler, Lee, Kuo, et al., 2002; Wechsler, Lee, Nelson, & Kuo, 2002).
Research indicates that students experience numerous behavioral consequences resulting from their drinking habits. The most commonly reported behavioral consequences are hangovers, driving while intoxicated, DUI arrests, performing poorly on a test or major project, skipping class, and becoming sick because of drinking behavior (Kuo et al., 2002; Sullivan & Risler, 2002; Wechsler, Lee, Kuo, et al., 2002; Wechsler, Lee, Nelson, et al., 2002). Despite a breadth of research that reports on binge drinkers’ experienced behavioral consequences (e.g., Kuo et al., 2002; Sullivan & Risler, 2002; Wechsler, Lee, Kuo, et al., 2002; Wechsler, Lee, Nelson, et al., 2002), an understanding of the deeper nature and effects of such consequences remains scantly charted. Although one study sought to validate a measure of behavioral consequences of drinking (Arriola et al., 2009), the authors reported on respondents’ expectations of what they would experience after alcohol consumption, which is unlike most other studies utilizing measures of behavioral consequences that comprise self-reports of behavioral experiences.
Another recent study investigated the factor structure of behavioral consequence of drinking, using data gathered from community college students (Derby & Smith, 2008). Supported by the research was a two-factor structure consisting of (a) personal consequences and (b) social consequences. The validity of this two-factor structure was later examined by Derby and Smith (2010) with data from 4-year university/college students using the 2001 Harvard School of Public Health College Alcohol Study (Wechsler, 2005). The authors found that the two-factor structure (which had described the community college student population) did not adequately explain the 4-year university/college student population, and they instead posited and cross-validated a four-factor model of behavioral consequences comprising Personal, Social, Academic, and Sexual consequences.
One notable gap in the literature is attentiveness to potential gender differences with regard to known behavioral consequences factor structures. When considering a globally fitted factor model, a hypothesis for a lack of factorial invariance between genders might be posited for several reasons. First, men and women perceive drinking and drinking problems differently (Karoll, 2002; Wechsler & Kuo, 2000). Second, the physiological effects of alcohol differ between men and women (Collins & McNair, 2002; Karoll, 2002; Mulligan & Bryant, 2000). Given empirical psychosocial and physiological differences between genders, assuming a global factorial model common to both genders seems unlikely to constitute a best practice approach (Karoll, 2002; Weitzman & Kawachi, 2000). Invariance testing is an examination of the extent to which score properties and interpretations are generalizable across population groups, settings, and tasks (Messick, 1995). Within the current study, invariance testing would illuminate if these constructs generalize to both men and women, or if the constructs are limited to only one of the population subgroups. Such an examination of potential invariance, or lack thereof, could assist counseling personnel with targeting developmental programs and initiatives.
The purpose of the current study was to assess the invariance of a hypothesized four-factor, behavioral consequences model of drinking among undergraduate college men and women in the United States, using a large, nationally representative sample. The aim of the current study is to investigate the validity of constructs resulting from data generalizable to 4-year undergraduate students. As such, the aim is not to provide conclusive validity evidence for either the constructs under study or for all the data resulting from the College Alcohol Study instrument.
Method
The present study utilized data from the 2001 Harvard School of Public Health College Alcohol Study (Wechsler, 2005). These data consist of a random sample of 10,904 surveyed full-time undergraduate students from 119 four-year colleges and universities in the United States. The mean age (based on weighted estimates) was 20.7 years, with 55.4% of the respondents identifying themselves as female. The survey focused on issues surrounding the use/abuse of alcohol as well as other high-risk behaviors (e.g., drug use, sexual activity). The initial analyses (Derby & Smith, 2010) were carried out using two randomly selected subsample of the data consisting of approximately one third of the total number of cases, with the second sample reserved for subsequent exploratory and cross-validation purposes, if the need arose (i.e., if the hypothesized two-factor structure did not adequately explain the data). In the present study, we used the third, reserved random sample of data (n = 2,872) to assess the measurement and structural invariance with the confirmed four factor model. Cases with missing data were omitted from the analysis.
Students responded to questions related to alcohol use and other high-risk behaviors. For the present study, we considered student responses to 12 queries about the frequency of occurrence for specific behavioral consequences of alcohol drinking (see Table 1). Response options for each behavioral consequence comprised a frequency rating scale 1 = not at all, 2 = once, 3 = twice, 4 = 3 times, and 5 = 4+ times. Because the number of categorical response categories exceeded four, robust maximum likelihood estimation was used to estimate the models (as per the recommendations of Rhemtulla, Brosseau-Liard, & Savalei, 2012) and the Satorra–Bentler SBχ2 statistic computed. Model invariance was assessed using the Satorra–Bentler scaled chi-square difference test (SBSΔχ2; Satorra & Bentler, 2001) and the change in comparative fit index (ΔCFI) model fit statistics, wherein “a negative ΔCFI lower than −.01” and significant Δχ2 “indicated a lack of model invariance” (Dimitrov, 2010, p. 127; see also Cheung & Rensvold, 2002). Because of the influence of sample size on the chi square statistic, where large samples may induce spurious statistical significance of SBSΔχ2 (see Brannick, 1995; Kelloway, 1995) an emphasis was placed on the value of ΔCFI. In the event of model misspecification, the modification indices assisted in respecifying the model, and a threshold of 10.0 was employed for these values. Because the aim of this study is to increase the practical utility of this type of data within various counseling situations, the authors adhered to Dimitrov’s suggestion that model modifications should involve less than 20% of the parameters. Criteria used for fit of the obtained models were CFI ≥ .95, TLI (Tucker–Lewis coefficient) ≥ .95, and root mean square error of approximation (RMSEA)]≤ .06 (Hu & Bentler, 1999). Data analysis transpired using Mplus 7 software.
Items (and Associated Factors) on the Behavioral Consequences of Drinking Scale.
This item did not load onto any of the a priori four factors and was omitted from the model.
Results
For the present study, we assessed configural invariance by fitting the four-factor model (see Figure 1) posited by Derby and Smith (2010) separately to the data from the males and the data from the females. This model consisted of four correlated behavioral factors: Personal, Social, Academic, and Sexual consequences of drinking. We then assessed measurement invariance by comparing successively constrained models.

Four correlated factor model of behavioral consequences of drinking for 4-year students.
Configural Invariance
When the four-factor model was fitted separately to the samples of men and women, the results showed relatively similar fit. The SBχ2 values for both men and women were statistically significant; however, this was likely because of the large sample sizes. For men, TLI exceeded the .95 threshold, whereas for women it approached this value. However, both men and women exceeded the CFI .95 threshold, and additionally exhibited RMSEA values lower than the .06 threshold. Thus, these results indicate that the models exhibited configural invariance. Table 2 reports the results for these analyses. For greater understanding of the factor analytic models for men and women, Table 3 provides the standardized factor loadings and robust standard errors for these models.
Fit Indices for Four-Factor Model Fitted to Male and Female Samples.
Note. SB = Satorra–Bentler test; CI = confidence interval; LL = lower limit; UL = upper limit; TLI = Tucker–Lewis index; RMSEA = root mean square error of approximation; SRMR = standardized root mean square residual; CFI = comparative fit index.
Standardized Factor Loadings, Robust Standard Errors, Item Reliability, and Construct Reliability by Gender.
Measurement Invariance
To assess measurement invariance (metric, scalar, and item uniqueness), we compared successively constrained models. First, evaluation of metric invariance occurred by comparing an unconstrained model (Model 0) to a model constraining all factor loadings to be equal (Model 1). Results for this comparison indicated a lack of invariance (SBSΔχ2(7) = 284.90, p < .001). However, for this comparison, it was possible that the large sample size may have induced spurious statistical significance. Therefore, for this comparison (and all subsequent comparisons), the change in CFI was evaluated as an indicator of invariance. This change in CFI (ΔCFI = −.018) also indicated a lack of metric invariance. Using the modification indices as a guide, the constrained model was thus revised (Model 1P) by freeing the loading for the item related to damaging property (Item 6 in Table 1), which led to a conclusion of partial metric invariance (with the exception of one factor loading) based on the change in CFI (ΔCFI = −.003).
Next, to examine scalar invariance, we compared the model with partial metric invariance (Model 1P) to a more constrained model with equal item intercepts (Model 2). This comparison yielded a ΔCFI = −.004, suggesting scalar invariance. Finally, assessment of uniqueness invariance occurred by comparing the model with both partial metric and scalar invariance (Model 2) to a more constrained model with equal item variances/covariances (Model 3). This comparison indicated a lack of invariance (ΔCFI = −.102). Using the modification indices as a guide, it was necessary to sequentially free 6 (55%) of the 11 item residuals (associated with Items 2, 3, 4, 6, 7, and 11 in Table 1) to obtain evidence for partial uniqueness invariance (ΔCFI = −.008). However, because Model 3P exceeded the recommended threshold for freeing parameters by nearly three times the “acceptable” rate for practical application (Dimitrov, 2010), we concluded that the four factor model exhibited strong, partial measurement invariance, with all but one factor loading being invariant and all intercepts being invariant.
Structural Invariance
Several seminal studies consider testing for uniqueness invariance an “overly restrictive” data assessment (Dimitrov, 2010, p. 128; see also Bentler, 2004; Byrne, 1988). Because testing for uniqueness invariance is inherently nested under Model 2 (scalar invariance), and because the four factor model exhibited strong partial invariance, we next used Model 2 as the baseline to test the four-factor model for structural invariance. A comparison of a model that posited partially invariant factor loadings, invariant intercepts (Model 2) to a model with invariant factor variances/covariances (Model 4) provided marginal evidence for structural invariance (ΔCFI = −.010). Tables 4 and 5 provide the statistics for the measurement invariance tests.
Measurement Invariance Statistics.
Note. SB = Satorra–Bentler test; TLI = Tucker–Lewis index; CFI = comparative fit index; RMSEA = root mean square error of approximation. Model 1P had 1/11 = 9% of factor loadings freely estimated. Model 3P had 6/11 = 55% of residual variances freely estimated.
Model Comparison Statistics.
Note. SB = Satorra–Bentler scaled chi-square difference test; CFI = comparative fit index.
Discussion
The purpose of the present study was to investigate gender invariance of an a priori four-factor model of behavioral consequences from drinking previously found by Derby and Smith (2010). A clear understanding of college students’ perceptions of the behavioral consequences of drinking, what structure these perceptions take, and how these perceptions relate to each other, is important to developing interventions and educational initiatives aimed at addressing these consequences. Results of the four-factor solution suggest that it may be important for these interventions and educational initiatives, at least at the 4-year college level, to take care to distinguish sexual and academic consequences from other personal and social consequences of drinking, as students appear to perceive them distinctly. That is, because these consequences appear distinctly perceived, initiatives that target these specific consequences may “imprint” more deeply and work more effectively to reduce the potential adverse outcomes.
The failure to disconfirm the four-factor solution of behavioral consequences (Derby & Smith, 2010) allowed for an investigation of gender invariance of the model. In other words, we examined the four-factor solution to understand if the underlying constructs possessed the same theoretical structure for both men and women. Given the differing psychosocial and physiological differences observed between men and women with respect to consuming alcohol (Karoll, 2002), the authors anticipated that they might find a variant model that could spur the development of new models for both genders. However, this was not entirely the case as the resulting analysis indicated that the four-factor behavioral consequences model exhibited strong partial measurement invariance, fulfilling the criteria for both metric and scalar invariance. This result corroborates the results found when separately fitting models for each gender, which suggested the four-factor model fit relatively well for each gender. Thus, the configural invariance for the four-factor model evidenced in this study indicates that the patterns of free and fixed model parameters are equivalent for men and women.
In the next steps of model testing, partial metric invariance, scalar invariance, and a lack of item uniqueness invariance indicated that the model possessed strong partial measurement invariance. Evidence of partial metric invariance indicated that, with the exception of the freed factor loading for the variable “damaging property,” the relationship between the indicators and the latent factors are comparable across men and women. In addition, the indication of scalar invariance suggests that item bias may not be evident between men and women. Although a more formal DIF analysis would be necessary to substantiate this indication, the results of the present study are encouraging in this regard. However, the present study did provide evidence against strict measurement invariance, because of a lack of item uniqueness invariance (i.e., error term invariance), suggesting that many of the items may have been measured with differing precision for male and females. This finding indicates that the unique variance for items was not accounted for by their respective latent construct (i.e., was not common variance).
In addition, the lack of item uniqueness, which resulted in strong partial measurement invariance (instead of strict measurement invariance) indicated differential degrees of reliability. 1 For instance, as reported in Table 3, items within the personal and academic consequences constructs exhibited relatively similar reliabilities at the item and construct levels. However, differential reliability appears evident for the social consequences construct, where the data for men were more reliable than women’s data. In addition, the data for women were more reliable than men’s data for the sexual consequences construct.
As previously noted, however, the test for item uniqueness invariance has been posited by some as an overly restrictive test, and may be indicative of greater precision than warranted for the behavioral measures signaled here. In addition, a number of researchers (e.g., Byrne, Shavelson, & Muthén, 1989; Marsh & Hocevar, 1985) have suggested that, if noninvariant loadings comprise only a small proportion of the total, then meaningful between-group comparisons may still be made. From this, we conclude that making meaningful between-group comparisons for men and women is appropriate given that we found only one noninvariant factor loading (Item 6: “damage property”).
Finally, testing for structural invariance was important given the authors’ belief that the correlational relationships and variability of the latent constructs are relevant for generalizability. Thus, given that the model possessed the prerequisite weak measurement invariance, the test for structural invariance that we conducted would be considered appropriate (see Dimitrov, 2010, for more information). Structural invariance testing indicated marginal invariance when compared against the strong, partial measurement model (Model 2). Marginal structural invariance indicates that, although there is not conclusive evidence against structural invariance, the factor variances and or covariances may possess differing degrees of dispersion between males and females (i.e., exhibit heterogeneity by gender). A finding of between gender heterogeneity appears to corroborate the finding of a lack of item uniqueness invariance, and could potentially signal the presence of differential affects across groups (see Bryk & Raudenbush, 1988). However, the findings of this study are inconclusive on this matter and further research and testing are required.
Limitations
One limitation of the current study is the self-reported nature of the data. Because the data were self-reported, they could contain certain inherent biases, although the anonymous nature of responses may mitigate such biases. A second limitation concerns the age of the data, which were collected in 2001. However, few if any other large-scale studies of behavioral consequences of drinking are available, and the data from the Harvard study consist of high-quality data employing random sampling. More important, given the lack of existing research on gender invariance in behavioral consequences of drinking, these data provide an important baseline for examining invariance in the structure of these behavioral consequences—a baseline that does not currently exist and against which future studies might compare. Another limitation might be that the items utilized within this study may not contain of the full range of items necessary to assess adequately the behavioral consequences of alcohol use in general. This may be particularly relevant for women, as the screening and detection of alcohol use issues have been reported as highly problematic (Karoll, 2002). Nonetheless, the data again provide an important baseline and future studies might consider additional indicators.
Finally, although we employed estimation methods (MLM estimation) intended to mitigate the impact of multivariate nonnormality, no existing research has indicated how such nonnormality might affect inferences based on differences in CFI values. Cheung and Rensvold (2002) provide Monte Carlo simulations that support the use of a negative ΔCFI value lower than −.01 as a criterion, but the simulations transpired under conditions of multivariate normality and, in the absence of similar simulations positing nonnormality, it is unknown how nonnormality might affect inferences made using this criterion.
Future Research
Two areas of future research are important. First, future research should demarcate counseling thresholds for the personal, social, academic, and sexual consequence factors within the current study to improve diagnostic and treatment efforts. For women specifically, who tend to seek counseling or psychiatric assistance for drinking instead of drug or alcohol abuse counseling (see Karoll, 2002, for more information), mental health professions should consider being on the front end of this movement.
Second, alcohol consumption represents a social convention predicated on a myriad of belief and perceptual structures. Examination of how these relate to the structures of behavioral consequences is necessary to gain a holistic view of the impact of drinking. For example, are variables or structures concerning peer valuations of drinking or of students’ alcohol expectancies related to the type or frequency of differing behavioral consequences?
Finally, invariance testing should occur across other known groups that might possess differing perceptions of alcohol. For example, one could examine invariance through the lens of group membership within a fraternity/sorority. The breadth of research on alcohol consumption and Greek membership, and the noted alcohol-related perceptual differences between fraternity/sorority members and nonmembers may suggest distinct theoretical structures for the underlying constructs. Testing of this hypothesis should occur in future research initiatives.
Conclusions
Because there are relatively few studies that examine the factor structure of this type of data, the current study represents an important preliminary investigation into the potential gender invariance concerning the factor structure of behavioral consequences. For gender, at least, the current investigation failed to disconfirm equivalent factor structures for men and women, meaning that the underlying constructs for men and women possessed the same theoretical structure. However, evidence from this study suggests that there may be varying levels of precision for the measurements across male and female groups.
The significance of this research is a confirmation of the measurement accuracy of the items across male and female groups. One of a surveyor’s primary tasks is the mitigation of the four sources of survey error: sampling, coverage, measurement, and nonresponse error (see Dillman, Smyth, & Christian, 2009). High rates of measurement error chip away at the validity and integrity of any resulting data, rending them undesirable for decision making. Findings from the current study underscore that the data resulting from the measurement items used are valid measures of behavioral consequences of drinking for both men and women and are appropriate for group comparisons with populations similar to the one examined herein. These findings should inspire mental health and substance abuse counselors to have greater confidence in the data collected from this behavioral consequences subscale.
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
