Abstract
The Multidimensional Behavioral Health Screen (MBHS) is a brief screening measure of behavioral health symptoms. Although the measure was first developed for primary care, it is likely to have clinical utility in other settings. This study examined the MBHS’s factor structure and psychometric properties with a university undergraduate and graduate student sample (n = 602, 58.6% female, 75.9% White, primarily aged 20–24) during the COVID-19 pandemic. MBHS subscale scores demonstrated internal consistency reliability and both convergent and discriminant relations with external, criterion variables. Confirmatory factor analyses supported a 7-subscale factor structure of the MBHS and did not find evidence of higher order factors. Clinical and theoretical implications, as well as future research directions, are discussed.
Keywords
Routine, population-based screening is essential for achieving integrative health care that bridges the gap between physical and behavioral health care providers. The Multidimensional Behavioral Health Screen (MBHS; McCord, 2020) is a screening measure of behavioral health symptoms intended for clinical use in primary care settings. The foundation of the MBHS is the Minnesota Multiphasic Personality Inventory-2-Restructured Form (MMPI-2-RF; Ben-Porath & Tellegen, 2008), which reflects the recent paradigm shift to a hierarchical-dimensional, as opposed to a categorical, conceptualization of psychopathology (e.g., Forbush & Watson, 2013; Kotov et al., 2018; Krueger & Markon, 2006). The MMPI-3 (Ben-Porath & Tellegen, 2020), which now supersedes the MMPI-2-RF, furthers this movement toward a hierarchical-dimensional model of assessment. The MBHS is made up of 27 items on nine scales that reflect the most common problems in the somatic/cognitive, internalizing, and externalizing domains: Somatization (SOM), Demoralization (DEM), Anhedonia (ANH), Anxiety (ANX), Suicidal Tendencies (SUI), Cognitive Issues (COG), Activation (ACT), Disconstraint (DSC), and Substance Misuse (SUB). The initial validation study by McCord (2020) indicated that the measure’s subscale scores demonstrated internal consistency reliability (α levels ranging from .61 to .81), convergent and discriminant relations with the expected MMPI-2-RF scales, as well as classification accuracy vis-ä -vis these same scales using a T-score cutoff of 65 or higher.
The MBHS was developed using a clinical sample from a single medical facility, limiting generalizability. Although McCord’s (2020) original study used college student samples to collect psychometric data for the preliminary version of the MBHS, these data have not been reported. The present study therefore extends examination of the measure’s convergent and discriminant relations with other measures using a college student sample. Furthermore, the factor structure of the MBHS has not yet been empirically validated and therefore the present study adds to the literature by testing its implied structure. Next, examinations of how the MBHS correlates with clinical variables have been limited thus far, focusing only on associations with the Patient Health Questionnaire-9 (PHQ-9; Kroenke et al., 2001; Mitchell et al., 2020) and the MMPI-2-RF (McCord, 2020). To establish the clinical utility of the MBHS, it is critical to examine its associations with other widely used measures of common, clinically and conceptually relevant mental health constructs. Finally, the present study is novel in its inclusion of strength-based measures and examination of how these measures relate to the MBHS. Assessment of strengths adds an important contribution to the holistic assessment of an individual’s well-being in clinical settings (Rashid & Ostermann, 2009; Seligman et al., 2006).
The primary purpose of the present study was to examine the psychometric properties of the MBHS with a university sample of undergraduate and graduate students. Although the MBHS is intended for use in primary care, validation studies should examine the usefulness of this measure in other settings, such as university health centers, which serve as the de facto primary care for many college students. The MBHS may be particularly useful in university health centers to facilitate screening of psychological problems among college students and, subsequently, appropriate referrals to college counseling centers. That is, screening instruments must be quick to administer and interpret and could result in referrals to providers that may conduct a more thorough assessment (e.g., a university counseling center). Indeed, validation studies on the MBHS have already begun to expand outside primary care. For example, Dodge (2022) utilized a sample of 551 college students to examine the ability of the MBHS 2.0, which is not yet publicly available, to predict suicide risk and found significant associations between classification levels on the MBHS 2.0 and suicide risk levels determined via clinical interview. In the present study, we collected data in a university setting as part of a larger study examining psychological distress in the context of the COVID-19 pandemic. We therefore included the MBHS as an opportunity to collect a validation sample during a real-world scenario.
Our aims for the present study were as follows: (a) to assess the factor structure of the MBHS in a university sample, (b) to determine the internal consistency reliability of the measure’s subscale scores, and (c) to evaluate the convergent and discriminant relations of the MBHS subscales with criterion measures. Finally, we sought to assess model fit of a three-factor higher order structure of the MBHS in an exploratory manner. Although McCord (2020) did not explicitly propose a three-factor solution, the MBHS was designed based on the MMPI-2-RF, which is hierarchically arranged and consists of five domains: Somatic/Cognitive Dysfunction, Emotional/Internalizing Dysfunction, Thought Dysfunction, Behavioral/Externalizing Dysfunction, and Interpersonal Functioning (Ben-Porath & Tellegen, 2008). The MBHS scales map onto three of these five domains: Somatic/Cognitive Dysfunction (SOM and COG scales), Emotional/Internalizing Dysfunction (DEM, ANH, ANX, and SUI scales), and Behavioral/Externalizing Dysfunction (ACT, DSC, and SUB scales). It is therefore reasonable to suspect that the MBHS may be similarly composed of the higher order factors of the MMPI-2-RF.
Expected convergent relations between MBHS subscales and external, criterion variables were as follows. We hypothesized that the Kessler Psychological Distress Scale (K6; Kessler et al., 2002) would be most strongly related to the Demoralization subscale, as both measure nonspecific psychological distress that is strongly related to internalizing disorder symptomatology (Sellbom et al., 2008). We hypothesized that the PHQ-9, a widely used measure of depressive symptoms based on the Diagnostic and Statistical Manual of Mental Disorders (4th ed.; DSM-IV; American Psychiatric Association [APA], 1994) criteria for major depressive disorder, would be most strongly related to the Demoralization subscale, but also related to the Anhedonia, Somatization, Anxiety, and Cognitive Issues subscales given previous research indicating associations between the PHQ-9 and the MMPI-2-RF demoralization, anhedonia, somatic complaints, anxiety, and cognitive complaints scales (Mitchell et al., 2020). The PHQ-9 was also expected to be related to the Suicidal Tendencies subscale due to its explicit assessment of suicidal ideation, although past research (e.g., Mitchell et al., 2020) suggests lower-than-expected correlations between the PHQ-9 and other measures of suicidality. We expected the PROMIS Anxiety Short Form (SF; Pilkonis et al., 2011), a measure of anxiety-related symptoms, to be most strongly associated with the Anxiety subscale.
Next, we hypothesized that the Insomnia Severity Index (ISI; Morin, 1993) would be most strongly related to the Cognitive Issues and Somatization subscales, as previous research shows associations between insomnia and both cognitive impairment (Fortier-Brochu et al., 2012) and somatic complaints (Zhang et al., 2012), though we also expected that the ISI would be related to other MBHS subscales measuring internalizing symptoms (e.g., Anhedonia, Anxiety) given that insomnia is highly prevalent among those with mood and anxiety disorders (Soehner & Harvey, 2012). We hypothesized that the Alcohol Use Disorders Identification Test (AUDIT; Babor et al., 1992; Saunders et al., 1993) would be most strongly associated with the Substance Misuse subscale. We also expected that the UCLA Loneliness Scale (Version 3; Russell, 1996) would be most strongly associated with the Anhedonia subscale given that depression, of which anhedonia is a primary symptom, has been found to be related to loneliness (Erzen & Çikrikci, 2018). Finally, negative correlations were predicted between the Anxiety subscale and both distress tolerance and resilience, in light of previous research on anxiety and its inverse relation with the ability to cope with stress effectively (e.g., Michel et al., 2016; Smith et al., 2008). No explicit hypotheses were made regarding relations between external variables and the Activation and Disconstraint subscales.
Method
Participants and Procedures
The study sample consisted of 602 undergraduate and graduate students enrolled at a public university in the Midwestern United States during the Spring 2020 semester. In the final sample, most participants identified as female (58.6%), 35.7% identified as male, and 5.32% identified as gender diverse (e.g., nonbinary, other). The age distribution was as follows: 15.6% aged 18 to 19, 47.0% aged 20 to 24, 16.5% aged 25 to 29, 7.14% aged 30 to 34, 4.49% aged 35 to 39, 3.49% aged 40 to 44, 3.16% aged 45 to 49, and 2.66% aged 50 or older. Most participants (75.9%) identified as White, with 8.31% identifying as Asian, 6.48% as Black or African American, 4.98% as some other race (e.g., biracial), and 0.17% as American Indian or Alaska Native. Some participants chose not to disclose their race (4.15%).
Data for the present study were collected in September and October of 2020 as part of a larger study examining levels of psychological distress and risk and protective factors for mental health problems during the COVID-19 pandemic. For the larger parent study, all individuals currently enrolled in an undergraduate or graduate course (n = 30,996) were emailed in March and April 2020 for recruitment. Responses were received from 8,574 individuals, 5,547 of whom provided valid and complete surveys. Of those 5,547 individuals, 2,979 individuals indicated they were willing to be contacted for a follow-up study in September and October of 2020 (i.e., the present study). Invitations to participate in the present study were sent to 1,788 individuals and 706 (39.5%) responded. Data from 93 participants were excluded due to being incomplete (<90% complete). Responses from 11 participants who identified as faculty or staff members enrolled in courses using their tuition benefit were also excluded to make the study sample more representative of the population likely to be seen at university and college health clinics, leaving the final sample size of 602. A proper response rate could not be calculated because the survey was closed when the planned sample size was exceeded (i.e., more people may have responded).
At the time of data collection in Fall 2020, the COVID-19 pandemic was ongoing, therefore the large majority of classes were delivered through remote or hybrid instruction, with very few classes offered in person. Courses offered in person were taught under modified conditions of limited capacity to maintain 6-foot distancing and mandated mask-wearing. Data for the present study were collected entirely online. Prior to completing the survey using QualtricsXM, all participants were provided information on the study (e.g., the study purpose, contact information for the study investigators) and consented to participate. Participants then completed the MBHS, K6, PHQ-9, PROMIS Anxiety SF, AUDIT, ISI, UCLA Loneliness Scale, Distress Tolerance Scale (DTS; Simons & Gaher, 2005), and the Brief Resilience Scale (BRS; Smith et al., 2008), as well as other measures not relevant to the present study.
Participants’ privacy and safety were protected using data collection and storage precautions, as well as by providing all participants with up-to-date information on safety precautions about the pandemic and information for COVID-19 resources from trusted sources (e.g., the Centers for Disease Control and Prevention’s COVID-19 information website). Ethics approval was obtained from the affiliated university’s Institutional Review Board (IRB). We report how we determined our sample size, all data exclusions, all manipulations, and all measures in the study.
Measures
MBHS
The MBHS (McCord, 2020) is a 27-item screening tool of clinically relevant behavioral health symptoms in primary care settings. The MBHS consists of nine scales with three items each, measuring somatization, demoralization, anhedonia, anxiety, suicidality, cognitive issues, activation, disconstraint, and substance abuse. Participants rate each item on a 4-point format (definitely false, mostly false, mostly true, definitely true). To score the measure, scores for the three items that comprise the scale are added together for each scale. Empirically determined cut-points for each scale, as described by McCord (2020), indicate whether the scale is clinically elevated. The psychometric properties of the MBHS have been previously described in this article.
Psychological Distress
The K6 (Kessler et al., 2002) is a brief, six-item screening tool used to measure levels of nonspecific psychological distress and assess the risk for serious mental illness. Participants rate, on a 5-point scale ranging from 1 (none of the time) to 5 (all of the time), how often they have experienced various distressing feelings during the past 30 days. Due to the rapidly evolving nature of the COVID-19 pandemic, instructions were modified and prompted participants to rate these feelings during the previous 7 days. Responses to each item are summed to generate a total score ranging from 6 to 30, with higher scores indicating greater levels of psychological distress. A score of 19 or greater indicates severe levels of psychological distress (Kessler et al., 2003). Scores on the K6 demonstrate internal consistency reliability (Cronbach’s α = .89 for nonclinical and .86–.87 for clinical samples; Kessler et al., 2003; Umucu et al., 2022) and show positive correlations with measures of psychological distress and negative correlations with overall health status and daily functioning (Umucu et al., 2022).
Depressive Symptoms
Depressive symptoms were measured using the PHQ-9 (Kroenke et al., 2001). The PHQ-9 is a nine-item, self-report measure of depressive symptoms corresponding to the DSM-IV (APA, 1994) criteria for major depressive disorder. On a 4-point rating scale ranging from 0 (not at all) to 3 (nearly every day), participants rate the frequency at which they experienced each depressive symptom during the previous 2 weeks. Ratings on each PHQ-9 item are summed to generate a total score ranging from 0 to 27, with higher scores indicating greater severity of depressive symptoms. Scores on the PHQ-9 demonstrate internal consistency reliability (Cronbach’s α = .86–.89) and test–retest reliability in both clinical and nonclinical samples (Beard et al., 2016; Kroenke et al., 2001). Scores on the PHQ-9 have been found to be associated with expected outcomes, including functional status, disability days, doctor office visits, and symptom-related difficulties (Kroenke et al., 2001).
Anxiety Symptoms
The Patient-Reported Outcomes Measurement Information System (PROMIS) profile instruments measure health-related quality of life in seven primary domains, including anxiety symptoms (Cella et al., 2010; Pilkonis et al., 2011). Short forms have been developed to maximize the specificity and efficiency of administration (Cella et al., 2019; Hays et al., 2009). Of these, the eight-item PROMIS Anxiety SF was used in the present study to measure anxiety-related symptoms during the previous 7 days. Each item is scored on a 5-point rating scale ranging from 1 (never) to 5 (always). A total score is calculated by summing ratings on each item, then transforming this total score into a T-score. Higher T-scores indicate greater severity of anxiety-related symptoms. Scores on the PROMIS short-form instruments show internal consistency reliability (Cronbach’s α = .90–.91) and are associated with ratings of general health and quality of life (Cella et al., 2019; Hays et al., 2018).
Alcohol Use Disorder Symptoms
The AUDIT (Babor et al., 1992; Saunders et al., 1993) is a 10-item, self-report screening measure of hazardous and harmful levels of alcohol use developed by the World Health Organization (WHO). The AUDIT assesses current alcohol intake, drinking behaviors, possible dependence, and alcohol-related issues. Each item is rated from 0 to 4 and is summed to generate a total score. Total scores on the AUDIT range from 0 to 40. Scores greater than 8 indicate hazardous or harmful alcohol use. The AUDIT was developed and validated cross-nationally, with scores on the measure demonstrating both internal consistency and test–retest reliability (Babor et al., 1992; Hall et al., 1993; Piccinelli et al., 1997; Saunders et al., 1993). The use of the AUDIT is therefore supported across genders, racial/ethnic groups, and settings (e.g., primary care clinics, college campuses). Scores greater than 8 on the AUDIT have been shown to predict alcohol-related difficulties 2–3 years later (Conigrave et al., 1995).
Insomnia
The ISI (Morin, 1993) is a seven-item self-report measure assessing the nature, severity, and impact of insomnia in the previous 2 weeks. Items are rated on a 5-point scale ranging from 0 (no problem) to 4 (very severe problem) and correspond with diagnostic criteria for insomnia disorder defined by the Diagnostic and Statistical Manual of Mental Disorders (5th ed.; DSM-5; APA, 2013). Specifically, the ISI assesses problems with sleep initiation, sleep maintenance, and/or early morning awakening, as well as satisfaction with one’s sleep, noticeability of sleep problems by others, and distress and daytime impairment caused by sleep problems. Total scores range from 0 to 28, with a score ranging from 0 to 7 indicating absence of insomnia, 8 to 14 indicating subthreshold insomnia, 15 to 21 indicating moderately severe insomnia, and 22 to 28 indicating severe insomnia. Scores on the ISI demonstrate internal consistency reliability (Cronbach’s α = .90 for nonclinical and .91 for clinical samples) and show expected convergent relations (Morin et al., 2011). It is a sensitive outcome measure often used during the treatment of insomnia among patients (Bastien et al., 2001; Morin et al., 2011).
Loneliness
The UCLA Loneliness Scale (Version 3; Russell, 1996) is a 20-item, self-report measure of feelings of loneliness and isolation. Items are rated on a 4-point scale ranging from 1 (never) to 4 (often). After reverse scoring nine items and summing them to generate a total score, total scores range from 20 to 80, with higher scores indicating greater feelings of subjective loneliness. Russell (1996) reported internal consistency reliability (Cronbach’s α = .89–.94) of the measure’s scores, as well as test–retest reliability of scores after 1 year (r = .73), positive correlations with other measures of loneliness, and negative correlations with measures of social support.
Distress Tolerance
The DTS (Simons & Gaher, 2005) is a 15-item measure of perceived (lack of) capacity to tolerate emotional distress. Items are reverse-keyed and rated on a 5-point Likert-type scale ranging from 1 (strongly agree) to 5 (strongly disagree) such that lower scores indicate poorer distress tolerance. One item (Item 6) is reverse-scored. The items are then summed to generate a total score ranging from 15 to 75. The DTS was originally proposed to be composed of a single higher order factor and four subscales: Tolerance (i.e., the ability to tolerate emotional distress), Appraisal (i.e., one’s assessment of emotional situations as acceptable), Absorption (i.e., the level of attention absorbed by negative emotions and relative interference with functioning), and Regulation (i.e., the ability to manage one’s negative emotions). However, recent psychometric analyses suggest that the four subscales are not distinct and recommend using only the total sum score (Rogers et al., 2020). Scores on the DTS have demonstrated internal consistency (α range = .81–.92; Kremyar et al., 2020; Simons & Gaher, 2005) and test–retest reliability after 6 months (r = .61), as well as positive correlations with measures of positive affect and negative correlations with measures of negative affect, emotion regulation difficulties, and coping-oriented substance use (Kremyar et al., 2020; Simons & Gaher, 2005).
Resilience
The BRS (Smith et al., 2008) is a six-item, self-report measure of the perceived ability to recover from stress. Items are rated on a 5-point Likert-type scale ranging from 1 (strongly disagree) to 5 (strongly agree). To score the measure, ratings on the six items are summed and divided by the total number of items rated. Higher mean scores indicate greater levels of resilience. Smith et al. (2008) reported that scores on the BRS show internal consistency (α range = .80–.91) and test–retest reliability (r = .62–.69 after 1–3 months) as well as positive correlations with measures of resilience, optimism, and positive affect and negative correlations with measures of pessimism, alexithymia, and mental health symptoms.
Data Analytic Plan
The present study sought to examine the psychometric properties of the MBHS in a university sample. First, descriptive statistics including mean, standard deviation, range, skewness, and kurtosis were generated for each MBHS item and proposed subscale. Skewness and excess kurtosis values between −2 and 2 were considered acceptable indicators of approximate normality (West et al., 1995).
Model fit for the proposed factor structure of the MBHS (i.e., nine subscales) and several alternate models were then tested using confirmatory factor analysis (CFA). All CFA models were estimated using maximum likelihood approximation with robust standard errors and the Satorra–Bentler correction (Satorra & Bentler, 1994). This estimation method adjusts for multivariate non-normality and is generally appropriate for use with ordinal indicators (Robitzsch, 2020; StataCorp, 2021). Goodness of model fit was evaluated using common fit indices. Model fit was considered good if the following criteria were met: a root mean square error of approximation (RMSEA) of .08 or less, a standardized root mean squared residual (SRMR) of .08 or less, and a comparative fit index (CFI) of .9 or above (Kline, 2016). The chi-square test was not used as it is known to be overly sensitive to sample size and is therefore not recommended as an index of model fit in large samples such as the one used in the present study (Kline, 2016). The standardized solution was used to estimate individual item loadings onto each subscale of the MBHS. Factor loadings of .4 or greater were considered good.
Internal consistency reliability was then estimated for each proposed subscale with both McDonald’s omega and Cronbach’s alpha (Hayes & Coutts, 2020). McDonald’s omega provides an estimate of internal consistency that takes into account the item loadings derived from the CFAs, whereas Cronbach’s alpha is computed using raw scores and the assumption that each item loads equally to the underlying construct being assessed by the subscale (i.e., essential tau-equivalence). Both estimates were generated to provide a more comprehensive assessment of internal consistency reliability. Subscale intercorrelations were then examined for signs of excessive overlap between subscales. Correlations among both observed sum scores and estimated latent factors were examined separately. As correlations between observed scores on symptom measures of psychopathology tend to be high and latent intercorrelations are expected to be even higher as a result of disattenuation due to unreliability, observed subscale intercorrelations above .8 and/or latent subscale intercorrelations above .9 were considered signs of potentially problematic collinearity (Rasmussen et al., 2019).
After evaluating the factor structure of the MBHS and internal consistency reliability of its subscales, evidence for convergent and discriminant validity was assessed by examining the relations between the MBHS subscales and theoretically related markers of psychopathology as assessed by self-report questionnaires. Specifically, following McCord (2020) and effect size conventions for interpreting correlation coefficients (Cohen, 1992), correlations between MBHS subscales and theoretically related criterion variables of .5 or greater were considered evidence of convergent validity in this sample. After identifying the MBHS subscale most strongly associated with each criterion variable, Steiger’s (1980) test was used to determine whether each other subscale association was significantly different from the strongest one at the p < .05 level. MBHS subscales were considered to show discriminant relations with criterion variables if they showed correlation coefficients that were both (a) of at least a value of |.5| with hypothesized external variables (i.e., evidenced convergent validity) and (b) were significantly higher than correlations with nonhypothesized criterion variables according to Steiger’s test.
MBHS sum scores were used in these analyses rather than factor scores for several reasons. Correlations with sum scores provide more generalizable evidence of convergent and discriminant validity than correlations with factor scores, which are estimated “using sample-specific weights applied to scores subjected to sample-specific standardization” (Widaman & Revelle, 2023, p. 794) and so are more susceptible to sampling error than sum scores. Interpretability of correlations between factor scores and criterion variables is further hampered by factor score indeterminacy (Grice, 2001; Waller, 2023; Widaman & Revelle, 2023). Finally, while sum scores and factor scores tend to be nearly identically correlated with criterion variables (Widaman & Revelle, 2023), extracted factor scores from a CFA model featuring multiple, correlated latent factors (as is the case with nine MBHS subscales) tend to show biased correlations with other variables (Logan et al., 2022).
Model fit and psychometric properties were assessed for a simplified model loading all MBHS items onto a single-factor as well as for exploratory, alternate models modified based on results of the original, nine-factor CFA and item characteristics. Additional models were also estimated to test for evidence of three higher order factors among the MBHS subscales. Models showing equivalent or superior fit to the proposed nine-factor model were further probed for evidence of convergent and discriminant validity.
Results
Preliminary Analyses
Descriptive statistics of the MBHS items are presented in Table 1. Most MBHS items showed acceptable skewness and excess kurtosis, except Items 14 (“I have tried to kill myself”), 23 (“I want to die”), and 25 (“I do dangerous things for thrills).
MBHS Item-Level Descriptive Statistics.
Note. MBHS = Multidimensional Behavioral Health Screen.
Proposed MBHS Subscales
Factor Structure
The first estimated CFA model was based on the proposed factor structure of the MBHS featuring nine latent subscales allowed to covary freely (see Figure 1). The model was a good fit to the data (RMSEA = .064, 90% confidence interval [CI] = [.059, .068]; SRMR = .052; CFI = .914; Akaike information criterion [AIC] = 37,935.3).

Proposed CFA Model With Nine MBHS Subscales as Latent Factors.
The factor loading of each item on its corresponding factor and subscale internal consistency coefficients are presented in Table 2. Factor loadings were substantial, with all being greater than .52 except for one item (Item 25: “I do dangerous things for thrills,” with a factor loading of .369). Most subscales displayed internal consistency (ωs > .73 and αs > .72), except the Somatization (ω = .685, α = .684) and Activation subscales (ω = .637, α = .570) demonstrated relatively lower internal consistency. Furthermore, no evidence was found to suggest that subscale internal consistency would substantially increase with the removal of any items except for Item 25, for which removal was estimated to improve the internal consistency of the Activation subscale by .048.
Internal Consistency Coefficients and Factor Loadings for the Proposed MBHS Subscales.
Note. MBHS = Multidimensional Behavioral Health Screen.
Descriptive statistics for each MBHS subscale sum score are presented in Table 3. Subscale intercorrelations are shown in Table 4 and were all positive and generally moderate to large in magnitude. Moreover, high observed score (r = .810) and latent score (r = 1.00) correlations between Demoralization and Anhedonia suggested that the subscales were not sufficiently distinct. Activation and Cognitive Issues also demonstrated a concerningly high latent score correlation (r = .965). The Substance Misuse subscale demonstrated notably lower correlations with the other MBHS subscales.
Descriptive Statistics of the Proposed and Modified MBHS Subscales.
Note. In the proposed MBHS subscales, MBHS = Multidimensional Behavioral Health Screen; SOM = Somatization; DEM = Demoralization; ANH = Anhedonia; ANX = Anxiety; SUI = Suicidal Tendencies; COG = Cognitive Issues; ACT = Activation; DSC = Disconstraint; SUB = Substance Misuse. In the modified MBHS subscales, DEM = demoralization; COG = cognitive activation.
Intercorrelations Among the Proposed and Modified MBHS Subscales.
Note. N = 602. Latent score correlations are presented in parentheses. All correlations are statistically significant at the p < .01 level. In the proposed MBHS subscales, MBHS = Multidimensional Behavioral Health Screen; SOM = Somatization; DEM = Demoralization; ANH = Anhedonia; ANX = Anxiety; SUI = Suicidal Tendencies; COG = Cognitive Issues; ACT= Activation; DSC = Disconstraint; SUB = Substance Misuse. In the modified MBHS subscales, SOM = Somatization; DEM = Demoralization; ANX = Anxiety; SUI = Suicidal Tendencies; COG = Cognitive Activation; DSC = Disconstraint; SUB = Substance Misuse. In the modified MBHS subscales, SOM = Somatization; DEM = Demoralization; ANX = Anxiety; SUI = Suicidal Tendencies; COG = Cognitive Activation; DSC = Disconstraint; SUB = Substance Misuse.
Validity
Correlations between MBHS subscales and criterion variables are shown in Table 5. Evidence for convergent relations was provided by large correlations (i.e., coefficients of >.5) among several hypothesized pairings between MBHS subscale scores with scores on self-report measures of psychopathology and strengths/resilience. Specifically, large correlations were found between the Somatization subscale and the K6, PHQ-9, PROMIS Anxiety SF, and the ISI; the Demoralization subscale and the K6, PHQ-9, PROMIS Anxiety SF, UCLA Loneliness Scale, and the BRS; the Anhedonia subscale and the K6, PHQ-9, PROMIS Anxiety SF, and the UCLA Loneliness Scale; the Anxiety subscale and the K6, PHQ-9, PROMIS Anxiety SF, DTS, and the BRS; the Suicidal Tendencies subscale and the PHQ-9; the Cognitive Issues subscale and the K6, PHQ-9, and the PROMIS Anxiety SF; the Activation subscale and the K6, PHQ-9, and the PROMIS Anxiety SF; and between the Substance Misuse subscale and the AUDIT.
Descriptive Statistics of Criterion Variables and Their Correlations With the Proposed and Modified MBHS Subscales.
Note. N = 602. The highest correlation in each row is bolded, as well as correlations in each row not significantly different from the greatest correlation at the p < .05 level. Correlations of |ρ| >.08 are significant at the p < .05 level. Correlations of |ρ| > .11 are significant at the p < .01 level. K6 = Kessler Psychological Distress Scale; PHQ-9 = Patient Health Questionnaire-9; PROMIS Anxiety SF = PROMIS Anxiety Short Form; ISI = Insomnia Severity Index; AUDIT = Alcohol Use Disorders Identification Test; DTS = Distress Tolerance Scale; BRS = Brief Resilience Scale. In the proposed Multidimensional Behavioral Health Screen (MBHS) subscales, SOM = Somatization; DEM = Demoralization; ANH = Anhedonia; ANX = Anxiety; SUI = Suicidal Tendencies; COG = Cognitive Issues; ACT= Activation; DSC = Disconstraint; SUB = Substance Misuse. In the modified MBHS subscales, SOM = Somatization; DEM = Demoralization; ANX = Anxiety; SUI = Suicidal Tendencies; COG = Cognitive Activation; DSC = Disconstraint; SUB = Substance Misuse.
Evidence for discriminant relations was also found, as the correlations between the K6 and Demoralization, the PHQ-9 and Demoralization, the PROMIS Anxiety SF and Anxiety, the AUDIT and Substance Misuse, and between both the DTS and BRS and Anxiety were also significantly higher than the correlations between these criterion variables and any other MBHS subscale. However, of the eight criterion variables, two showed at least one association not significantly different from the highest association, and correlations were generally high with other MBHS subscales even in those cases where a single correlation significantly greater than the rest can be identified. Taken together, the MBHS subscales showed more evidence for convergent than discriminant relations.
Modified MBHS Subscales
Due to the high observed and latent correlations among the nine MBHS subscales, an alternate model was specified in which all MBHS items loaded onto a single, general psychopathology factor (presented in Figure S1 in the online supplemental materials). However, this single-factor model demonstrated poorer fit than the original nine-factor model (RMSEA = .113, 90% CI = [.109, .117]; SRMR = .089; CFI = .694; AIC = 39,677.5). As previously noted, extremely high latent correlations between Demoralization and Anhedonia (r = 1.00) and between Activation and Cognitive Issues (r = .965) suggested that the subscales within each of these pairs are not meaningfully distinct. These subscale pairs were combined to test whether a simpler, seven-factor model may provide a better fit to the data. An examination of the item content in the Demoralization and Anhedonia subscales suggested that the three items in each subscale could be combined into a single, six-item Demoralization subscale. The decision was made to retain the “Demoralization” label for this modified subscale as demoralization is posited to reflect nonspecific, general distress and negative affect (McCord, 2020; Tellegen et al., 2003). It therefore seems reasonable to consider demoralization to be a simpler or more basic construct, with the burden of proof being required to show that a scale measures anhedonia or low-positive emotionality distinct from demoralization.
A preliminary CFA was run using the modified, six-item Demoralization subscale and a new subscale combining the items from Activation and Cognitive Issues in a single, six-item subscale to examine internal consistency via McDonald’s omega. Cronbach’s alpha was also computed for these combined subscales. The modified Demoralization subscale showed excellent internal consistency (ω = .890, α = .887). However, while evidence for internal consistency was found for the combined Activation and Cognitive Issues subscale (ω = .839, α = .822), Item 25 (“I do dangerous things for thrills”) again did not seem to belong, with its removal estimated to improve internal consistency by .025. Item 25 was removed from the combined Activation and Cognitive Issues subscale and consideration was initially given to adding it to the Disconstraint subscale, which is designed to index impulsivity (McCord, 2020). However, impulsivity is made up of four distinct personality facets: urgency, (lack of) premeditation, (lack of) perseverance, and sensation-seeking (Whiteside & Lynam, 2001). Item 25 appears to index sensation-seeking, whereas the three items of the Disconstraint subscale, Items 8 (“I often make impulsive decisions”), 17 (“I often break rules, regardless of the consequences”), and 26 (“I don’t think before I act”) appear to each index (lack of) premeditation. Furthermore, low premeditation has been shown to be related to disconstraint, whereas sensation-seeking is more closely related to extraversion (Sharma et al., 2014). Thus, rather than adding Item 25 to another MBHS subscale, the decision was made to drop it from the scale entirely.
The subscale made up of the remaining five items from the proposed Activation and Cognitive Issues subscales was named “Cognitive Activation” to reflect more accurately the common nature of its constituent items. The Somatization, Anxiety, Suicidal Tendencies, Disconstraint, and Substance Misuse subscales were retained unaltered. Descriptive statistics for the modified Demoralization and Cognitive Activation subscales are presented in Table 3.
Factor Structure
Model fit for this modified, seven-factor model was good (RMSEA = .058, 90% CI = [.054, .063]; SRMR = .046; CFI = .928; AIC = 36,742.0; see Figure 2). Furthermore, every model fit index was better for this model than for the original model based on the proposed nine-subscale structure of the MBHS. Estimates of internal consistency reliability and factor loadings for the modified MBHS subscales are presented in Table 6. Cognitive Activation (ω = .849, α = .847) showed evidence of internal consistency that would not benefit from the removal of any items. Factor loadings were also substantial, with all being greater than .52. Intercorrelations among the modified and unaltered subscales are presented in Table 4 and were again generally moderate to large in magnitude. Importantly, unlike the nine-factor model, no intercorrelation was so high as to raise concerns about possible collinearity, with the highest observed score (r = .672) and latent score (r = .820) correlations falling below the thresholds suggested by Rasmussen and colleagues (2019). The Substance Misuse subscale was also correlated with the modified subscales at relatively low levels.

Modified CFA Model With Seven MBHS Subscales as Latent Factors.
Reliability Coefficients and Factor Loadings for the Modified MBHS Subscales.
Validity
Correlations between the modified MBHS subscale scores and criterion variables are presented in Table 5. The pattern of correlations was similar to that of the proposed subscales, providing similar evidence for convergent and discriminant relations. There were two notable differences. First, the correlation between the K6 and the modified Demoralization subscale and the correlation between the K6 and Anxiety are not significantly different. Second, the UCLA Loneliness Scale is now clearly more strongly correlated with the modified Demoralization subscale than with any other subscale. Taken together, the modified, seven-subscale model of the MBHS appears to be a better fit to the data in this sample.
Higher Order Factors
The next set of models was estimated to determine whether evidence exists to support higher order latent factors in the MBHS. First, CFA models were estimated including a single, general psychopathology higher order factor indicated by all MBHS subscales. The inclusion of this single latent factor did not improve model fit for either the original, nine-subscale model (RMSEA = .075, 90% CI = [.071, .079]; SRMR = .067; CFI = .869, AIC = 38,267.3; see Figure S2) or the modified, seven-subscale model (RMSEA = .062, 90% CI = [.058, .067]; SRMR = .057; CFI = .913; AIC = 36,841.9; see Figure S3). Thus, while model fit indices were generally good for both of these single-factor, higher order models, the simpler, correlated subscale models are preferred.
CFA models in which the MBHS subscales were modeled as indicators of the hypothesized latent variables of Somatic/Cognitive Dysfunction, Emotional/Internalizing Dysfunction, and Behavioral/Externalizing Dysfunction were also run. As before, the inclusion of these three latent factors did not improve model fit for either the original, nine-subscale model (RMSEA = .075, 90% CI = [.071, .079]; SRMR = .084; CFI = .869; AIC = 38,269.2; see Figure S4) or the modified, seven-subscale model (RMSEA = .062, 90% CI = [.057, .066]; SRMR = .059; CFI = .916; AIC = 36,823.2; see Figure S5). The results of the present study, therefore, do not support the existence of higher order factors in the MBHS. The modified, seven-subscale model in which the subscales freely covary is the best supported model.
Discussion
The present study sought to examine the psychometric properties of the MBHS in a large, university sample of primarily young adults. The nine-subscale model of the MBHS proposed by McCord (2020) showed good model fit but excessively high correlations between the Demoralization and Anhedonia subscales and between the Activation and Cognitive Issues subscales. Therefore, the items in the Demoralization and Anhedonia subscales were combined into a single, modified Demoralization subscale and the items in the Activation and Cognitive Issues subscales were combined to form a new subscale labeled Cognitive Activation (with the exception of Item 25, “I do dangerous things for thrills,” which was dropped entirely from the MBHS). This new, seven-subscale model of the MBHS demonstrated better model fit than the nine-subscale model as well as high factor loadings, internal consistency reliability of the measure’s subscales, and no excessively high subscale intercorrelations. The overall pattern of intercorrelations among the seven subscales and their correlations with criterion variables support the use of the subscales as measures of distinct (though substantially related) psychological constructs. Evidence for convergent and some evidence for discriminant validity was found based on relations between MBHS subscale scores and these criterion variables. However, exploratory analyses did not support the existence of any higher order factors among the MBHS subscales. Taken together, our results support the use of MBHS scores as a promising and efficient measure of common, psychopathology-related aspects of behavioral health in a primarily young adult sample and its continued use in primary care settings.
A few findings are particularly notable. First, the preferred, seven-subscale model did not support separate Anhedonia and Activation subscales, suggesting that the MBHS may be a less specific screener than intended for anhedonia/low-positive emotionality and hyperactivity/hypomanic activation, respectively. New or substantially altered items may be needed that more specifically capture these constructs. However, these findings may be due to the nature of the primarily college student sample. College students tend to exhibit low base rates of mania/hypomania (Auerbach et al., 2018), and it could be that demoralization and anhedonia are less distinct in this population.
Second, the Substance Misuse subscale exhibited low associations with the other MBHS subscales. This is in line with past findings that alcohol use is less associated with psychopathology in young adult university and college students than in the general adult population (Dawson et al., 2005). This may be because relatively frequent alcohol use and binge drinking are often normative and viewed as part of the “college experience” (Colby et al., 2012; Winograd et al., 2012; Wrye & Pruitt, 2017). Specifically, alcohol use increases in the early college years but has been found to decrease steadily throughout and after graduation (Carter et al., 2010). Relatively reduced associations between alcohol use and mental health concerns in undergraduates and graduate students may also be potentially due to greater access to social support and psychological services on university and college campuses, which decreases the likelihood of turning to alcohol to self-medicate symptoms of psychopathology (Dawson et al., 2005). These findings suggest that a high score on the Substance Misuse subscale of the MBHS may not reliably distinguish between healthy and problematic alcohol use in a university student sample, whereas a low score implies that substance use is infrequent and likely not problematic. The Substance Misuse subscale may function best as a screener, with a high score suggesting the need to conduct a more thorough assessment of substance use.
Additional limitations of the present study should be noted. This study did not assess the convergent and discriminant relations of the MBHS subscale scores with the MMPI-2-RF or MMPI-3 scales or with any external variables specifically designed to assess disconstraint or hypomanic activation, limiting our ability to assess convergent relations for the proposed Activation and Disconstraint subscales. Future research may seek to replicate McCord’s (2020) findings with the MMPI-2-RF or MMPI-3 in other samples and settings. Another limitation is the relatively limited generalizability of our sample composed of undergraduate and graduate students, although recruiting our sample from the entire university community suggests it is likely to be representative of the population of individuals most likely to seek treatment or services at a university or college health clinic. In addition, the data for this study were collected relatively early during the COVID-19 pandemic. Results should be interpreted while keeping in mind that at least some participants were likely to be experiencing increased stress as a result, which may have affected reported estimates. Finally, given that the onset of psychotic symptoms often occurs in young adulthood with steep increases until 25 years old (Häfner et al., 1993; Kessler et al., 2007), one important limitation of using the MBHS in a university or college setting is the absence of assessment of psychotic symptoms. Many individuals who are at risk for developing psychosis often seek help for other mental health problems, primarily mood-related disturbances (Falkenberg et al., 2015; Fridgen et al., 2013), so mental health providers in university and college health clinics may be uniquely positioned to identify early signs of psychosis and activate appropriate services. Therefore, future versions of the MBHS would likely benefit from the addition of one or two psychosis screening questions, at least when used in settings with a large proportion of young adults.
Replications of this study’s findings and examinations of the psychometric properties of the MBHS subscale scores in other samples are critically needed to determine whether the differences in our assessment of the MBHS factor structure from that of the proposed structure reflect the true nature of the measure or stem from measurement noninvariance, perhaps based on the degree of self-perceived impairment or treatment-seeking. More broadly, further research using the MBHS is needed to establish the psychometric properties of its subscales. In particular, due to the cross-sectional nature of the current study, we were unable to examine test–retest reliability or the ability of MBHS subscales to predict clinical variables prospectively. Future studies would do well to examine the psychometric properties of the MBHS subscales and test for measurement and structural invariance in a sample seeking treatment at a university or college health clinic and to examine whether the modified, seven-factor structure identified in this study improves model fit in primary care, treatment-seeking samples as well. The MBHS holds great potential to be useful as an efficient screener for common mental health problems in military populations, community mental health centers, emergency departments, and other health care settings beyond primary care once reliability and validity of the subscale scores are established in these populations. As the factor structure found in the current student sample differed from the proposed factor structure developed for a primary care setting, it could be that unique versions of the MBHS tailored to the needs and characteristic ways of interpreting items of different populations may need to be developed. In addition, future researchers may wish to explore the incremental validity of the MBHS to determine whether its predictive ability is stronger than that of combining domain-specific screening measures (e.g., the PHQ-9 and K6).
The MBHS has high clinical utility as a screening tool for behavioral health symptoms and our findings suggest that the measure is psychometrically sound for use in a new setting, on college and university campuses. Similar to primary care settings, there is a need for brief, effective mental health screening in college health clinics given that rates of mental health problems among young adults are steadily increasing and many experience onset of these problems while in college. One fourth of college students presenting for routine health care at college health clinics screen positive for depressive symptoms and one in 10 report suicidal thoughts (Auerbach et al., 2018; Eisenberg et al., 2013; Mackenzie et al., 2011). Furthermore, age of onset of mental health disorders has been found to predict trajectory and prognosis, making early identification and referral to treatment critical for preventing outcomes such as academic impairment, college dropout, and even suicide (Pedrelli et al., 2015; Wang et al., 2007). Because the majority of students utilize university and college health clinic services for routine care, medical providers in these settings are crucial to screening and identifying mental health problems and facilitating appropriate referrals to counseling services, especially during current times when mental health resources on college campuses are strained. The use of the MBHS provides an excellent opportunity to bridge this gap, as it is efficient, easy to score, and interpret (including cut-points for referral) and can be utilized by healthcare providers of all specialties. In addition to these clinical implications, the MBHS holds the potential to advance research and clinical practice toward a more dimensional model of psychopathology, which is especially important for the goal of achieving integrative health care.
Conclusion
The present study found evidence for internal consistency reliability and convergent and discriminant validity based on relations with external variables in a nonprimary care, university setting. The MBHS, which is based on a dimensional paradigm, is therefore well-suited to bridge the gap between physical and behavioral health care. This measure may assist health care providers with determining referrals to behavioral health care providers and advance research in dimensional models of psychopathology.
Supplemental Material
sj-docx-1-asm-10.1177_10731911231205547 – Supplemental material for Psychometric Properties and Factor Structure of the Multidimensional Behavioral Health Screen (MBHS) With a University Sample
Supplemental material, sj-docx-1-asm-10.1177_10731911231205547 for Psychometric Properties and Factor Structure of the Multidimensional Behavioral Health Screen (MBHS) With a University Sample by Brooke R. Leonelli, Christian A. L. Bean and Joel W. Hughes in Assessment
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: The parent project from which the data were obtained was supported by a Rapid Response COVID-19-Related Pilot Funding Research Opportunity titled, Risk and protective factors for lasting mental health impact of the COVID-19 pandemic, and was funded by the Cleveland Brain Health Institute and the Brain Health Research Institute of Kent State University.
Supplemental Material
Supplemental material for this article is available online.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
