Abstract
The purpose of the current study was to identify Minnesota Multiphasic Personality Inventory-2–Restructured Form (MMPI-2-RF) correlates of police officer integrity violations and other problem behaviors in an archival database with original MMPI item responses and collateral information regarding integrity violations obtained for 417 male officers. In Study 1, we estimated MMPI-2-RF scores from the MMPI item pool (which includes approximately 80% of the MMPI-2-RF items) in a normative sample, a psychiatric inpatient sample, and a police officer sample, and conducted analyses that demonstrated the comparability of estimated and full scale scores for 41 of the 51 MMPI-2-RF scales. In Study 2, we correlated estimated MMPI-2-RF scores with information about subsequent integrity violations and problem behaviors from the integrity violation data set. Several meaningful associations were obtained, predominately with scales from the emotional, thought, and behavioral dysfunction domains of the MMPI-2-RF. Application of a correction for range restriction yielded substantially improved validity estimates. Finally, we calculated relative risk ratios for the statistically significant findings using cutoffs lower than 65T, which is traditionally used to identify clinically significant elevations, and found several meaningful relative risk ratios.
Keywords
The most comprehensive investigation of police officer integrity violations and the Minnesota Multiphasic Personality Inventory (MMPI; Hathaway & McKinley, 1943) was conducted by Boes, Chandler, and Timm (1997), who attempted to survey over 4,000 police departments nationwide in order to identify officers with a substantiated history of 1 of 10 predefined integrity violations that resulted in disciplinary action. This study is commonly known as the Defense Personnel and Security Research Center (PERSEREC) police integrity study. The authors reported that MMPI scales were generally not associated with integrity violations, with the exception of the Psychopathic Deviate Clinical Scale and the Lie Validity Scale, with both demonstrating a small effect size in the prediction of violator status.
Despite these and other studies contributing to the substantial extant literature on the MMPI, the test would come to be dated by the 1980s, necessitating the development and subsequent release of the MMPI-2 (Butcher, Dahlstrom, Graham, Tellegen, & Kaemmer, 1989). The MMPI-2 authors sought to improve the test and also placed a heavy emphasis on continuity with the original version so that findings from the MMPI literature, such as those just described, could be applied to interpretation of the MMPI-2. New norms were collected and items added, used primarily to score a new set of content scales. In spite of some significant, long-documented psychometric shortcomings of the original Clinical Scales of the MMPI (cf. Norman, 1972), only minimal changes were made to them to preserve continuity.
Responding to concerns regarding the psychometrics of the Clinical Scales, Tellegen et al. (2003) restructured them using modern test development techniques, yielding the Restructured Clinical (RC) scales. Numerous studies have documented comparable to improved convergent validity for the RC scales, as well as markedly improved discriminant validity (e.g., Arbisi, Sellbom, & Ben-Porath, 2008; Ben-Porath & Tellegen, 2008; Forbey & Ben-Porath, 2007; Handel & Archer, 2008; Sellbom, Ben-Porath, & Graham, 2006). Ben-Porath and Tellegen (2011) subsequently created a restructured form of the entire MMPI-2 using similar test development strategies. The MMPI-2–Restructured Form (MMPI-2-RF) has two new sets of scales to complement RC scale interpretation (i.e., the three Higher-Order scales and 23 Specific Problems scales), as well as revised and improved versions of the validity scales and Personality–Psychopathology-5 scales from the MMPI-2. The structure of the MMPI-2-RF is consistent with contemporary hierarchical models of personality and psychopathology (Krueger & Markon, 2006; Sellbom, Ben-Porath, & Bagby, 2008), and it has a growing, modern peer-reviewed research literature of over 240 publications at the time of this writing (University of Minnesota Press, 2014).
Several of the publications on the MMPI-2-RF utilized archival data sets of MMPI-2 responses, which was possible because all 338 MMPI-2-RF items are included in the 567-item MMPI-2 booklet and the two test booklets yielded comparable MMPI-2-RF scores (Tarescavage, Alosco, Ben-Porath, Wood, & Luna-Jones, 2014; Tellegen & Ben-Porath, 2011; Van der Heijden, Egger, & Derksen, 2010). Because of this feature of the test, research on the MMPI-2-RF has likely accumulated at a faster rate than if new data collection was required for all studies. However, only 81% of the MMPI-2-RF items are included in the original MMPI booklet, a feature that may have impeded similar use of thousands of original MMPI data sets in MMPI-2-RF research. Thus, researchers with access to MMPI databases of potential utility, such as the one collected by Boes et al. (1997), would seemingly be unable to use this information to investigate the psychometrically improved MMPI-2-RF.
Despite potential concerns about the impact of item loss, some researchers have attempted to investigate the reliability and validity of MMPI-2-RF scales scored from original MMPI responses. Using a data set collected by Hathaway and Monachesi (1963)—who administered the MMPI to over 15,000 ninth-grade students in Minnesota and subsequently assessed their adjustment and psychological functioning through early adulthood—Bolinskey, Trumbetta, Hanson, and Gottesman (2010) calculated truncated raw scores for RC4 (Antisocial Behavior). The resulting scale score, which included 73% of the items from the scale, was found to be associated with the development of psychopathy during early adulthood. Trumbetta, Bolinskey, and Gottesman (2013) later utilized the Hathaway and Monachesi (1963) data set to show that truncated scores from the MMPI-2-RF Higher-Order and RC scales generally demonstrated adequate internal consistency reliability and temporal stability. Overall, these studies provide preliminary support that MMPI-2-RF scores from MMPI responses are reliable and associated with criteria in expected ways, despite item loss concerns.
Current Study
As described earlier, Boes et al. (1997) obtained prehire MMPI data and found that original MMPI scores generally did not meaningfully correlate with criteria in their study. Recent research using the MMPI-2-RF has demonstrated that the test’s anchoring RC scales have improved predictive validity across settings, including police officer screenings (Sellbom, Fischler, & Ben-Porath, 2007), relative to the MMPI and MMPI-2 Clinical Scales. Therefore, a study investigating the predictive validity of the MMPI-2-RF scales in the Boes et al. (1997) integrity violation data set may be informative, but these authors collected original MMPI protocols.
Consequently, the purpose of the current study was twofold. First, we sought to examine the utility of a new method for estimating MMPI-2-RF scores from the subset of original MMPI items. We report in Study 1 the procedures and fidelity checks of this method. In Study 2, we investigated associations between estimated MMPI-2-RF scores (on scales that passed our Study 1 fidelity checks) and criteria from the seminal Boes et al. (1997) data set. We sought to examine in the context of discovery (Reichenbach, 1938) whether the score estimates of the psychometrically improved MMPI-2-RF scales were meaningfully correlated with integrity violations and other problem behaviors among police officers.
We also sought to disattenuate the correlations for range restriction using formulas derived by Hunter and Schmidt (2004). Correction for range restriction is needed when correlation coefficients are artificially diminished (insofar as serving as validity estimates is concerned) owing to decreased variance, which, in police officer candidate studies, results primarily from the rigorous screening process. Several components of the police officer application process are intended to screen out candidates at risk for integrity problems. These include background checks, interviews with previous supervisors, polygraph testing, drug screens, credit checks, and criminal record checks, as well as psychological evaluations (Scharf, 2006). Underreporting is also common in this setting, given the incentive to appear well-adjusted and mentally healthy to obtain employment (Carpenter & Raza, 1987; Hiatt & Hargrave, 1988). These factors combine to artifactually attenuate the magnitude of correlation coefficients by decreasing variance of psychological test scores, necessitating the use of disattenuation procedures to obtain more accurate validity estimates (see Anastasi & Urbina, 1997, for a review). Past research with the MMPI-2 (Sellbom et al., 2007) and Personality Assessment Inventory (Lowmaster & Morey, 2012) has used corrections for attenuation due to range restriction in prehire police officer samples as well.
Finally, we sought to investigate the practical utility of associations between estimated preemployment MMPI-2-RF scale scores and subsequent integrity violations by calculating relative risk ratios (RRRs), which quantify the increased likelihood of a negative outcome at preselected T-score cutoffs. In addition to the traditional clinically focused interpretive cutoff of 65T, we examined the utility of lower cutoffs because, for reasons just discussed, hired police officers tend to produce scores that are meaningfully lower than those found in the general population. Based on the prior findings of Sellbom et al. (2007) with the RC scales as well as Tarescavage, Corey, and Ben-Porath (2015) with the full MMPI-2-RF, we expected that lower cutoffs would produce selection ratios that are consistent with the base rates of integrity violations.
Study 1
The first study aimed to examine a new prorating method for estimating MMPI-2-RF scores from MMPI responses. To examine the fidelity of these prorated scores as estimates of full MMPI-2-RF raw scores, we first calculated correlations between similarly derived prorated scores and actual raw scores in three samples in which all 338 MMPI-2-RF item responses were available (the MMPI-2-RF normative sample, a psychiatric inpatient sample, and a large sample of law enforcement candidates). We next calculated MMPI-2-RF T-score means and standard deviations using the prorated scores in these three groups and compared the estimates with the descriptive statistics obtained for the full scale scores.
Method
Participants
We utilized three samples. The normative sample and psychiatric inpatients comparison group are presented in the MMPI-2-RF technical manual (Tellegen & Ben-Porath, 2011). The normative sample includes 1,138 males and 1,138 females (N = 2,276). The average age was 41.1 years (SD = 15.3). Most individuals were Caucasian (81.8%), with the remaining ethnic groups comprising African Americans (12.6%) or other ethnicity (6.7%). The psychiatric inpatients comparison group includes 659 males and 498 females (N = 1,157). The average age was 34.1 years (SD = 10.9). Most individuals were Caucasian (76.3%), with the remaining groups comprising African Americans (17.0%) or another ethnicity (6.7%). The male police officer candidate comparison group (Corey & Ben-Porath, 2014) comprises 1,037 police officer candidates. The average age is 27.5 years (SD = 7.4). Information on racial composition was not available.
Measures
MMPI-2-RF
The MMPI-2-RF (Ben-Porath & Tellegen, 2011) is a 338-item broadband measure of personality–psychopathology that has 51 scales and utilizes a true–false item response format. Nine scales assess threats to protocol validity, such as noncontent-based invalid responding (i.e., responding randomly, acquiescently, or counteracquiescently), overreporting, and underreporting. The remaining substantive scales are organized hierarchically. The three Higher-Order scales are consistent with contemporary conceptualizations of psychopathology (e.g., Krueger & Markon, 2006; Sellbom et al., 2008), as these scales measure internalizing psychopathology, externalizing psychopathology, and thought dysfunction. The nine RC scales measure major and distinctive psychopathological constructs embedded within the original MMPI Clinical Scales. Twenty-three Specific Problems scales measure narrower constructs focusing on somatic and cognitive complaints, internalizing problems, externalizing behaviors, and interpersonal functioning. Finally, the Personality–Psychopathology-5 scales measure features of personality-related psychopathology delineated by Harkness and McNulty (1994).
Procedures
As mentioned earlier, the MMPI booklet includes 274 of the 338 MMPI-2-RF items. Of the 274 MMPI items scored on the MMPI-2-RF, 35 (12.8%) were reworded on the MMPI-2, owing to outdated content (14 items), sexist language (5 items), sentence flow problems (5 items), grammatical problems (4 items), complex wording (3 items), awkward wording (2 items), or ambiguous wording (2 items). Ben-Porath and Butcher (1989) demonstrated psychometric comparability for the vast majority of the rewritten items with few exceptions, only five (1.8%) of which are scored on the MMPI-2-RF. Of these five items, only two (0.7%) showed diminished psychometric properties among men, who entirely comprise the Study 2 sample.
Thirteen MMPI-2-RF scales have all items included in the MMPI booklet (see Table 1). For the remaining scales, we estimated prorated MMPI-2-RF raw scores from MMPI responses by determining the proportion of available MMPI-2-RF items answered in the keyed direction and multiplying this value by the number of items for the full scale. For example, RC3 has 15 items, but only 12 of these are included in the MMPI item pool. If an individual in the current sample responded to all 12 of these items (i.e., 100%) in the keyed direction, the prorated RC3 score would be 15 (15 × 1.00). For a police candidate who responded to 4 of the 12 items (i.e., 33.3%) in the keyed direction, the prorated RC3 score would be 5 (15 × .333). If an individual responded to none of the 12 available items in the keyed direction, the prorated RC3 score would be 0 (15 × 0). For items that have duplicates in the MMPI booklet, we used only the first presented item.
Minnesota Multiphasic Personality Inventory-2–Restructured Form (MMPI-2-RF) Scales Descriptives and Correlations Between Full Scale Scores and Estimates From MMPI Item Pool.
Note. Full normative sample means and standard deviations are 50 and 10, respectively, for all scales.
VRIN-r = Variable Response Inconsistency; TRIN-r = True Response Inconsistency; F-r = Infrequent Responses; Fp-r = Infrequent Psychopathological Responses; Fs = Infrequent Somatic Responses; FBS-r = Symptom Validity; RBS = Response Bias Scale; L-r = Uncommon Virtues; K-r = Adjustment Validity; EID = Emotional/Internalizing Dysfunction; THD = Thought Dysfunction; BXD = Behavioral/Externalizing Dysfunction; RCd = Demoralization; RC1 = Somatic Complaints; RC2 = Low Positive Emotions; RC3 = Cynicism; RC4 = Antisocial Behavior; RC6 = Ideas of Persecution; RC7 = Dysfunctional Negative Emotions; RC8 = Aberrant Experiences; RC9 = Hypomanic Activation; MLS = Malaise; HPC = Head Pain Complaints; NUC = Neurological Complaints; COG = Cognitive Complaints; HLP = Helplessness/Hopelessness; SFD = Self-Doubt; NFC = Inefficacy; STW = Stress/Worry; AXY = Anxiety; ANP = Anger Proneness; BRF = Behavior-Restricting Fears; MSF = Multiple Specific Fears; JCP = Juvenile Conduct Problems; SUB = Substance Abuse; AGG = Aggression; ACT = Activation; FML = Family Problems; IPP = Interpersonal Passivity; SAV = Social Avoidance; SHY = Shyness; DSF = Disaffiliativeness; AEC = Aesthetic–Literary Interests; MEC = Mechanical–Physical Interests; AGGR-r = Aggressiveness–Revised; PSYC-r = Psychoticism–Revised; DISC-r = Disconstraint–Revised; NEGE-r = Negative Emotionality/Neuroticism–Revised; INTR-r = Introversion/Low Positive Emotionality–Revised.
Data Analyses
To rely on prorated scores, the reduced subset of items must measure the same construct as the full set. To test this assumption, we conducted two fidelity checks. First, we calculated correlations between the reduced MMPI-2-RF scale subsets and the full scales in three MMPI-2-RF comparison groups with different levels of psychopathology (normative, psychiatric inpatients, and male law enforcement officers; Corey & Ben-Porath, 2014; Tellegen & Ben-Porath, 2011). This process is analogous to testing the similarity of short forms to full forms. For example, the 12-item subset of RC3 items available on the MMPI was correlated with the full 15-item scale. Subsets that shared >80% variance (r > .90) with the full scales in all three samples were deemed to substantially measure the same construct and passed this first fidelity check.
Next, in the same samples, we calculated uniform T-scores (Tellegen & Ben-Porath, 1992) using prorated raw scores derived from the reduced subsets and compared the resulting means and standard deviations with the actual descriptive statistics for these groups. For example, as noted in Table 1 and described in the next section, the prorated raw scores on RC3 yielded an estimated mean of 51T and a standard deviation of 11T in the normative (standardization) sample, which, by definition, has means of 50T and standard deviations of 10T for all scales. Estimated means and standard deviations that were within 2T of actual scores in all three samples passed this second fidelity check. This criterion was based on Cohen (1992), who defined a difference in distributions of 0.20 standard deviations as a small effect.
Results and Discussion
To examine the fidelity of prorated scores as estimates of full MMPI-2-RF raw scores, we calculated correlations between prorated scores and actual raw scores in three samples (the normative sample, a psychiatric inpatient sample, and a large sample of male law enforcement candidates; see Table 1). The procedure just described yielded correlations between full and prorated MMPI-2-RF scales ≥.90 for 44 of 51 scales (86.2%). Exceptions included Variable Response Inconsistency (VRIN-r), Cognitive Complaints (COG), Suicidal Ideation (SUI), Helplessness/Hopelessness (HLP), Stress/Worry (STW), Substance Abuse (SUB), and Disaffiliativeness (DSF). We then used prorated raw scores to calculate estimated uniform T-scores (see Ben-Porath & Tellegen, 2011, for a review of calculations). This process yielded estimated means and standard deviations that were within 2 T-score points of actual scores for the majority of scales in the three samples, except for True Response Inconsistency (TRIN-r), Behavioral/Externalizing Dysfunction (BXD), SUI, HLP, and SUB (see Table 1).
Because VRIN-r, TRIN-r, BXD, COG, SUI, HLP, STW, SUB, and DSF did not meet the criteria outlined for both fidelity checks, they were excluded from analyses in Study 2. However, based on these findings from three samples with varying levels of psychopathology, the prorating procedure yielded score estimates generally comparable to the full scales for the remaining 28 MMPI-2-RF scales. Thus, this study indicates that actual scores (13 scales) and reliable estimates (28 scales) can be obtained from MMPI responses for 41 of the 51 MMPI-2-RF scales.
Study 2
The purpose of the second study was to examine the ability of the MMPI-2-RF scales to predict integrity violations and problem behaviors in the Boes et al. (1997) data set. This data set only includes MMPI item information. We calculated prorated MMPI-2-RF scores from the MMPI data and examined the resulting MMPI-2-RF score estimates’ associations with criteria from the Boes et al. (1997) study. We disattenuated these correlations for range restriction, owing to concerns discussed earlier, and calculated RRRs to examine the practical utility of the results.
Method
Participants
Four-hundred and sixty of the 878 police officers in the Boes et al. (1997) study were administered the original MMPI as part of their preemployment evaluation. We obtained data for all but one of these individuals. We excluded two individuals who had 15 or more missing item responses (equivalent to the 18-item MMPI-2-RF cutoff based on 338 items) or scored greater than a raw score of 14 on the original MMPI F scale, which is a cutoff associated with increased possibility for symptom exaggeration or noncontent-based invalid responding (i.e., responding randomly, acquiescently, or counteracquiescently). We also excluded female officers (n = 40), because we did not have an adequate sample size to conduct analyses by gender. The final sample included 417 male police officers that ranged in age from 20 to 50 years (M = 31.8, SD = 11.8). The majority had some college education (51.5%), with 31.2% having a high school diploma, 14.8% having a bachelor’s degree, and 2.5% having a General Educational Development certification. About half were of Caucasian descent (51.8%), with the remaining ethnicities comprising African American (32.6%), Hispanic (14.7%), and Asian (1.0%). The average years of employment was 3.6 years (SD = 3.0).
Measures
MMPI-2-RF
The MMPI-2-RF was described in Study 1.
Procedure
The data for this study were drawn from the PERSEREC police integrity study, an investigation of personality inventory predictors of police officer integrity violations collected in the mid-1990s from 69 U.S. police departments (Boes et al., 1997). Officers who had 1 of 10 integrity violations that had been designated for study based on input from expert panels and a literature review were matched on several characteristics (gender, age, ethnicity, and length of employment) to officers without a history of these 10 integrity violations. However, Boes et al. (1997) reported that approximately one quarter of the nonviolator group had a history of “disciplinary actions” (p. 87), which involved primarily alcohol abuse, firearms misuse, use of excessive force, supervisory problems, substandard work, conduct unbecoming, procedural violations, loss/private use of equipment, falsification of time worked, and failure to attend court. Because the violations just listed were not among the 10 designated for the study by Boes et al. (1997), individuals with histories of disciplinary actions were assigned to the nonviolator group. This group was consequently not optimal as a nonviolating control group.
For the purposes of the current study, we removed these individuals from the nonviolator group and subsequently identified the group of 161 police officers in the Boes et al. (1997) sample with no history of the 10 integrity violations as well as no history of other serious problem behaviors. For each of the integrity violations designated by Boes et al. (1997) for their study, as well as other problem behaviors that led to disciplinary actions, we created a dichotomous dummy variable differentiating between those who committed the violation (assigned a value of 1.0) and the 161 officers without any violations (assigned a value of 0). For example, as presented in Tables 2 through 5, the 8 individuals with alcohol abuse problems were compared with the 161 nonviolators (n = 169). For the next variable, firearms misuse, 6 individuals with this issue were compared with the 161 nonviolators (n = 167), and so on for the remaining variables. We excluded variables that had a base rate lower than 2.0% to mitigate the influence of outliers, which resulted in the exclusion of one integrity violation defined by Boes et al. (1997; information breach endangering officers) and seven other counterproductive work behaviors (illegal drug use, discrimination, poor evidence control, release of unauthorized information, criminal violation, use of authority to obtain, and moonlighting).
Male Officer Estimated Minnesota Multiphasic Personality Inventory-2–Restructured Form Higher Order, Restructured Clinical, and Somatic–Cognitive Specific Problems Scales Correlations (Disattenuated) With Integrity Violations and Problem Behaviors.
Note. L-r = Uncommon Virtues; K-r = Adjustment Validity; EID = Emotional/Internalizing Dysfunction; THD = Thought Dysfunction; RCd = Demoralization; RC1 = Somatic Complaints; RC2 = Low Positive Emotions; RC3 = Cynicism; RC4 = Antisocial Behavior; RC6 = Ideas of Persecution; RC7 = Dysfunctional Negative Emotions; RC8 = Aberrant Experiences; RC9 = Hypomanic Activation. Behavioral/Externalizing Dysfunction is not included because it demonstrated inadequate score comparability. Bolded correlations are statistically significant and ≥ |.15|. No disattenuated correlations are provided for the underreporting scales because they are not restricted.
p < .05. **p < .01.
Data Analyses
As mentioned earlier, because VRIN-r, TRIN-r, BXD, COG, SUI, HLP, STW, SUB, and DSF did not meet the criteria outlined for both fidelity checks in Study 1, they were excluded from analyses in Study 2. For the remaining scales, we first compared the estimated means and standard deviations for this sample with those of the normative sample (who by definition produced mean T-scores of 50 and standard deviations of 10 on all scales) and to the MMPI-2-RF male police officer candidate comparison group (Corey & Ben-Porath, 2014). Consistent with traditional benchmarks for MMPI-2 and MMPI-2-RF interpretation, we used a difference of 5 T-score points to demarcate a meaningful difference across samples (Graham, 2012).
Next, we calculated zero-order correlations between the prorated MMPI-2-RF raw scores and violation criteria. We excluded most validity scales because they do not assess constructs that would be expected to have associations with the criteria; however, we did include Uncommon Virtues (L-r) and Adjustment Validity (K-r) because their MMPI counterparts were associated with criteria from the police integrity study from which the current data are derived (Boes et al., 1997).
In line with prior research, we considered a cutoff of r > |.15| as indicative of a practically meaningful effect (Graham, Ben-Porath, & McNulty, 1999; Sellbom et al., 2007). Only statistically significant correlations that met this criterion were interpreted. This cutoff was considered a conservative and appropriate reference for our study given that the sample had range restricted scores, leading to artifactually diminished correlation coefficients, and past research with police officers indicates that correlations of .15 and higher produce practically meaningful findings. For example, Tarescavage, Corey, and Ben-Porath (2015) reported that police officers in their sample who scored >50T on RC2 were at over three times greater risk of commitment problems (p < .05) than those who scored below the T-score cutoff. The correlation between these two variables was .18.
We did not implement alpha-adjustment procedures to correct for potential familywise error rates because this study was conducted in the context of discovery (Reichenbach, 1938). Ellis (2010) also presents an argument against such procedures, which reduce statistical power and complicate integration of the research literature. In the current study, which is the most comprehensive investigation of integrity violations among police officers, we utilized archival data that were limited to a median of 175 officers per variable, which yields a power value of .51 to detect correlational magnitudes of .15 using an alpha of .05 (Faul, Erdfelder, Lang, & Buchner, 2007). Decreasing the alpha to .01 reduces the power value to .28, and further adjusting the alpha to .001 would yield a power value of .09. As demonstrated by this example, such corrections would substantially limit the amount of practically meaningful information that could be gained from this study.
In addition to zero-order correlations between estimated MMPI-2-RF scores and criteria, we also calculated range restriction-corrected correlations using formulas derived from Hunter and Schmidt (2004), an approach that has been used in previous investigations (Lowmaster & Morey, 2012; Sellbom et al., 2007). Three pieces of information are required to apply the formula: (a) the zero-order correlation between the prorated MMPI-2-RF scale score and criterion, (b) the estimated standard deviation of the MMPI-2-RF scale score in this sample, and (c) the standard deviation of the MMPI-2-RF scale score in the general population (i.e., the unrestricted standard deviation). For example, as seen in Table 2 and discussed later, the zero-order correlation between prorated MMPI-2-RF Thought Dysfunction (THD) and an alcohol abuse violation is .20. The estimated standard deviation for THD was 7.0 in this sample. The general population standard deviation is 10T (by definition). Applying the range-restriction correction using these values yields a corrected correlation of .28, which is the unbiased (in terms of range restriction) validity estimate for prorated THD scores as predictors of alcohol abuse violations in this sample. Other studies have used normative information in this manner to estimate the unrestricted standard deviation (Hoffman, 1995; Ones & Viswesvaran, 2003; Sackett & Ostgaard, 1994).
Finally, to investigate the practical utility of estimated MMPI-2-RF scale scores, we calculated RRR with the criteria using cutoffs of ≥65T, 60T, 55T, 50T, and 45T for positive correlations (as well as <39T and 33T for negative associations). RRR values are calculated by dividing the risk of a violation or problem behavior for individuals who score at or above the cutoff by the risk of a violation or problem behavior for individuals who score below the cutoff. We calculated 95% confidence intervals for the RRRs, which indicate nonsignificant findings if the range overlaps with 1.0 (meaning we cannot reject the null hypothesis that there is an equal risk of negative outcome for both groups). We only calculated RRRs for prorated scales that were significantly correlated with the criteria. Finally, only RRRs that yielded selection ratios between 2.0% and 20% were calculated in order to reduce the risk of outliers affecting the results and to decrease the likelihood of false positives.
Results and Discussion
Descriptive Findings
Estimated MMPI-2-RF scale score means and standard deviations are presented in the right columns of Table 1 next to the actual scores of the male police officer candidate comparison group (Corey & Ben-Porath, 2014). Both groups showed evidence of range restriction, as the median standard deviations in the current sample and comparison group were 6.8 and 6.3, respectively; approximately two thirds of the general population standard deviation of 10.
The current sample scored meaningfully lower than the general population on the majority of interpreted scales, but produced higher estimated scores on the underreporting validity scales L-r and K-r, as well as Mechanical/Physical Interests. The sample’s remaining estimated scores were at least 5 T-score points lower than the general population on all but the following scales: Infrequent Psychopathological Responses (Fp-r), Symptom Validity (FBS-r), Cynicism (RC3), Persecutory Ideation (RC6), Gastrointestinal Complaints (GIC), Juvenile Conduct Problems (JCP), Activation (ACT), Social Avoidance (SAV), Aggressiveness (AGGR-r), Disconstraint (DISC-r), and Introversion (INTR-r).
Correlations
Zero-order and disattenuated correlations between prorated MMPI-2-RF underreporting and substantive scale scores with integrity violations are presented in Table 2 for the underreporting, Higher-Order, Restructured ClinicalRestructured Clinical, and somatic/cognitive Specific Problems scales, in Table 3 for the Internalizing and Externalizing Specific Problem Scales, and in Table 4 for the interpersonal Specific Problems scales and Personality–Psychopathology-5 scales. In order to facilitate interpretation, the findings are summarized in reference to MMPI-2-RF domains that include (a) Underreporting, (b) Emotional Dysfunction, (c) Thought Dysfunction, (d) Behavioral Dysfunction, (e) Somatic/Cognitive Complaints, and (f) Interpersonal Functioning. In keeping with our goal to interpret the findings of this study in the context of exploration, the correlate results and discussion from each domain are structured into three paragraphs: a straightforward description of the findings, a discussion of correlates that converge with the broader research literature (e.g., Tarescavage, Brewster, Corey, & Ben-Porath, 2014; Tarescavage, Corey, & Ben-Porath, 2015; Tarescavage, Corey, Gupton, & Ben-Porath, 2015; Tarescavage, Fischler, et al., 2014), and identification of correlates that may be an artifact of Type I error or would need to be replicated before integrating into the literature base.
Male Officer Minnesota Multiphasic Personality Inventory-2–Restructured Form Somatic/Cognitive, Internalizing, and Externalizing Specific Problem Scales Correlations (Disattenuated) With Integrity Violations and Problem Behaviors.
Note. MLS = Malaise; GIC = Gastrointestinal Complaints; HPC = Head Pain Complaints; NUC = Neurological Complaints; COG = Cognitive Complaints; SFD = Self-Doubt; NFC = Inefficacy; AXY = Anxiety; ANP = Anger Proneness; BRF = Behavior-Restricting Fears; MSF = Multiple Specific Fears; JCP = Juvenile Conduct Problems; AGG = Aggression; ACT = Activation. Cognitive Complaints, Suicidal Ideation, Helplessness/Hopelessness, Stress/Worry, and Substance Abuse are not included because they demonstrated inadequate score comparability. Bolded correlations are statistically significant and ≥ |.15|.
p < .05. **p < .01.
Male Officer Estimated Minnesota Multiphasic Personality Inventory-2–Restructured Form Interpersonal Functioning Specific Problem Scales, Interest Scales, and Personality–Psychopathology-5 Scales Correlations (Disattenuated) With Integrity Violations and Problem Behaviors.
Note. FML = Family Problems; IPP = Interpersonal Passivity; SAV = Social Avoidance; SHY = Shyness; AGGR-r = Aggressiveness–Revised; PSYC-r = Psychoticism–Revised; DISC-r = Disconstraint–Revised; NEGE-r = Negative Emotionality/Neuroticism–Revised; INTR-r = Introversion/Low Positive Emotionality–Revised. Disaffiliativeness was not included because it demonstrated inadequate score comparability. Bolded correlations are statistically significant and ≥ |.15|.
p < .05. **p < .01.
Underreporting
The underreporting scales include Uncommon Virtues (L-r) and Adjustment Validity (K-r). We did not present disattenuated correlations for these scales, because, as seen in Table 1, they are not range restricted in this setting, in which there is incentive to underreport psychological symptoms to obtain employment. L-r was positively associated with information breach aiding criminals and receiving protection money, whereas K-r was negatively associated with firearms misuse and use of excessive force. K-r was also positively associated with information breach aiding criminals.
Boes et al. (1997) concluded from the observed association between Lie scale scores and integrity violations that “individuals prone to trust betrayal will be particularly concerned with trying to appear honest” (p. 46). This observation invokes an essential paradox found among at least a subpopulation of individuals who claim high levels of moral virtuousness (L-r) and adaptation (K-r): namely, that those to whom the L-r/K-r items least apply are, at least in a personnel selection context, most likely to endorse them. Findings from the current study provide a more nuanced illumination. In contrast to the other behaviors in the Integrity Violation domain (e.g., bribes/shakedowns, theft on duty, embezzlement/fraud), which may involve impulsive, opportunistic motives, the integrity violations associated in this study with L-r—namely, information breach aiding criminals and receiving protection money—may involve more interpersonal motives. Notably, Sellbom and Bagby (2008) reported that L-r elevations tend to be associated with contexts in which interpersonal demands are emphasized. Similarly, Detrick and Chibnall (2014) concluded in a study of police candidates under high-demand conditions (personnel selection contexts in which test results influenced suitability determinations) and low-demand conditions (police academy contexts in which the results were used solely for research) that L-r elevations are associated primarily with an orientation toward “moralistic bias (social communion).” These collective findings support the hypothesis that some portion of L-r elevations occurs in circumstances in which the test taker is motivated to project a prosocial orientation in order to mask a socially proscribed one.
The correlate findings for K-r, particularly the negative associations, are more difficult to reconcile with the current literature. Further complicating the matter, most studies of the MMPI-2-RF in this population do not provide correlations with the validity scales of the test, because they are not intended to measure substantive constructs associated with police officer behavior. As it stands, though the negative associations with K-r may result from genuine positive adjustment, which confounds moderate elevations on this scale, it is not reasonable at this juncture to infer functional outcomes (i.e., less firearms misuse and decreased use of excessive force) from higher scores on an underreporting scale. Of note in this context, the more robust substantive scale findings relative to the underreporting scale findings support that the latter are best used to assess protocol validity rather than to infer future police officer behavior.
Emotional Dysfunction
The interpreted scales in the Emotional Dysfunction domain include Higher-Order scale Emotional/Internalizing Dysfunction (EID), Restructured Clinical Scales Demoralization (RCd), Low Positive Emotions (RC2), and Dysfunctional Negative Emotions (RC7), Specific Problems Scales Self-Doubt (SFD), Inefficacy (NFC), Anxiety (AXY), Anger Proneness (ANP), Behavior-Restricting Fears (BRF), and Multiple Specific Fears (MSF), and Personality–Psychopathology-5 scales Negative Emotionality/Neuroticism–Revised (NEGE-r) and Low Positive Emotionality/Introversion–Revised (INTR-r). Prorated scale scores from this domain demonstrated positive associations with use of excessive force (EID and RCd), embezzlement (RCd), and failure to attend court hearings (EID and INTR-r). Other positive associations included supervisory problems (MSF), substandard work (BRF), conduct unbecoming (RC2 and INTR-r), procedural violations (MSF), and loss/private use of equipment (MSF). Some scales from this domain had negative associations with problem behaviors, including NFC (theft on duty), NEGE-r (off-duty violations), and INTR-r (alcohol abuse).
A comprehensive job analysis undertaken for the California Commission on Peace Officer Standards and Training (Spilberg & Corey, 2014) demonstrates correspondence between deficits in emotional control and “excessive, unrestrained use of force,” “reactions to job stress, both near-term (anxiety, worry) and long-term (e.g., physical symptoms, burnout, substance abuse),” and “[allowing] personal problems and stressors to bleed into behavior on the job,” among others, reflecting concordance with an array of problem behaviors found in this study to correlate with MMPI-2-RF scores on the Emotional Dysfunction domain scales.
This consideration notwithstanding, some of the correlations from this domain does not converge with prior literature. Broadly speaking, findings from several studies on associations between the MMPI-2-RF scales from this domain and police officer behavior (Tarescavage, Brewster, et al., 2014; Tarescavage, Corey, Gupton, et al., 2015; Tarescavage, Fischler, et al., 2014) have neither supported that BRF and MSF are meaningfully correlated with future problems among police officers, nor have they indicated that scores from the emotional dysfunction domain are negatively associated with problem behaviors. However, there is one exception relevant to the current study. Tarescavage, Brewster, et al. (2014) also found that lower scores on INTR-r were associated with future alcohol use. As described by Ben-Porath and Tellegen (2011), low scores on INTR-r reflect high degrees of energy and positive emotions (i.e., components of extraversion). Ruiz, Pincus, and Schinka (2008) found that though extraversion (as delineated in the five-factor model of personality) was negatively associated with substance use disorders, the facet of excitement seeking demonstrated a positive correlation. A subset of the INTR-r items describe the propensity to seek out exciting social activities, such as parties and places where there is a crowd (the denial of which would indicate introversion). It may be that this subset of items is driving the negative association between INTR-r and alcohol use that was observed in this study and by Tarescavage, Brewster, et al. (2014).
Thought Dysfunction
The interpreted scales in the thought dysfunction domain include Higher-Order scale THD, Restructured Clinical scales Persecutory Ideation (RC6) and Aberrant Experiences (RC8), and Personality–Psychopathology-5 Scale Psychoticism–Revised (PSYC-r). Prorated scale scores in this domain were positively associated with failure to attend court hearings (THD, RC6, RC8, and PSYC-r), alcohol abuse (THD, RC8, and PSYC-r), supervisory problems (THD and RC8), and use of excessive force (RC6).
Although scale scores in the thought dysfunction domain may not be intuitively associated with the problem behaviors identified in this study, Spilberg and Corey (2014) reported that impaired decision making and judgment in police officers (which would be reflected by moderate and higher elevations on the thought dysfunction scales) are associated with engaging in “unnecessary physical force before verbal control methods [are] exhausted.” Sellbom et al. (2007), in their sample of Midwestern police candidates, reported that RC6 was the most robust predictor of problem behaviors, including excessive force, and that RC8 correlated with missing court appearances and other counterproductive behaviors. Relatedly, Tarescavage, Brewster, et al. (2014) found that higher scores on all thought dysfunction scales were related to supervisor ratings indicating a decreased ability to predict situational outcomes.
As of the time of this writing, only one study in this setting has investigated associations between any of the MMPI-2-RF scales and substance abuse (Tarescavage, Brewster, et al., 2014), and no studies have sought to identify problems between MMPI-2-RF scales and supervisory problems. Tarescavage, Brewster, et al. (2014) found no meaningful associations between the THD scales and supervisor ratings of police officers’ substance use problems. Forbey and Ben-Porath (2008) did find modest associations (.20 < r < .30) between RC6 and RC8 with substance abuse, as measured by the Drug Abuse Screening Test (Skinner, 1982) and Michigan Alcohol Screening Test (Selzer, 1971), in a nonclinical sample of college undergraduates. However, given the lack of research in this area among police officers, further investigations of the associations between MMPI-2-RF measured thought dysfunction with supervisory problems and alcohol abuse are warranted.
Behavioral Dysfunction
The interpreted scales in the Behavioral Dysfunction domain include Restructured Clinical scales Antisocial Behavior (RC4) and Hypomanic Activation (RC9), and the Externalizing Specific Problems scales Juvenile Conduct Problems (JCP), Aggression (AGG), and Activation (ACT). This domain also includes Personality–Psychopathology-5 scales Aggressiveness–Revised (AGGR-r) and Disconstraint–Revised (DISC-r). Scales in this domain were positively associated with supervisory problems (JCP and AGG), off-duty violations (RC4 and JCP), use of excessive force (RC4 and JCP), and conduct unbecoming (RC4). They were also positively associated with falsification of time worked (AGG) and unlawfully dropping a case (AGG). Some prorated scale scores also had negative associations with use of excessive force (AGGR-r), information breach aiding criminals (RC9 and DISC-r), receiving protection money (DISC-r), and fix of testimony (ACT).
The problem behaviors found in this study to be positively correlated with the BHD domain scales (namely, supervisory problems, off-duty violations, use of excessive force, conduct unbecoming, falsification of time worked, and unlawfully dropping a case) correspond well with the constructs measured by them. Forbey and Ben-Porath (2007) reported positive associations between scores on RC4 and low constraint, low behavioral control, low agreeableness, low consciousness, and increased impulsiveness and anger. The negative association between AGGR-r and use of excessive force found in the current study adds to a robust research literature in this setting demonstrating this seemingly paradoxical finding with the MMPI-2-RF externalizing scales—particularly those associated with RC9 (Tarescavage, Brewster, et al., 2014; Tarescavage, Corey, & Ben-Porath, 2015; Tarescavage, Corey, Gupton, et al., 2015). These findings suggest that counterproductive behaviors in police officers may be potentiated by deficits in psychological activation and instrumental orientation, which would be reflected by low scores on RC9, ACT, AGG, and AGGR-r. In the case of use of excessive force, police officers with these deficits may confront suspects too slowly and inadequately, thereby permitting the suspect’s behavior to escalate to a point that ultimately triggers an overreaction by the officer.
The negative associations between the MMPI-2-RF externalizing scales and information breach endangering officers, receiving protection money, and fixing testimony are more difficult to reconcile with the current literature. However, this does not necessarily indicate that low scores on these scales are not associated with serious integrity violations, as these criteria are difficult to study because of their very low base rates. For example, Tarescavage, Fischler, et al. (2014) note that several variables were excluded from their study of MMPI-2-RF predictors of integrity violations for this reason, including unlawful activity, using position for personal advantage, and accepting gratuities. Similarly, Tarescavage, Corey, Gupton, et al. (2015) excluded the following variables from their analyses of MMPI-2-RF predictors of police officer problem behaviors: abuses authority, uses position for personal advantage, and unlawful activity. Of note in this context, Tarescavage, Corey, Gupton, et al. (2015) did find that low scores on AGG and ACT were associated with problems related to police integrity, defined as the maintenance of high standards of personal and professional conduct, including honesty, impartiality, trustworthiness, and compliance with laws, regulations, and policies. However, further replication is necessary before integrating these findings into the broader research literature.
Somatic/Cognitive Complaints
The interpreted scales in the Somatic/Cognitive Complaints domain include Restructured Clinical scale Somatic Complaints (RC1) and the Somatic/Cognitive Specific Problems Scales, which are composed of Malaise (MLS), Gastrointestinal Complaints (GIC), Head Pain Complaints (HPC), and Neurological Complaints (NUC). Associations from this domain included conduct unbecoming (MLS), procedural violations (GIC), use of excessive force (HPC), and embezzlement (HPC).
The problem behaviors correlated with the Somatic/Cognitive Complaints domain may help broaden understanding of the constructs measured by its constituent scales. A number of researchers have reported that elevations on scales in this domain often occur without a clear somatic or other origin and without concomitant evidence of impaired cognitive functioning (Ben-Porath, 2012). Tarescavage, Corey, and Ben-Porath (2015) also found positive associations between MMPI-2-RF scores in the Somatic/Cognitive Complaints domain, particularly MLS, and problems in the following police officer performance domains: Emotional Control and Stress, Routine Task Performance, Decision Making and Judgment, and Social Competence and Teamwork. These problems may result from other difficulties correlated with the Somatic/Cognitive scale scores, including preoccupation with poor health, complaints of sleep disturbance and of fatigue and low energy, and problems concentrating (Tellegen & Ben-Porath, 2011).
The correlations between GIC and HPC, in particular, with problem behaviors in this study do not converge with the broader literature. In the case of GIC, this scale is markedly range restricted in police officer samples, making it difficult to examine any associations with the scale (Tarescavage, Brewster, et al., 2014; Tarescavage, Fischler, et al., 2014). Regarding the HPC scale, past research has generally failed to demonstrate meaningful associations between this scale and extratest criteria among police officers. Rather, among the five somatic/cognitive Specific Problems Scales, only the MLS and COG scales have demonstrated replicable associations with problem behaviors in police officer samples. Of note in this context, whereas HPC, NUC, and GIC are relatively straightforward subsets of RC1 items, the other two Somatic/Cognitive Specific Problems scales (MLS and COG) are drawn from the broader item pool and are also closely related to the test’s emotional/internalizing and thought dysfunction domain scales, respectively (Tellegen & Ben-Porath, 2011).
Interpersonal Functioning
The interpreted scales in the Interpersonal Functioning domain include Restructured Clinical scale Cynicism (RC3) and the Interpersonal Specific Problems scales, which are composed of Family Problems (FML), Interpersonal Passivity (IPP), Social Avoidance (SAV), and Shyness (SHY). In this domain, RC3 was correlated with firearms misuse. FML was associated with unlawfully dropping a case, embezzlement, and falsification of time worked. SAV was negatively associated with alcohol abuse. Finally, SHY was positively associated with use of excessive force.
The association between RC3 and firearms misuse adds to the wide array of problem behaviors reported in the police psychology literature to correlate with RC3. Sellbom et al. (2007) found in their longitudinal study of police candidates that RC3 was significantly correlated with counterproductive behaviors as broad and diverse as citizen complaints, inappropriate language, rude behavior, uncooperativeness toward peers, deceptiveness, unlawful activity, failure to take responsibility for mistakes, use of one’s position for personal advantage, inappropriate sexual attitudes or behavior, conduct unbecoming, global deficits in integrity, and unwillingness to hire the officer again. That RC3 is correlated in the present study with only one problem behavior may derive from the comparatively low proportion of RC3 scale items (80%) contained in the original MMPI version. The correlations between FML with unlawfully dropping a case, embezzlement, and falsification of time worked converge with the results of Tarescavage, Fischler, et al. (2014), who found that this scale was meaningfully correlated with deceptiveness. Of note in this context, though FML is included in the Interpersonal Functioning domain, it is also meaningfully correlated with the MMPI-2-RF externalizing scales, particularly RC4 (Tellegen & Ben-Porath, 2011). Finally, the negative association between SAV and alcohol abuse converges with the findings of Tarescavage, Brewster, et al. (2014). Ben-Porath and Tellegen (2011) note that low scores on SAV are associated with enjoying social situations and events. Relatedly, the item content of SAV overlaps considerably with a subset of items scored on INTR-r that, as discussed earlier in this section, likely contribute to negative associations between the INTR-r scale and alcohol abuse.
At this juncture, there is no previously published correlate data to suggest that the SHY scale is associated with use of excessive force among police officers. As discussed earlier, however, similar to officers who lack an instrumental orientation (as indicated by low AGGR-r), it may be that police officers with social anxiety may confront suspects too slowly and inadequately, thereby allowing their behavior to escalate to a point that ultimately triggers an overreaction by the officer. Nevertheless, further investigation of this possible association is needed.
Relative Risk Ratios
In Table 5, we present RRRs that met our previously described selection criteria (i.e., those that had statistically significant zero-order correlations and yielded selection ratios ranging from 2.0% to 20%). In the interest of space, we only report the 33 statistically significant findings from the total of 77 RRR analyses. To assist the reader with interpretation, we provide a description of the RRR for estimated EID and use of excessive force (i.e., the first row of Table 5). The selection ratio (SR) indicates that 3.4% of the sample scored at or above a cutoff of 50T. The base rate (BR) indicates that 8.5% of the sample was found to have used excessive force. The risk for this outcome if EID is ≥50T is 33.3%, and if EID is <50T the risk is 7.6%. Dividing the risk if elevated by the risk if not elevated yields an RRR of 4.359. Because the 95% confidence interval for this analysis (1.25 to 15.16) does not include 1.0, the finding is statistically significant.
Male Officer Estimated Minnesota Multiphasic Personality Inventory-2–Restructured Form Statistically Significant Score Relative Risk Ratios With Violations and Problem Behaviors.
Note. SR = selection ratio; BR = base rate; RRR = relative risk ratio; CI = confidence interval; EID = Emotional/Internalizing Dysfunction; THD = Thought Dysfunction; RCd = Demoralization; RC2 = Low Positive Emotions; RC4 = Antisocial Behavior; RC6 = Ideas of Persecution; RC8 = Aberrant Experiences; MLS = Malaise; GIC = Gastrointestinal Complaints; HPC = Head Pain Complaints; BRF = Behavior-Restricting Fears; MSF = Multiple Specific Fears; AGG = Aggression; FML = Family Problems; SHY = Shyness; AGGR-r = Aggressiveness–Revised; PSYC-r = Psychoticism–Revised; DISC-r = Disconstraint–Revised. For all analyses, ns range from 167 to 247. All RRRs are significant at an alpha of .05. Table 5 presents the 33 statistically significant RRRs (77 total).
Overall, the RRR analyses demonstrated meaningful findings for a variety of scales at cutoffs of 50T and 55T, which yielded selection ratios ranging from 3.4% to 19.7%. Moreover, the RRRs indicated substantially increased risk for problems when scales were elevated, with RRRs ranging from 1.9 to 12.0. For example, individuals with elevations at 55T or above on RC4 were 1.9 times more likely to have off-duty violations and those with elevations at 55T on PSYC-r were 12.0 times more likely to have a violation involving alcohol abuse.
The selection ratios (also known as elevation rates) for the presented RRR analyses were generally consistent with the base rates of violations, which is necessary to minimize false positive findings. For example, 27% of individuals scoring 45T or higher on HPC had a subsequent violation relating to use of excessive force compared with 7% of individuals scoring below the cutoff. Thus, police officers who, in their preemployment evaluation scored at or above 45T on HPC, were at a 3.9 times greater risk of subsequently using excessive force when compared with those who scored below the cutoff. Despite use of the relatively low cutoff of 45T, the elevation rate at this level was 8.5%, which was exactly the same as the base rate for excessive force violations.
General Discussion
The purpose of this study was to extend the findings of the seminal Boes et al. (1997) investigation to the MMPI-2-RF using a new method for estimating MMPI-2-RF scores from MMPI responses. In Study 1, we found that for most scales, prorated scores were highly correlated with actual scores in the MMPI-2-RF normative sample, as well as the psychiatric inpatient and male police candidate comparison groups. We also found that estimated T-scores were generally comparable to actual T-scores in these samples. In Study 2, descriptive analyses indicated that there were minimal differences between officers’ estimated MMPI-2-RF T-scores in the current sample and the male police officer candidate comparison group’s T-scores. MMPI-2-RF scale scores, particularly in the emotional, behavioral, and thought dysfunction domains, were associated with a number of integrity violations and other problem behaviors. These associations were generally of a meaningful magnitude, especially after they were disattenuated for range restriction. These correlations led to the identification of statistically significant RRRs, obtained for cutoffs ranging from 45T to 65T. Several aspects of these findings warrant discussion.
From an assessment science perspective, Study 1 demonstrates that most of the MMPI-2-RF scales can be estimated from MMPI responses for the purpose of evaluating the construct validity of the instrument. We relied on scores derived from the original MMPI because of the benefits of using the unique data collected by Boes et al. (1997). We believe that fidelity checks presented in the current study and past research on MMPI/MMPI-2/MMPI-2-RF score comparability outweigh two potentially important objections to using data collected with the original version of the inventory. Namely, the MMPI items appear in a different order than they do on the MMPI-2-RF, and 35 of the 274 MMPI items that are included in the MMPI-2-RF booklet have been reworded owing to outdated content, sexist language, and sentence structure problems. Both of these factors could affect score comparability. However, we can infer from past MMPI-2/MMPI-2-RF comparability research that item order does not affect scores on broadband measures (Tarescavage, Alosco, et al., 2014; Tellegen & Ben-Porath, 2011; van der Heijden, Egger, Rossi, van der Veld, & Derksen, 2013), and MMPI/MMPI-2 item comparability research demonstrated that the vast majority of rewritten items have similar psychometric properties when compared with the originally worded items (Ben-Porath & Butcher, 1989).
Nevertheless, future research should explore the issue of score comparability directly, using a within-subjects design, in which MMPI-2-RF scores are calculated from both MMPI and MMPI-2-RF administrations to the same individuals in a counterbalanced or random design. Similar research designs have been used to demonstrate the comparability of MMPI-2-RF reliability estimates and scores from MMPI-2 and MMPI-2-RF administrations (Tellegen & Ben-Porath, 2011; Van der Heijden et al., 2010). Future investigations of the comparability of MMPI-2-RF validity coefficients from the two booklets would also be informative.
The correlations calculated in Study 2 yielded meaningful associations between all measurement domains of the MMPI-2-RF and integrity violations, although the effect sizes were modest. However, at least two extraneous factors may have attenuated the observed correlations. First, although the prorating method yielded comparable scores for 29 of 38 reduced subsets used in Study 2 (along with 13 scales with all items in the MMPI booklet), these scores were less reliable because the subsets had fewer items, which attenuates correlation coefficients. Second, in Study 2 we found that the current sample produced range restricted scores. The average standard deviation across the estimated MMPI-2-RF substantive scales was approximately two thirds that of the general population. Though the correction for range restriction yielded significant increases in estimated validities of preemployment MMPI-2-RF scores as predictors of subsequent police officer integrity violations and other negative outcomes, significance testing was conducting on the zero-order correlations to reduce Type I error. Therefore, the number of interpreted associations was reduced. The effect of these extraneous factors on the findings of the current study, notwithstanding the magnitude and volume of interpretable correlation coefficients, is consistent with other studies where problem behaviors and integrity violations are utilized as criteria among police officer candidates (Sellbom et al., 2007; Tarescavage, Corey, & Ben-Porath, 2015).
Range restriction similarly affected the MMPI Clinical Scales in the Boes et al. (1997) study, from which our data were derived. This may, in part, account for the minimal number of meaningful associations between integrity violations/problem behaviors and the MMPI Clinical Scales reported by Boes et al. (1997), who did not correct for range restriction. However, in this context it is worth noting that Sellbom et al. (2007) did report correlations with outcome variables corrected for range restriction in both the MMPI-2 clinical and RC scales and found substantially stronger predictions of negative outcomes for the psychometrically improved RC scales. Further confounding the analyses of Boes et al. (1997) was that, as discussed earlier, their “nonviolator” group included a substantial number of individuals who had committed other violations that were not the focus of their investigation. These individuals were removed from our analyses.
Consistent with Boes et al. (1997), who found associations between Lie scores and integrity violations, L-r was associated with two integrity violations (information breach aiding criminals and receipt of protection money). However, contrary to the findings of Boes et al. (1997), the MMPI-2-RF substantive scales clearly outperformed L-r in predicting negative outcomes. Because the L-r scale is a measure of underreporting and there is incentive for police officers to minimize psychological problems to obtain employment, this scale is not range restricted, which may partly account for the relatively stronger findings for this scale in the Boes et al. (1997) study in comparison to the substantive scales (i.e., which were range restricted). Relatedly, Boes et al. (1997) concluded that Clinical Scale 4 (Psychopathic Deviate) was the only clinical scale that showed promise for identifying future violators. The current study’s findings were stronger for the restructured version of this scale and, in particular, JCP—especially after accounting for range restriction. However, it is important to note that the JCP RRR analyses were nonsignificant.
Our study design has limitations that warrant discussion. First, we did not calculate correlations or RRRs between estimated MMPI-2-RF scale scores and criteria with less than a 2.0% base rate, owing to concerns that outliers may lead to spurious findings. Some of these low base rate criteria involved more severe integrity violations, including discrimination, criminal violations, and information breaches that endanger officers. Relatedly, even variables with higher base rates (e.g., Alcohol Abuse: 4.7%) had relatively low absolute numbers of violators (8 out of 169 in the just mentioned example). This demonstrates one difficulty associated with field research in this area, where some problem behaviors are uncommon owing to the intensive screening process. Thus, future research with larger samples and/or higher base rates of these and similar issues is needed. Along the same lines, we were unable to conduct analyses using a female sample because of the small number of women with MMPIs in the Boes et al. (1997) data set. Because of our modifications to the research design of Boes et al. (1997), which included removing individuals with any history of integrity violations or disciplinary actions from the previously contaminated nonviolator control group, the sample size decreased meaningfully. Though this study characteristic limited our statistical power, and relatedly our ability to implement alpha-adjustment procedures to mitigate the possibility of Type I errors, we suggest that the benefits of using criteria with improved validity outweigh these costs. Relatedly, because of these issues and the exploratory nature of the current study, further integration with other studies in this area is needed. Finally, the Boes et al. (1997) data were collected in the early 1990s and an updated study of this magnitude is needed.
These limitations notwithstanding, the current study extends the findings of Boes et al. (1997) by using a method to score estimated MMPI-2-RF scales from MMPI item responses. Our results demonstrate that the MMPI-2-RF scales are associated with integrity violations, especially after correcting for range restriction, which attenuates correlation coefficients. More broadly, the test provides information relevant to the goal of preemployment psychological evaluations, which is identification of police candidates at increased risk for problems that may jeopardize public safety and trust.
Footnotes
Authors’ Note
Points of view expressed in this article are those of the authors and do not necessarily represent the official positions or policies of the U.S. Departments of Justice or Defense.
Declaration of Conflicting Interests
The author(s) declared the following potential conflicts of interest with respect to the research, authorship, and/or publication of this article: Yossef Ben-Porath and David Corey receive research funds from the MMPI-2-RF publisher and, as coauthors of the MMPI-2-RF Police Candidate Interpretive Report, they receive royalties on sales of the report. Yossef Ben-Porath is a paid consultant to the MMPI publisher, the University of Minnesota, and distributor, Pearson. As coauthor of the MMPI-2-RF, he receives royalties on sales of the test.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: The data collection was partially funded by Grant No. 96-IJ-CX-A056 awarded by the National Institute of Justice, Office of Justice Programs, U.S. Department of Justice.
