Abstract
The current article reports on the development, psychometric properties, and external validity of an informant-report form of the Personality Inventory for DSM-5 (the PID-5-IRF). Using data from two nationally representative samples, as well as an elevated-risk community sample, we report on the PID-5-IRF item characteristics, scale properties, superordinate factor structure, and correlations with other measures. The PID-5-IRF replicates the factor structure of the self-report form and has relationships with other measures (including the PID-5 self-report form and a widely used Big Five measure) that are consistent with previous research and theory. We believe that the PID-5-IRF is a useful measure for a number of scenarios, such as when additional sources of information are desired, where informant measures are expected to provide incremental validity over self-report, where relationships or social perception is a focal interest, or when response bias is a salient concern. Areas for future research are also discussed.
Like many personality pathology inventories, the Personality Inventory for DSM-5 (PID-5; Krueger, Derringer, Markon, Watson, & Skodol, 2012) was initially developed as a self-report instrument. This is arguably a reasonable starting point, as patient self-report in one way or another forms the basis for most clinical assessment, and is integral to current clinical and research practice. Fundamentally, self report provides a sample of behaviors from the individual being assessed, which can be interpreted purely in terms of its predictive value (Meehl, 1945/2000). At another level, the content of self report itself becomes critically relevant in certain scenarios—for example, when assessing narcissistic personality features, assessment of self-perception would arguably be central to a valid assessment protocol (cf. Haynes, Richard, & Kubany, 1995). Relatedly, evidence suggests that self report might be more valid than reports from other sources when assessing certain constructs, such as those involving mood or other internal states that are not easily observed by others (Funder, 1995; Vazire, 2010; Vazire & Carlson, 2011).
Although self report is critically important for numerous reasons, there are occasions when informant-report information is helpful or even critical. For example, considered solely as another source of information, informant report provides another measurement on the individual being assessed, which can be used to increase the reliability and generalizability of conclusions through aggregation (Rushton, Brainerd, & Pressley, 1983). In this regard, an informant-report measure can be treated as another important indicator of a construct of interest, which may be useful in certain scenarios.
Substantial literature suggests, however, that informant-report measures also provide important unique information about individuals. Just as self-report measures of internal states may be more accurate than informant-report measures, informant-report measures may be more accurate indices of traits that are highly observable by others, or involve a substantial evaluative component (e.g., intelligence; Carlson, Vazire, & Oltmanns, 2013; Vazire, 2010; Vazire & Carlson, 2011). Relatedly, meta-analyses suggest that informant reports may be more externally valid than self reports for certain constructs or criteria, especially those criteria involving substantial social factors (Connelly & Ones, 2010; Duckworth & Kern, 2011; Oh, Wang, & Mount, 2011). Although both self report and informant report possess comparable levels of validity overall (Vazire, 2010), there may be certain constructs or settings where informant report might be more accurate or informative.
Furthermore, informant-report information might also be important in itself, as it may provide important information about how particular individuals of importance to a target perceive them and their relationship (e.g., spousal reports of a client’s personality when the relationship is a focus). Social cognitive research, for example, suggests that accurate perceptions of a partner’s personality, as well as biases (e.g., related to perceived positivity or as similarity), both predict relationship satisfaction (e.g., Luo & Snider, 2009). Obtaining information from informants is critical to understanding social processes and functioning, in tandem with and independently of self-report data.
Finally, informant report may also help provide information in the presence of response style or response bias (e.g., impression management). Research findings have raised important questions about the interpretability of traditional validity indices in personality assessment, in that these indices do not always moderate the prediction of external criteria, leading some authors to challenge their validity or utility (e.g., McGrath, Mitchell, Kim, & Hough, 2010; Piedmont, McCrae, Riemann, & Angleitner, 2000). Although these findings remain controversial and alternate interpretations remain possible (e.g., Edens & Ruiz, 2006; Markon, 2012; Rohling et al., 2011; Ziegler, MacCann, & Roberts, 2012), having additional means of assessing for response biases is desirable. This is particularly critical in the case of the PID-5, which is freely available and therefore particularly susceptible to coaching. For these reasons, an informant-report form of the PID-5 would be particularly useful as a way of assessing individuals when response biases are of concern.
Here, we describe the development and psychometric characteristics of the PID-5 informant-report form (PID-5-IRF), an informant-report version of the PID-5. Like the PID-5 self-report form (PID-5-SRF), the PID-5-IRF was developed under the auspices of the American Psychiatric Association, and the APA holds the copyright on this instrument; however, like the PID-5-SRF, the PID-5-IRF will be freely available for clinical and research use through the DSM-5 website. We provide information on the method of its development, its factor structure and other psychometric characteristics, and correlates with the self-report form and other measures.
Method
Participants and Procedure
Three samples were used to develop the informant form: two normative samples, and one elevated-risk community sample. The two normative samples were designed to be representative of the general U.S. population. The primary normative sample was used for measure development; the secondary sample was used to obtain information on self–informant concordance. The two normative samples also differed in informant recruitment strategies—in the primary sample, informants were asked to identify targets, and in the secondary sample, informants were identified based on targets—and thus were intended to provide generalizability across different mechanisms of informant recruitment. The community sample was used for measure development (i.e., to increase the total sample size available, and to cross-validate results), and to examine characteristics of the measures with regard to external variables of interest. Summaries of the samples are provided here; additional details regarding the samples are presented in supplementary Table A1 in the appendix at http://asm.sagepub.com/supplemental.
Normative Samples
Data from the normative samples were collected by Knowledge Networks (KN), an online survey and social research firm involved in development of the PID-5-SRF (Krueger et al., 2012). Participants in the normative sample were recruited from KN’s online research panel of approximately 50,000 individuals. This panel, representative of the U.S. national population, was recruited using a combination of random-digit dialing and addressed-based sampling, with individuals being recruited by mail and phone. Individuals on the panel participate in exchange for various financial incentives; those without Internet access are provided with Internet access in exchange for participating in the panel. Further details are available in technical articles provided at the KN website (e.g., Dennis, 2010; Knowledge Networks, 2011).
Primary sample
The primary normative sample included 320 adult respondents who were asked to respond about “an adult they know well.” Respondents were 51.9% female and 50.363 years old on average (SD = 17.175), and the targets were 50.3% female. Respondents reported knowing targets 22.441 years on average (SD = 15.554). In all, 38.1% were friends; 36.9% were spouses, partners, or significant others; 16.6% were relatives or family members; 3.8% were coworkers or acquaintances; and 4.6% had some other form of relationship or declined to describe the nature of the relationship.
Secondary sample
The secondary normative sample included 40 adults who responded about an individual who was part of the PID-5-SRF development sample. Informants were identified as KN panel members in households of PID-5-SRF normative development sample participants; the original SRF participants were asked for their permission to be rated by the informants (26.9% of eligible SRF participants declined to be rated by the identified informant). Respondents were 57.5% female and 51.125 years old on average (SD = 16.223), and the targets were 47.5% female. Respondents reported knowing targets 28.975 years on average (SD = 14.814). In all, 60.0% were spouses, partners, or significant others; 32.5% were relatives or family members; 2.5% were friends; and 5.0% had some other form of relationship or declined to describe the nature of the relationship.
Community Sample
The elevated-risk community sample included 221 adult targets, recruited via online and print advertisements for an investigation of “gambling and personality” as well as via the research registry of both patients and nonpatients interested in future research participation at the Centre for Addiction and Mental Health in Toronto, Canada. Targets attended two assessment sessions to complete interview and questionnaire measures of personality and psychopathology; targets identified and brought an informant of at least 1 year’s acquaintance to the second session. Targets were 45.2% female and 41.727 years old on average (SD = 12.755); informants were 48.4% female and 41.348 years old on average (SD = 14.033). Informants reported knowing targets 13.927 years on average (SD = 12.928). Overall, 23.1% were romantic partners of targets, 22.6% were family members, 52.5% were friends, and the remainder described themselves as having another association with the target (e.g., coworker). Perceived accuracy of ratings and closeness of acquaintance were rated on a 5-point Likert-type scale ranging from 1 to 5, and were generally high: accuracy M = 4.322, SD = 0.669; closeness M = 4.442, SD = 0.671. Current Axis I diagnoses as assessed by the Structure Clinical Interview for DSM-IV, Axis I Disorders, Patient Form (First, Spitzer, Gibbon, & Williams, 1995) were found in 45.7% of targets (n = 101), and included the following conditions: mood disorders (n = 30), psychotic disorders (n = 8), substance use disorders (n = 10), anxiety disorders (n = 33), and eating or adjustment disorders (n = 3). In addition to the PID-5, targets and informants completed Form S (self-rated) and Form R (observer-rated) of the Revised NEO Personality Inventory (NEO PI-R; Costa & McCrae, 1992), respectively.
Item Development
Initial candidate PID-5-IRF items were constructed by modifying the 220 self-report items to be in the third person (e.g., replacing “I” with “he or she”). Items had the same four-point response format as the self-report form. This initial set of items was then reviewed by members of the research team and clinical psychology graduate students, who rated each item for suitability. Based on these ratings, items deemed to be unsuitable (e.g., because of odd phrasing, problems with perspective, or other potential difficulties in responding to the item) were revised or replaced in a consensus process. For example, the self-report item “I’ve achieved far more than almost anyone I know” was replaced with “[He or she] exaggerates their own achievements.” In addition, one candidate self-report item that was eventually dropped from the self-report form was retained in its third person form (“[He or she] finds it rarely worth it to take risks”) as a possible informant-report item. This resulted in a potential pool of 221 candidate items.
Results
Scale Construction and Item Selection
Item-Level Scale Structures
Items were assumed a priori to be structurally organized into item sets based on their scale membership in the self-report form. The dimensionality of each item set was verified by permutation-based parallel analyses (Buja & Eyuboglu, 1992) in the primary normative and community data sets. The results of these parallel analyses supported a one-factor structure for all scales in both data sets, with two exceptions. The first exception was the Risk Taking scale, which parallel analyses suggested had a two-factor structure in both data sets. The second exception was the Depressivity scale, which in the normative sample parallel analyses suggested had a one-factor structure, but in the community sample suggested had a two-factor structure.
Exploratory factor analyses (EFAs) with maximum likelihood (ML) estimation and geomin rotation suggested that the Risk Taking factors reflected item keying factors. EFAs suggested that the second Depressivity factor in the community sample was a suicidality-specific factor; however, this factor was not clearly evident in the normative sample, even in a two-factor model. We did run bifactor item response theory (IRT) models with distinct keying factors for all scales with reverse-keyed items, and a suicidality factor for the Depressivity scale. IRT parameter estimates for the superordinate factors in these models were extremely similar to those under one-factor models, however, and did not change conclusions about item selection. For the sake of simplicity, we present results for the one-factor models here. As discussed later, however, the presence of minor keying and suicidality factors raises the possibility of scoring for response style factors and self-harm risk, which may be of interest in various settings.
Item Selection
IRT analyses were conducted in Mplus (Muthén & Muthén, 1998-2011) using ML estimation of a cumulative logit model (i.e., graded response model). Analyses were conducted separately in the primary normative sample and community samples, as well as in a combined multiple-group model. In this multiple-group model, item parameters were fixed equal across the normative and community samples; latent means and variances were fixed to 0 and 1, respectively, in the normative sample, and were freely estimated in the community sample.
Results of the IRT analyses suggested that all items performed satisfactorily, with the exception of three items (on the Anxiousness, Suspiciousness, and Risk Taking scales). These three items all had unacceptably low discriminations (e.g., less than 0.5 on a logit scale or 0.3 on a standardized scale) in both samples. All were reverse-keyed and included the word “rarely.” Because of their low discriminations, they were dropped from the final inventory. This left a final informant-report form containing 218 items; these items are presented in the appendix at http://asm.sagepub.com/supplemental.
Psychometric Properties
Item Properties
IRT parameter estimates from the multiple-group model are presented in supplementary Table A2 in the appendix at http://asm.sagepub.com/supplemental. In general, within a scale, reverse-keyed items tended to have lower discriminations (i.e., loadings) than non–reverse-keyed items. Also, in general, the most severe items (i.e., those with the largest thresholds, indicating the most rare levels of greater endorsement) were those in the Psychoticism and Antagonism domains, especially the Callousness and Deceitfulness subscales, although certain items in other subscales (notably within the Depressivity subscale) were comparably severe. Overall, across subscales, greater item severity was associated with greater discrimination (i.e., across subscales, the correlation between discrimination and each of the three thresholds was .482, .642, and .829).
Scale Properties
Scale properties in the primary normative and community samples, including sample means and standard deviations, are presented in Table 1. Means were calculated as the mean item response across items within a scale, where the items were coded 0, 1, 2, 3. Means and standard deviations in the primary normative sample were calculated using sampling weights provided by KN, to improve estimates of means and standard deviations in the U.S. population. Also included in Table 1 are d values comparing the community with normative samples, using the normative sample standard deviation in calculation of d. The final column includes the p value corresponding to the d statistic, testing against the null hypothesis of no difference between the samples. As can be seen, consistent with what one might predict, the community sample means were generally greater than the normative sample means. This is especially true of scales related to disinhibition and to a lesser extent psychoticism.
Scale Descriptive Statistics and Psychometric Properties.
Note. Values in table are normalized minimum reduction in uncertainty (NMRU), mean and standard deviation of item responses (with options coded 0-3), and ω and α as indices of reliability. The final columns give the difference between the primary normative and community samples as a d statistic, together with the corresponding p value.
Table 1 also includes reliability statistics in the two samples, specifically omega (ω) and alpha (α). Omega is a model-based reliability statistic reflecting the ratio of trait variance to total variance (e.g., Revelle & Zinbarg, 2009). As can be seen, all the scales had adequate reliabilities in both samples. Reliabilities for each scale were similar across samples. Omega ranged from .719 to .952 in the normative sample and .740 to .948 in the community sample, with means of .887 and .886, respectively. Alpha ranged from .718 to .952 in the normative sample and .738 to .948 in the community sample, with means of .884 and .882, respectively. As one might predict based on its relatively small number of four items, the Submissiveness scale had the smallest reliabilities in both samples. Eccentricity had the largest reliabilities in both samples.
Finally, Table 1 includes normalized minimum reduction in uncertainty (NMRU) statistics (Markon, 2013). NMRU provides an index of the total measurement information a test can provide, a summary of the test information function. Like the test information function, and in contrast to reliability, the NMRU is not sample specific. It varies from 0 to 1, with larger values representing more information; values of .45 to .50 or greater are generally typical of well-validated and reliable measures. The NMRU quantifies the reduction in uncertainty about a person’s trait level when a test is providing the most information possible. In this regard, it can be thought of as representing a test’s “measurement potential.” As shown in Table 1, all the NMRU statistics were adequate, and most were greater than .50; the mean NMRU value was .727. The rank ordering of the NMRU statistics is generally consistent with the reliabilities, moreover, with the smallest value for Submissiveness and the largest for Eccentricity.
Scale Structure
Given that the five-factor structure of the PID-5-SRF has replicated in a number of samples (e.g., De Fruyt et al., 2013; Krueger et al., 2012; Wright et al., 2012), we hypothesized that the informant-report form would as well. Parallel analysis and Velicer’s minimum average partial statistic (MAP) both generally suggested a three-factor structure (both suggested three factors in both samples, except for MAP in the community sample, which suggested four factors). However, the estimated five-factor model of the informant form was similar to that of the self-report form (congruence coefficients between the informant- and self-form loadings were .874 and .864, .899 and .905, .832 and .841, .864 and .878, and .910 and .908 for the normative and community samples, respectively, for the standardized loadings reported in Table 2). Because the five-factor structure was similar to that observed in previous studies, and was hypothesized a priori, we present results for it here.
Scale Structure: Five-Factor Exploratory Structural Equation Modeling Parameter Estimates.
Note. Estimates were obtained using maximum likelihood (ML) estimation with equamax rotation, for a multiple-group model as described in the text. Values in boldface are largest loadings in each row for each sample.
This five-factor superordinate structure is illustrated in Table 2. The loading estimates in the table were obtained under multiple-group EFAs using exploratory structural equation modeling. In this approach, EFA models were fit under the constraint that the loadings and intercepts were equal across groups, with latent factor means and variances fixed to 0 and 1 in the normative sample, respectively, and free in the community sample. Estimates were obtained using ML estimation with equamax rotation (for comparison with the SRF, which was reported using equamax rotation; Krueger et al., 2012). The root mean square error of approximation (RMSEA) for the estimated model was .095 and root mean square residual (RMSR) was .044.
As is evident in Table 2, the five factors are essentially the same as those underlying the self-report form. The first factor, Negative Affect, was characterized by loadings on subscales such as Anxiety, Emotional Lability, and Depressivity. The second factor, Antagonism, was characterized by loadings on subscales such as Attention Seeking, Manipulativeness, Grandiosity, Deceitfulness, and Callousness. The third factor, Disinhibition, was characterized by loadings on subscales such as Distractibility, Impulsivity, and Irresponsibility. The fourth factor, Psychoticism, primarily loaded on the Perceptual Dysregulation, Unusual Beliefs, and Eccentricity subscales. The fifth factor, Detachment, was characterized by loadings on the Withdrawal, Anhedonia, Restricted Affectivity, and Intimacy Avoidance subscales.
Self–Informant Congruence
Self–Informant Correlations
Self–informant correlations are reported in Table 3. Facet- and domain-level correlations were calculated using unit-weighted sum scores; domain scores were calculated using those facets that had the greatest loading in the domain for both the IRF and SRF. The self–informant correlations were relatively large for the Disinhibition and Antagonism domains, and smaller for the Psychoticism domain. Correlations involving the Detachment domain were relatively moderate, as were those involving the Negative Affect domain.
Self–Informant Correlations.
Note. Values are self–informant correlations in secondary normative and community samples, and p values for a test of the null hypothesis that the correlations are equal.
In general, patterns of observed scale correlations were similar between the normative and community samples (the correlation between correlations was .501 for the facets and .986 for the domains). However, the community correlations were generally slightly smaller than the normative correlations (mean correlations were .355 vs. .468 for the facets and .290 vs. .494 for the domains). The greatest differences between the normative and community sample self–informant correlations were generally in the Disinhibition domain, and for the Emotional Lability, Callousness, Separation Insecurity, Restricted Affectivity, Perfectionism, and Withdrawal facets.
The self–informant scale correlations also provided evidence of discriminant validity. For example, in the normative sample, the median absolute values of the convergent and discriminant self–informant correlations were .481 and .300 and .503 and .215 for the facets and domains, respectively. Similarly, in the community sample, the median absolute values of the convergent and discriminant self–informant correlations were .336 and .200 and .299 and .175 for the facets and domains, respectively. The patterns of discriminant validity were also consistent with theoretical predictions, moreover, in that the larger discriminant validities tended to be within domains or between related domains, especially within the Antagonism and Disinhibition domains. The largest absolute discriminant facet correlation in the normative sample (.590), for example, was between self-rated Deceitfulness and informant-rated Manipulativeness; in the community sample, the largest absolute discriminant facet correlation (.548) was between self-rated Impulsivity and informant-rated Irresponsibility. Similarly, in both the normative and community samples, the largest discriminant domain correlations (.508 and .398, respectively) were between Antagonism and Disinhibition. At the domain level, Psychoticism demonstrated the least discriminant validity, in that in both samples, the median discriminant correlations for that domain (.324 and .271 for the normative and community sample) were more similar to the convergent correlations (.379 and .183) than was the case for the other domains.
Self–Informant Profile Similarities
Self–informant profile similarities were examined to determine the degree of idiographic congruence between self and informant ratings. Standard self–informant correlations include the effects of response styles and scale use factors, and might obscure the tendency for targets and informants to agree on the degree to which a trait is elevated relative to other traits for that individual, even if they disagree on that individual’s standing on the trait relative to other individuals. For example, an informant and target might agree that hostility is the most severe problem for the target, even though they disagree on the target’s level of hostility relative to other individuals.
Self–informant profile similarity was examined by calculating intrapair profile correlations (i.e., the correlation between the self and informant ratings across traits for a single individual). This produced a correlation for each target–informant pair. The descriptive statistics for these correlations are presented in Table 4. In general, as with the nomothetic self–informant correlations, the intrapair correlations were generally larger for the normative sample (mean correlations of .499 and .518 for the facets and domains) than community sample (.374 and .470).
Intrapair Profile Correlations: Descriptive Statistics.
External Validity
PID-5-IRF correlations with NEO PI-R Forms R (observer-rated) and S (self-rated) domain scores are reported in Table 5. Correlations ≥|.300| are presented in boldface (Cohen, 1988), and correlations between analogous NEO PI-R and PID-5 domains and their associated facets are underlined to facilitate review. NEO PI-R self–informant correlations in the community sample ranged from .401 to .539 for domains (M = .474, SD = .055), and from .251 to .546 for facets (M = .473, SD = .075). In comparison, NEO PI-R self–informant correlations reported in the manual ranged from .36 to .53 for domains (M = .43, SD = .06), and from .16 to .52 for facets (M = .35, SD = .08; Costa & McCrae, 1992). There was no significant difference between the average domain and facet correlations in the community sample as compared with the NEO PI-R (Fisher’s z = .54, p = .59, and z = 1.55, p = .12, respectively).
Correlations Between PID-5-IRF Scale Scores and Revised NEO Personality Inventory Domain Scores in the Community Sample.
Note. R = Revised NEO Personality Inventory Form R (observer-rated); S = Revised NEO Personality Inventory Form S (self-rated). Correlations >|.13| are significant at p < .05, correlations > |.18| are significant at p < .01, Correlations ≥|.30| are given in boldface, correlations between analogous NEO PI-R and PID-5 domains (and associated facets) are underlined.
Form R correlations with IRF scales were larger than those between the Form S domains and the IRF scales (mean correlations .310 vs. .205 for Form R vs. Form S). To evaluate the similarity in the pattern of correlations across Forms R and S, we computed column–vector correlations for each of the five domains. Specifically, we transformed the correlations using Fisher’s r-to-z formula and then computed the correlation between the two columns of transformed correlations. The pattern of correlations for Neuroticism (.914), Extraversion (.896), and Conscientiousness (.913) were very high across the versions of the NEO PI-R. In contrast, patterns differed across the forms for Openness (.344) and Agreeableness (.272).
Neuroticism demonstrated medium to strong correlations with Negative Affect and its associated facets; this domain further demonstrated associations of this magnitude with the other PID-5-IRF domains and numerous facets (18 of 25 facets for Form R and 13 of 25 for Form S). Similarly, Agreeableness was moderately to strongly correlated with Antagonism, Callousness, Deceitfulness, and Grandiosity; and Conscientiousness with Disinhibition, Distractibility, Impulsivity, and Irresponsibility. These domains were also associated with all PID-5-IRF domains with at least a medium effect size; Form R Agreeableness and Conscientiousness were further associated with 19 and 16 facet scales, whereas Form S domains were associated with a more restricted set of facet traits. In contrast, Form R Extraversion demonstrated a more specific pattern of association with Detachment and the facets that have been associated with this domain. Openness was minimally associated with Psychoticism, and Form R demonstrated medium associations with several scales associated with Detachment. Thus, evidence for concurrent validity was evidenced by the fact that Neuroticism, Extraversion, Agreeableness, and Conscientiousness all had their strongest relationships with analogous and theoretically relevant scales (e.g., NEO-PI-R Form R Neuroticism was most strongly correlated with PID-5-IRF Anxiousness and Depressivity; Extraversion was most strongly correlated with PID-5-IRF Withdrawal and Anhedonia; Agreeableness was most strongly related to PID-5-IRF Hostility and Callousness; and Conscientiousness was most strongly related to Irresponsibility and Distractibility). Evidence for discriminant validity was varied, however (e.g., PID-5-IRF Grandiosity and Risk Taking were moderately and specifically associated with Agreeableness and Conscientiousness, respectively, whereas other scales such as Anhedonia and Irresponsibility were associated with three or four domains of the five-factor model).
Discussion
Although self-report will continue to occupy a central role in assessment of personality pathology in clinical and research settings, there are situations in which informant-report measures become critically valuable. The PID-5-IRF in this way complements the self-report form of the PID-5, providing informant-report measures that parallel the PID-5 in content and structure. As has been demonstrated, the PID-5-IRF scales generally have adequate psychometric properties and demonstrate external validity in their relationships with other scales and internal validity in their superordinate structure, which mirrors that of the self-report form.
Implications for Understanding PID-5 Structure
In addition to having numerous applied measurement uses, the PID-5-IRF provides another replication of the superordinate structure of the PID-5, but in the context of informant ratings. As has been observed in multiple studies of self-report (e.g., De Fruyt et al., 2013; Krueger et al., 2012; Wright et al., 2012), the five factors of Negative Affect, Antagonism, Disinhibition, Detachment, and Psychoticism were clearly evident in the PID-5-IRF structure. This finding replicates and extends the superordinate structure of the PID-5 scales, and underscores the utility of the PID-5 across multiple sources of information.
The current results also support the subordinate structure of the PID-5 scales. Factor-analytic results reported here support the coherence and unidimensionality of the PID-5 scales, as parallel analyses strongly supported a unidimensional interpretation of all but two of the subscales. These findings support interpretation of the subscales in terms of reasonably homogenous content domains, across multiple informant modalities.
There were certain important novel features of our results that merit further discussion. For instance, the cross-loading patterns of some PID-5-IRF scales with regard to the superordinate factors differed from what was reported in the PID-5-SRF development article. In the current results, for example, PID-5-IRF Depressivity, had loadings primarily from the negative affectivity and secondarily from detachment, in contrast to previous studies (Krueger et al., 2012; Wright et al., 2012) where the opposite pattern was observed. Restricted Affectivity, similarly, had prominent loadings from detachment, and minimal loadings from negative affect, which contrasts with previous studies (Krueger et al., 2012; Wright et al., 2012).
Further research is needed to determine the extent to which these differences in cross-loading patterns reflect sampling variation or replicable differences in the determinants of self- versus informant-report. One limitation of the current study is that sample-size considerations prevented more sophisticated multitrait multimethod modeling to evaluate these sorts of hypotheses (e.g., formal measurement invariance testing across self and informant forms; cf. Schimmack, 2010). The fact that some PID-5-IRF cross-loading patterns (e.g., for Depressivity) were observed across two samples suggests that there may be meaningful differences in how superordinate factors contribute to responses in self-report versus informant-report measures. Depression, for example, is related to both positive and negative emotions, and it may be that the salience of these two emotion factors, through detachment and negative affect, respectively, changes across self- and informant-report, with positive emotionality being a more salient determinant of Depressivity responses in self-report, but negative emotion being a more salient determinant of Depressivity for informant report. Perfectionism, similarly, appeared interstitial in nature, having cross-loadings from the Negative Affect and Psychoticism factors, as well as Disinhibition to some extent. However, the relative magnitudes of the loadings from these factors appeared somewhat different for the informant- and self-report versions of the PID-5 (e.g., with Perfectionism having greater loadings from Psychoticism and lower loadings from Antagonism, in informant report). The multidimensionality of PID-5 Perfectionism mirrors findings in the literature, which have found that Perfectionism and obssessive-compulsive–related traits variously behave as indicators of anxiety (Egan, Wade, & Shafran, 2011; Naragon-Gainey, 2011), extreme conscientiousness (Samuel & Widiger, 2008; Saulsman & Page, 2004), or thought disorder (Chmielewski & Watson, 2008). However, it may be that the relative importance of these traits in interpreting Perfectionism scores may differ in meaning between the informant- and self-report versions of the scale (cf. Clifton, Turkheimer, & Oltmanns, 2004). Overall, these results underscore observations (Hopwood & Donnellan, 2010) that unidimensional subordinate-level scales often exhibit heterogeneous relationships with superordinate factors. Cross-loadings are meaningful, and may change in their relative weight across different instruments, modalities, and information sources.
Also, although nearly all the subscales evidenced unidimensional structure, two evidenced some evidence for some multidimensionality. In the case of Risk Taking, analyses suggested the presence of keying factors, with reverse-keyed items having loadings from a keying subfactor. In the case of Depressivity, analyses in one sample suggested the presence of a suicidality subfactor (indexed by items 81, 118, and 176). Item parameter estimates with regard to superordinate factors were extremely similar in hierarchical (e.g., bifactor) and unifactorial analyses, suggesting that responses to those scales can be interpreted in terms of dominant Depressivity and Risk Taking factors. However, evidence for these subfactors suggests that supplementary scales or indices might be constructed, to reflect response styles or suicidal ideation. Reverse-keyed items are interspersed throughout the PID-5-IRF, but the Risk Taking scale comprises more of them than any of the other scales, which likely explains why keying is salient for that scale but not others. Indices of response style would most likely benefit from using items throughout the entire inventory. Further research is needed to determine how well these response style and suicidality factors replicate, and to determine how reliable and valid indicators of them can be optimally constructed.
Magnitude and Patterns of Self–Informant Congruence
In general, the self–informant scale correlations reported here were similar in magnitude or slightly larger than what is typically observed in the personality literature (Connelly & Ones, 2010), and comparable with what is typically observed in the psychopathology literature (Achenbach, Krukowski, Dumenci, & Ivanova, 2005). The median facet and domain correlations in the normative sample (.503 and .481, respectively) were slightly greater than what has been observed in other self–informant studies of personality pathology; the median facet and domain correlations in the community sample (.336 and .299) were slightly less (e.g., Klonsky, Oltmanns, & Turkheimer, 2002, reported median self–informant correlations of .36 and .47 across studies, depending on the type of personality pathology construct).
The intrapair correlations were also generally similar in magnitude to those that have been previously reported for the Big Five (Pelham, 1993; Vogt & Colvin, 2005). For example, compared with the average facet and domain profile correlations in the normative (.499 and .518) and community (.374 and .470) samples, Pelham (1993) reported an average intrapair Big Five profile correlation of .48; Vogt and Colvin (2005) reported average profile correlations ranging from .25 to .49. The intrapair correlations were also consistent with existing literature in that they were slightly larger on average than the typical nomothetic correlations (Pelham, 1993). This suggests that targets and informants may agree slightly more on the relative importance of different traits in describing an individual than they do that individual’s standing on the traits relative to other people. It is also possible that the larger intrapair profile correlations reflect the effects of controlling for response style, suggesting that to some extent nomothetic self–informant correlations are decreased by raters’ idiosyncratic use of response options (Pelham, 1993).
Although in the more normative personality literature, self–informant agreement is often greatest for measures related to extraversion or positive temperament (Connelly & Ones, 2010; Ready & Clark, 2002), the largest self–informant correlations observed here were in the Disinhibition and Antagonism domains. It is worth noting, however, that in studies of personality pathology, the greatest self–informant correlations are often observed for antisocial personality features (Klonsky et al., 2002), and in studies of general psychopathology self–informant concordance is often greater for externalizing psychopathology (e.g., Achenbach et al., 2005). The relatively large self–informant concordance for the Callousness scale (.667), for example, is similar to results that have been reported for aggression (.61; Ready & Clark, 2002) and psychopathy (.64; Miller, Jones, & Lynam, 2011). Considered in the context of previous findings, it might be speculated that the relative magnitude of concordance for self and informant reports depends on the relative observability and interpersonal consequences of the traits. In contexts where aggressive, disinhibited personality is more observable and has greater interpersonal consequences, self–informant correlations may be greater. Although many models of self–other agreement assume that source agreement might be lower for highly evaluative traits due to socially desirable responding (e.g., Funder, 1995), it should be emphasized that the motive to respond in a social desirable way is a function of the rater. In this regard, high disagreeableness may paradoxically not motivate socially desirable responding among individuals who are highly disagreeable and, therefore, have little regard for the consequences of their behavior (Miller et al., 2011). In this regard, the moderating effect of trait evaluativeness on self–informant agreement may differ across levels of the trait, depending on the trait.
The relatively low self–informant concordance observed among some scales is also consistent with previous literature. Psychoticism, for example, exhibited relatively low self–informant concordance, as well as relatively low discriminant self–informant validity compared with other scales. This finding, consistent with previous literature on schizotypal personality pathology (e.g., Klonsky et al., 2002; South, Oltmanns, Johnson, & Turkheimer, 2011), might be explained by a number of factors. One explanation, for example, is that the variance in actual positive psychotic experiences might be restricted in the current samples, decreasing the magnitude of observed correlations. This would be consistent with the observation that the Perceptual Dysregulation and Unusual Beliefs and Experiences scales, despite having NMRUs and reliabilities comparable with other scales, had the smallest variances among the measures overall. Another explanation is that the internal nature of positive psychotic experiences, coupled with their extreme low social desirability and idiosyncratic nature, may make them especially difficult for informants to detect, leading informants to rely on assumptions about the target’s experiences and behavior that nevertheless may not be accurate. Finally, it is possible that both targets and informants form reliable impressions based on cues that are only weakly related to the core processes underlying psychosis. This is consistent with some findings that neither self-reports nor informant reports of psychotic functioning correlate with behavioral functioning as well as clinician reports (Sabbag et al., 2011), as well as observations that positive psychotic symptoms (which constitute the PID-5 Psychoticism measures) do not correlate with neuropsychological status as well as negative psychotic symptoms (De Gracia Dominguez, Viechtbauer, Simons, Van Os, & Krabbendam, 2009).
Relationships With Other Personality Measures
The NEO PI-R has been linked to the domains of the PID-5-SRF (De Fruyt et al., 2013; Thomas et al., 2012). The results presented here extend understanding of those relationships to informant-report. Overall, NEO PI-R Form R scales demonstrated correlations of a greater magnitude with PID-5-IRF scales as compared with Form S, which was sensible in light of the shared source variance between the PID-5-IRF and the Form R.
Neuroticism demonstrated meaningful associations with traits across all PID-5-IRF domains; Extraversion and Conscientiousness were similarly broadly associated with pathological personality traits. This pattern of results may reflect the saturation of PID-5-IRF scales with personality impairment and distress, and the central importance of interpersonal difficulty and behavioral dysregulation to personality pathology. Indeed, previous research has documented the prominent role of Neuroticism and Conscientiousness in psychopathology in general (Kotov, Gamez, Schmidt, & Watson, 2010) and personality pathology in particular (Carlson et al., 2013). It is possible, however, that these “externalizing” personality traits are most reliably observed and rated by informants. Previous research suggests that the observability as well as the social desirability of traits can affect the utility of self- versus informant ratings (Funder, 1995; Vazire, 2010). It is further possible that the recruitment of the community participants, which was intended to ensure a sample with a fulsome range of gambling involvement and pathology, simultaneously ensured a sample with a fulsome range of personality traits relevant to impulse control and addiction and perhaps a more restricted range of other traits. Of note, although the association between Psychoticism and Openness might be predicted on the basis of some previous empirical findings regarding Openness and psychosis (e.g., Carlson et al., 2013; Saulsman & Page, 2004; Watson, Clark, & Chmielewski, 2008), it is discrepant in some ways from other models and findings relating Openness to psychosis and thought disorder (e.g., De Fruyt et al., 2013; DeYoung, Grazioplene, & Peterson, 2012; Markon, Krueger, & Watson, 2005; Thomas et al., 2012).
The associations between Neuroticism, Agreeableness, and Conscientiousness on one hand, and PID-5-IRF Negative Affect, Antagonism, and Disinhibition domains and facets on the other, do support the convergent validity of these scales. Yet numerous PID-5-IRF scales were broadly associated with NEO PI-R. Many of these results are consistent with extant evidence for the secondary loadings of PID-5 facets as well as the hierarchical structure of higher order personality traits (e.g., Krueger et al., 2012; Wright et al., 2012). Other results may reflect the limited capacity of the NEO PI-R to adequately assess the concurrent and discriminant validity of each PID-5-IRF facet scale. For example, Unusual Beliefs are not well represented in the NEO PI-R scales, and Intimacy Avoidance is likely to be only indirectly or partially captured by the content of the Agreeableness domain.
Directions for Future Research
Certain aspects of our results contribute to a growing body of evidence that can inform future revisions to the PID-5. Relative to most of the other PID-5-IRF scales, for example, in our results Submissiveness demonstrated a smaller NRMU, smaller reliabilities, greater residual variance in superordinate structural analyses, and smaller self–informant and criterion correlations. Similar phenomena have been observed in previous studies of the SRF (e.g., De Fruyt et al., 2013). One obvious explanation for the relative performance of the Submissiveness scale is the scale’s length, being the shortest of the scales at only four items. On the other hand, previous studies of self–informant correlations have found that constructs related to submissiveness (e.g., dependency; Morgan & Clark, 2010) tend to demonstrate relatively low self–informant associations (Klonsky et al., 2002), suggesting that perceptions of interpersonal submissiveness might be somewhat contextual in nature. Future revisions of the PID-5 might consider lengthening the Submissiveness scale, and possibly broadening its construct representation (Bornstein, 2012; Gore, Presnall, Miller, Lynam, & Widiger, 2012; Morgan & Clark, 2010). Another possibility is to psychometrically link the PID-5 to related measures of personality pathology, allowing access to a broader item pool via methods such as computer adaptive testing (Simms et al., 2011).
One important general area for further research is how to integrate self- and informant-report information on the PID-5 and other measures of personality pathology. This is a broad issue in many ways, which can be approached from multiple perspectives. The self–informant correlations observed here were similar in magnitude to those that have been reported for other personality and psychopathology instruments (Achenbach et al., 2005; Connelly & Ones, 2010; Klonsky et al., 2002). As such, the magnitude of associations between self- and informant-report are substantial enough to suggest that the two sources of information reflect common variance in target behavior, but not so great as to suggest that they are both entirely overlapping in their sources of reliable variance.
Multiple approaches might be taken to understanding how to integrate self- and informant-reports on the PID-5 and other instruments. For example, it may be that self–informant report discrepancy moderates the external validity of self-reports in some way, such that informant reports can be used to assess patterns of impression management or response style (Piedmont et al., 2000). It may also be that different informants provide unique, valid information about the behavior of the target in different contexts or relationships, such that informant-reports can be used to gauge information about a target’s behavior in a particular setting (e.g., the behavior of a client with his or her spouse vs. friend). Similarly, independently of the target’s actual behavior, informant reports might provide important information about the perceptions of the informant per se, which might be useful in certain contexts (e.g., family therapy). Finally, it may be that the informant and self-reports can be integrated in some way to estimate the target’s standing on some situationally generalizable trait dimension (e.g., by using the informant report as a prior in Bayesian computerized adaptive testing).
Footnotes
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: Data collected at the University of Toronto were supported in part by Ontario Problem Gambling Research Centre Grant 2662; PID-5-SRF normative data were collected using funds from the American Psychiatric Institute for Research and Education.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
