Abstract
“Sexually violent predator” (SVP) legislation requires, in part, that an individual has a mental abnormality that causes difficulty in controlling sexual behavior. Previous research has found paraphilia not otherwise specified (NOS) as one of the most prevalent diagnoses proffered in SVP evaluations. However, the fifth edition of the Diagnostic and Statistical Manual (DSM-5) modified paraphilia NOS diagnosis in two ways. First, this diagnosis was divided into two new diagnostic categories: other specified paraphilic disorder (OSPD) and unspecified paraphilic disorder. Second, OSPD required an added specifier to indicate the individual’s source of sexual arousal. To date, no study has systematically explored how the revision to paraphilia NOS has affected diagnoses within SVP evaluations. The current study explored the frequency and diagnostic reliability of paraphilic disorders in a sample of 190 adult men evaluated for SVP civil commitment using the DSM-5. Results indicated that OSPD was the second most common paraphilic disorder, next to pedophilic disorder. However, there was poor to fair agreement (kappa = 0.21, p < .01) between independent evaluators in providing this diagnosis. Additionally, the two most common OSPD specifiers were non-consent and hebephilia, despite recent debate and rejection of these constructs from the DSM-5. While these constructs were the most prevalent, the specifiers contained quite varied terminology, suggesting vague diagnostic tendencies within these evaluations. Given that the presence of a mental abnormality is the cornerstone to the constitutionality of SVP commitment, diagnostic practices should be based in reliable and valid techniques.
Keywords
Introduction
To date, “Sexually Violent Predator” (SVP) statutes have been implemented by twenty states and the federal government. These statutes allow for the indefinite, post-sentence civil commitment of certain individuals deemed to be “high-risk sexual offenders” following the completion of their prison sentence (see also pre-trial commitment for District of Columbia §22-3801 – 3811). The purpose of SVP commitment is both to protect the community by detaining the highest risk offenders who have committed sexually based crimes, and to provide mental health treatment aimed at rehabilitating individuals should they eventually return to the community (18 U.S.C. § 4248, 2024). Although the specific criteria for SVP civil commitment vary by state, in general, for an individual to be eligible for post sentence civil commitment they must have a history of sexual offending; must be diagnosed with a mental abnormality or personality disorder which makes it difficult or impossible to control their sexual behavior; and must be deemed as having a high likelihood to sexually re-offend (Kansas v. Crane, 2002; Kansas v. Hendricks, 1997).
The SVP commitment process is initiated when an incarcerated individual is reaching the end of their prison sentence. Based on a criminal history of sexual offenses, an individual may be referred to the state or federal government (petitioner) for commitment consideration. Forensic evaluators, typically a psychologist or psychiatrist (Felthous & Ko, 2018), are hired by the petitioner to conduct an SVP evaluation. In certain scenarios, the petitioner may request the individual to undergo more than one evaluation. Based upon the results of the SVP evaluation(s), the petitioner may then file for a commitment proceeding. If the court determines that probable cause exists, a civil commitment proceeding would be instituted (Felthous & Ko, 2018). Prior to the proceeding, the defense may also hire an expert to conduct an additional SVP evaluation. While there are legal criteria requisite for SVP commitment, there are no legally-mandated guidelines for conducting SVP evaluations nor is there a widely agreed upon and specific set of best practices. The Association for the Treatment and Prevention of Sexual Abuse (ATSA), an internationally recognized organization dedicated to the prevention and intervention of sexual abuse, provided a policy statement about the civil commitment of sexually violent predators (Association for the Treatment and Prevention of Sexual Abuse, 2010). Their position emphasized that SVP evaluations should be “conducted using empirically validated risk assessment instruments, measures, and methods” (https://www.atsa.com/civil-commitment-sexually-violent-predators). Likewise, ethical guidelines that govern the practice of forensic psychologists (e.g., The Specialty Guidelines for Forensic Psychology, 2013) emphasizes the use of multiple sources of information. A practice survey and additional research suggest SVP evaluations typically include a clinical interview with the defendant, a comprehensive review of records, and results from empirically validated risk assessment measures (Jackson & Hess, 2007; Kelley et al., 2020).
However, there has been debate about the criteria for SVP civil commitment, particularly surrounding the requirement that the individual suffer from a mental abnormality or personality disorder that impairs volition and predisposes the person to commit a sexual offense. Part of the confusion likely stems from the fact that there are no legal guidelines as to what this mental abnormality or personality disorder should entail or how it should be defined or diagnosed. Additionally, states vary in the terminology, sometimes exchanging ‘mental abnormality’ for ‘mental illness,’ ‘mental disorder,’ or ‘behavioral disorder’ (DeMatteo et al., 2015). Overall, most states’ definition is modeled after Hendricks (Kansas v. Hendricks, 1997), requiring that an individual suffer from “a congenital or acquired condition affecting the individual’s emotional or volitional capacity which predisposes the person to commit sexually violent offenses to a degree constituting that such a person is a menace to the health or safety of others” (for a review of each state’s definition, see DeMatteo et al., 2015). Although state guidelines do not require that this mental abnormality be construed based upon a disorder from the Diagnostic and Statistical Manual (DSM; APA, 2013), ethical guidelines emphasize that forensic psychologists base their opinions upon “scientific foundation, and reliable and valid principles and methods…” (American Psychological Association, 2013b, p. 9). Therefore, SVP evaluators “are obligated to use established diagnostic categories, such as those in the DSM or an alternative set of psychological concepts, that have been operationalized and validated.” (Wollert, 2007, p. 168).
Further, confinement under SVP law occurs subsequent to an individual successfully completing a criminal sentence; and is, most strikingly, indefinite 1 . As such, there are serious, long-term consequences in classifying an individual as a sexually violent predator. Notwithstanding the ethical obligations of forensic psychologists, the clinical practice of SVP evaluations should be conducted in accordance with the utmost scientific integrity given its potential for a significant imposition on one’s civil liberties. Likewise, the constitutionality of SVP commitment is established by the nexus between a mental illness, or defect, and sexually violent behavior. Thus, recommendation for civil commitment must be based in scientific principles and not based on uninformed clinical, legal, or political judgment (Prentky et al., 2006). That said, there have been numerous legal challenges to the constitutionality of the legislation as well as the question of the ethics surrounding the practice of SVP evaluations (Izzi, 2017; Jeglic & Calkins, 2016). Central to this debate has been the argument that the diagnoses often utilized to support SVP commitment criteria have weak empirical support establishing their reliability and validity (Frances & First, 2011a).
Frequency of Diagnoses
Several studies have explored the frequency of diagnoses among individuals who were evaluated for SVP civil commitment (Perillo et al., 2014) or ultimately committed (Becker et al., 2003; Elwood et al., 2010; Jackson & Richards, 2007; Janus & Walbek, 2000; Jumper et al., 2011; Levenson, 2004; Lu et al., 2015; McLawsen et al., 2012). These studies found that the most commonly utilized diagnoses for SVP civil commitment were paraphilic disorders; specifically, pedophilia and paraphilia not otherwise specified (NOS). That said, research has suggested there is questionable reliability of the diagnoses provided in SVP evaluations—particularly for paraphilia NOS (Packard & Levenson, 2006; Perillo et al., 2014). For example, Perillo and colleagues (2014) examined the inter-rater reliability of diagnoses provided by independent evaluators in 375 New Jersey SVP cases. Their results suggested poor (.23) to moderate (.55) agreement across the diagnoses given by clinicians in this context. Further, evaluators’ ability to reliably diagnose paraphilia NOS was modest (kappa = .35; PPV = 0.52). Similar findings were also demonstrated by Packard and Levenson (2006) in 540 Florida SVP cases wherein evaluators’ consistency in diagnosing paraphilia NOS was also modest (kappa = .36; PPV = .65).
Diagnostic Criteria for Paraphilia NOS, OSPD, and UPD.
Updated Diagnostic Criteria
Since publication of the above-mentioned studies which examined the frequency and reliability of SVP diagnoses, a new edition of the DSM has been published. That is, previous research was based on diagnoses from the fourth edition of the DSM (DSM-IV and its respective text revision, DSM-IV-TR; American Psychological Association, 1994; 2000) which utilized the term paraphilia NOS. The fifth edition of the DSM, which was published in 2013 (DSM-5; American Psychological Association, 2013b), no longer carries a paraphilia NOS category. Rather, this classification has been modified into two categories with the intention to improve diagnostic specificity and clinical communication (First, 2014). The two new categories include ‘other specified paraphilic disorder (OSPD),’ and ‘unspecified paraphilic disorder (UPD).’ All three of these categories (e.g., NOS; OSPD; UPD) are largely similar (see Table 1 for DSM-5 description) in that they allow the clinician to provide a diagnosis for an abnormal sexual interest that causes marked impairment in functioning but does not meet any of the eight DSM-5 paraphilic diagnoses (e.g., pedophilic disorder, sexual sadism disorder, frotteuristic disorder). The main difference between OSPD and UPD (and paraphilia NOS) is the option for the clinician to choose to communicate or specify the atypical sexual interest that is causing significant distress or impairment (First, 2014). Specifically, the DSM-5 states, “The other specified paraphilic disorder category is used in situations in which the clinician chooses to communicate the specific reason that the presentation does not meet the criteria for any specific paraphilic disorder. This is done by recording ‘other specified paraphilic disorder’ followed by the specific reason” (APA, 2013, p. 705). The DSM-5 then provides a non-exclusive list of examples (e.g., necrophilia; zoophilia). UPD can therefore be utilized when the clinician chooses not to specify the atypical paraphilic interest (First, 2014). UPD also provides an opportunity for the clinician to infer that a paraphilic interest is present; however, there is not enough information to provide a more definitive diagnosis (e.g., the individual is not forthcoming in his or her self-report; APA, 2013; First, 2014).
To date, research has not systematically explored which specifiers are utilized with the OSPD diagnosis as it relates to the evaluation of sexually violent predator legislation. Interestingly, however, three of the previous studies (Elwood et al., 2010; Jackson & Richards, 2007; McLawsen et al., 2012) parsed out diagnoses of paraphilia NOS into ‘paraphilia NOS, non-consent’ and ‘paraphilia NOS, excluding non-consent.’ Likewise, a review of SVP cases between 2008 and 2011 revealed that paraphilia NOS, non-consent has been used with increased frequency (King et al., 2014). More recently, several Frye hearings in the State of New York (e.g., Matter of State of New York v. Ralph P., 2016; and Matter of State of New York v. Jason C., 2016) have suggested the terms ‘non-consent’ and ‘hebephilia’ are being used as specifiers for the OSPD diagnosis.
What are “Hebephilia” and “Non-Consent”?
Hebephilia does not have a formal definition, however, it generally refers to a sexual preference for pubescent-aged adolescents. It is categorically distinct from pedophilia, which is the sexual preference for prepubescent-aged children, and from teleiophilia—the sexual preference for adults. Stephens et al. (2017) note that hebephilia has been conflated with an interest in older adolescents (e.g., 15 – 17-year-olds), however, hebephilia specifically refers to the interest in youth who are in Tanner stages 2-3. The Tanner stages describe the primary and secondary physical features of sexual development from childhood to adulthood (e.g., size of breasts or testes, development of pubic hair; Tanner, 1990). Those in the 2nd and 3rd stages are beginning to show some secondary sexual characteristics which would indicate the initial growth of pubic hair as well as budding breasts; versus older adolescents whose sexual development more closely resembles that of an adult. Typically, these Tanner stages refer to those who are around 11 – 14 years old, however, age is not a definitive factor as sexual development varies among individuals.
Paraphilia non-consent is a construct most in line with the crime of rape. Like hebephilia, paraphilia non-consent has no standard definition. In general, it typically refers to sexual arousal to coercive, sexual contact with non-consenting individuals (e.g., Wakefield, 2011). This construct has been referred to by several different terms. Paraphilia non-consent appears to be the most used term, however, paraphilic coercive disorder and biastophilia have also been used. By some, these terms appear to be referring to the same overall construct and have even been used interchangeably in the same paper (e.g., Knight, 2010). Others have defined differences between the two. Paraphilic coercive disorder has been explained as a sexual arousal to the coercive nature of the rape (Thornton, 2010). Whereas biastophilia has been described as a sexual arousal to the coercive nature and to the victim’s terror and resistance (Money, 1999). For this study, however, paraphilia non-consent will be the only term utilized but represents more broadly the sexual arousal to coercive sexual interactions. More recently, paraphilic coercive disorder has been theorized to fall along an “agnostic continuum” of sexual aggression, with no coercive fantasies at the ‘low end,’ to fantasies and behaviors of hurting, humiliating, torturing, and killing during sex at the most extreme end. This theory challenges the idea of paraphilic coercive disorder and sexual sadism disorder being distinct constructs or diagnoses (see for example, Longpre et al., 2020).
Both constructs—hebephilia and paraphilia non-consent—have stirred much debate. In short, scholars have argued for one—whether diagnostic criteria for these constructs can reliably distinguish individuals with such pathological sexual interests; two—scrutinized the methodology conducted in these respective areas (e.g., utilization of penile plethysmograph as a measure of sexual arousal); and three—questioned whether failure to adapt to cultural expectations should be considered disordered 2 . Likewise, some have expressed concern or their perception that the study of paraphilic disorders has occurred predominately within the legal arena as opposed to broader clinical treatment settings (e.g., Frances & First, 2011b; Franklin, 2010; King et al., 2014; Moser, 2009; Tromovitch, 2009; Zander, 2008, 2009; to the contrary, see for example, Seto et al., 2016). All in all, the DSM-5 Sexual and Gender Identity Disorders Work Group, Paraphilias Subgroup (see, Zucker, 2010) and its advisors determined to reject the constructs of hebephilia and paraphilia non-consent as unique mental disorders, as well as from specifiers under the OSPD diagnosis and rejected in the appendix as an area warranting further research (First, 2014).
The Use of Hebephilia and Paraphilia Non-Consent in SVP Evaluations
Despite these controversies, a 2012 survey of United States’ case law revealed that the use of paraphilia NOS, non-consent was, with one exception, used solely in matters pertaining to sexually violent predator legislation (King et al., 2014). Further, the findings generally showed an increase in usage of this disorder from 1998 through 2011. Interestingly, in over half of these cases, an opposing expert was present, and two-thirds of those experts testified about the insufficient evidentiary support surrounding paraphilia NOS, non-consent. Nevertheless, of the 126 cases, only 16 sought to challenge the admissibility of this disorder. Ultimately, all of the courts in these cases (albeit two courts found insufficient support for the diagnosis) ruled paraphilia NOS, non-consent to be admissible (King et al., 2014).
Further, it was argued that findings supporting the reliability and validity of these constructs is weak and any evidence relative to the etiology and prevalence of these constructs are in the nascent stages. Interestingly, however, recent Frye 3 hearings have ruled that a diagnosis of hebephilia (Matter of State of New York v. Ralph P., 2016) and paraphilia non-consent (Matter of State of New York v. Jason C., 2016; Matter of State of New York v. Kareem M., 2016) was inadmissible for SVP commitment not because it was not incorporated into the DSM-5 (as the petitioner also argued in Matter of State of New York v. Jason C., 2016), but because there was not a general consensus among those in the field of the acceptability of this construct as a diagnosis.
Although the court in these specific cases ruled that hebephilia and paraphilia non-consent were not generally accepted diagnoses within the field of psychology, other courts have accepted the admissibility of such diagnoses (e.g., The People of the State of Illinois, v. Kevin Stanbridge, 2012; see also, King et al., 2014). Further, pilot data has suggested that paraphilia non-consent was the most commonly utilized specifier within SVP civil commitment evaluations—with hebephilia being the second most prevalent. Thus, there may be some disagreement in the field about whether these constructs are ‘generally accepted.’ However, current literature does not address this possibility.
Current Study
Previous research has suggested that paraphilia NOS is one of the most frequently used categories in SVP evaluations (e.g., Elwood et al., 2010; Perillo et al., 2014). However, since these studies have been published, the DSM-5 modified the paraphilia NOS categories into two categories: other specified paraphilic disorder (OSPD) and unspecified paraphilic disorder (UPD). Research that has explored the frequency of diagnoses within SVP evaluations has not been updated since the addition of these modified diagnostic categories. Thus, the first aim of this study is to explore the frequency in which OSPD and UPD are used in SVP evaluations. Given the novelty of these diagnoses, this initial aim is exploratory in nature.
Previous research has evidenced poor reliability of diagnoses utilized in SVP evaluations; this has been specifically apparent for the paraphilic diagnoses (Packard & Levenson, 2006; Perillo et al., 2014). The second aim of this study it is to understand whether the modified paraphilia NOS categories demonstrate better inter-rater reliability (IRR) compared to paraphilia NOS. Given that the reliability of paraphilia NOS in the SVP context has been poor to modest in previous studies, it is anticipated that OSPD and UPD will continue to demonstrate poor reliability given vague diagnostic criteria in the DSM-5.
With the newly created OSPD diagnostic category, the DSM-5 allows clinicians to identify the source of problematic sexual arousal for the evaluee. There is no data driven guide or consensus opinion about what specifiers could be used. Thus, the third aim of this study is to explore which specifiers are being used within SVP evaluations. Based on anecdotal evidence, court hearings, and preliminary data, it is hypothesized that paraphilia non-consent and hebephilia will be the most commonly utilized specifiers.
Method
Data for the present study were provided by the Florida Department of Children and Families, Sexually Violent Predator Program (herein, SVPP). The SVPP houses all records of individuals who were evaluated under Florida’s SVPP. Individuals who have a “sexually violent offense” are referred by the Department of Corrections, Department of Juvenile Justice, and the Department of Children and Families to the SVPP multidisciplinary team (MDT) 545 days before their release (or as soon as possible if the incarceration period is shorter). All referrals are screened by the MDT. The team reviews each referral. In their review, the team examines several factors to determine whether the individual should be referred for a face-to-face evaluation. These factors include, but are not limited to, “the defendant’s pattern and severity of sexual offenses, evidence of a paraphilic disorder, evidence of a severe personality disorder with a sexually violent focus, evidence of a psychotic disorder with sexualized content, pattern of institutional violations of a sexual nature, refusal to cooperate or early termination from treatment, limited time in the community without a sexual offense, and an imminent risk for sexually violent behavior” (S. Lewis, personal communication, May 6, 2020). These contracted evaluations are usually performed by doctoral level psychologists who are required to opine whether the individual meets the state’s definition of a sexually violent predator. Specifically, Florida State Law defines a “sexually violent predator” as someone who “has been convicted of a sexually violent offense; and suffers from a mental abnormality or personality disorder that makes the person likely to engage in acts of sexual violence if not confined in a secure facility for long-term control, care, and treatment” (Florida § 394.912). The MDT initially requests one evaluation. If that evaluation indicates that the individual does meet commitment criteria, then a request for a second independent evaluation is made. Occasionally the MDT will request a second evaluation even if the first evaluation results in the opinion that the individual does not meet criteria. This may occur when the MDT concludes the first evaluation did not sufficiently answer an important question or perhaps when new information surfaces 4 . This second evaluation is performed by another contracted evaluator. The psychologist performing the second evaluation is not formally informed a first evaluation has been conducted, nor are they privy to the first evaluation report. However, it is possible the second psychologist becomes anecdotally aware (e.g., the individual reports it) a first evaluation occurred. The MDT reviews these evaluations and makes a final recommendation. If a recommendation for civil commitment is made, then this opinion is sent to the state attorney and a petition may be filed.
Sample
Given the aims of this study, there were several inclusion criteria. First, given that one aim of this study is to explore the frequency of the paraphilic categories adopted by the DSM-5, only those evaluations conducted after this edition was published (May 2013) were included. That is, this sample only included evaluations conducted between May 2013 and June 2017 (n = 611). As mentioned, there are instances in which the individual only receives one evaluation; however, given that one of the aims of this study was to explore the IRR in mental health diagnoses between evaluators, we used only those cases in which two evaluations were conducted. Overall, 42% (n = 255) of those individuals who were referred by the MDT for an SVP evaluation during the study time period were evaluated by two evaluators. A further 65 cases (25%) were excluded from the final sample because at least one evaluator did not use the DSM-5; that is, they reported diagnoses based on the DSM-IV TR. This resulted in a final sample of 190 cases, or 380 evaluations. The two evaluations were conducted within a range of time between one another, spanning from within one day of each other upwards of 410 days. On average, there was about one month (36 days) between the two evaluations (median was 27 days).
The entire sample comprised males who were convicted of sexual offenses (n = 190). They were all above the age of 18. These individuals were identified as White (n = 100; 53%), Black (n = 80; 42%), Hispanic (n = 8; 4%), or other ethnic minorities (n = 2; 1%). Individuals were incarcerated for a variety of sexual offenses and had a history of a sexually violent offense per Florida statute (e.g., sexual battery; lewd or lascivious act with or in presence of child; kidnapping or false imprisonment of a child involving sexual battery or lewd or lascivious acts; murder while engaged in sexual battery; Florida § 394.912). Further, 84% (n = 160) of the sample had a reported history of prior sexual offenses. The majority of individuals in this sample (70%, n = 133) were ultimately referred for civil commitment by the MDT.
Of the 380 evaluations that were conducted, there were 21 distinct evaluators. All of the evaluators were licensed psychologists in the state of Florida; most (n = 14; 67%) held a Ph.D., while a third held a Psy.D. Thirteen of the evaluators were male (n = 13; 62%) and eight were female (n = 8; 38%). Within this dataset, evaluators conducted a range of evaluations (Range 3- 39, M = 18).
Procedure
The current study was part of a larger study that was approved by the Institutional Review Board of the primary investigator’s host institution, as well as by the Florida Department of Children and Families. Four trained M.A. level psychology graduate students extracted data from each SVPP evaluation based on an established coding manual. Data were collected using an established spreadsheet that matched the coding manual. For all categorical variables, drop-down options were provided in the spreadsheet for coders to select the appropriate response; this helped to increase consistency in data collection. Offender demographic data were obtained from each file. Each SVPP evaluation was coded for the evaluator name, his or her gender and educational degree. Evaluator names were later assigned a non-identifiable specifier (e.g., letters A-U) and evaluators’ names were removed from the dataset. Evaluations were coded for the date of evaluation, whether a face-to-face evaluation was conducted—as some individuals rejected the interview, what diagnoses were provided, and whether the evaluator recommended the individual for civil commitment. Evaluators may have reported several diagnoses. The first five diagnoses reported in the evaluator’s report were coded. To note, paraphilic diagnoses were never reported as the sixth or greater diagnosis. Therefore, any diagnosed paraphilic disorder in this sample of offenders was captured. If a diagnosis of OSPD was provided, the specifier linked to this diagnosis was coded. In some instances, the evaluator provided a specifier for UPD, this was also recorded. Reliability across coders was explored in 10% of the sample. There was 100% agreement between coders in collecting offender demographics and SVPP evaluation variables as identified above.
Results
Aim 1. Frequency of Paraphilic Diagnostic Categories
Prevalence of Paraphilic Diagnoses.
Note. OSPD = Other specified paraphilic disorder; UPD = Unspecified paraphilic disorder.
Aim 2. Reliability of Paraphilic Diagnoses
Diagnostic Reliability Across Evaluators.
Note. Non-Consent Comb. and Hebephilia Comb. represent the combined specifiers. The Bloom et al. (1999) standard for kappa agreement was used (“poor” = below 0.60, “fair” = 0.60 – 0.74, and “good” = 0.75 and above) to assess for poor to good reliability. PA = Proportion of agreement, overall; PA+ = Proportion of agreement diagnosis is present; PA- = Proportion of agreement diagnosis is not present.
*p < .05; **p < .01; ***p < .001.
Positive predictive values (PPV) were used to demonstrate the probability that both evaluators agreed on the presence of a given diagnosis or specifier, given that the first evaluator provided that diagnosis. There was a high probability (90%) that if the first evaluator diagnosed any paraphilic disorder, the second evaluator would likely do the same. That said, PPVs ranged from 0.00 – 0.91 for the paraphilic disorders. Of the paraphilic disorders, PPV (91%) was strongest for pedophilia. For a diagnosis of OSPD, however, there was less than chance agreement if the first evaluator rendered this diagnosis that the second evaluator would do the same (PPV = 0.48). To note, PPV could not be calculated for frotteuristic disorder or transvestic disorder as there were no cases in which both evaluators provided this diagnosis.
Negative predictive values (NPV) indicate the probability that both evaluators agree the diagnosis is not present, given that the first evaluator did not provide said diagnosis. NPV trends were consistent across all the paraphilic categories; when the first evaluator did not provide a specific paraphilic diagnosis, the second evaluator was also unlikely to diagnosis this specific disorder (NPVs ranged 0.75 – 0.99). Regarding OSPD, if evaluator one did not diagnose OSPD, there was about 75% chance the second evaluator also would not diagnose OSPD. PPV and NPV values, however, are also sensitive to base rates of a disorder—such that diagnoses with a low prevalence will have a lower PPV and a higher NPV (Riddle & Stratford, 1999).
Proportion of agreement between evaluators was also explored. Proportions of agreement are descriptive statistics that compute the percent of times the evaluators agreed overall (i.e., both evaluators agreed in diagnosing or not diagnosing a specific disorder); or agreed on the presence (positive proportion of agreement; PA+) or absence (negative proportion agreement; PA-) of a disorder. Overall, evaluators were likely to agree that a paraphilic disorder, of some type, was present (PA+ = 0.82). However, agreement on specific disorders ranged from 0 (i.e., no agreement at all) to 0.85. Consistent with the aforementioned findings, evaluators were most likely to agree on the presence of pedophilic disorder (PA+ = 0.85). Evaluators agreed on the presence of OSPD 43% of the time (n = 46). There was no agreement (PA+ = 0.00) for diagnoses of sexual sadism disorder, frotteuristic disorder, and transvestic disorder. Evaluators were more consistent in opining when a specific paraphilic disorder was not present (this proportion ranged from 0.78 – 0.99).
Paraphilia Not Otherwise Specified (NOS)
To compare the results of this study to previous studies, a paraphilia NOS variable was computed by combining OSPD and UPD diagnoses. Compared to OSPD, this computed paraphilia NOS variable demonstrated improved diagnostic consistency between evaluators. While kappa was still considered poor (kappa = 0.27, p < .001), evaluators were much more likely to agree on the presence of this diagnosis (PA+ = 60%) compared to OSPD (PA+ = 43%) or UPD (PA+ = 30%) alone.
Discrepant OSPD Diagnostic Tendencies
After computing the primary analyses, we sought to further understand the discrepancy between evaluators’ diagnostic tendencies as it related to the use of OSPD. Descriptive statistics for the cases including an OSPD diagnosis were further explored. There was a wide range in the frequency with which individual evaluators used this diagnosis. For example, two (9%) clinicians proffered this diagnosis in 50% or more of the evaluations they conducted, and an additional seven evaluators (33%) provided it more than a third of the time. Eight (38%) of the clinicians provided this diagnosis in less than a quarter of the evaluations they conducted, and one clinician never used this diagnosis in any of the evaluations they conducted.
As noted above, OSPD was diagnosed on 107 occurrences. Within those instances, 57% of the time (61 occurrences) evaluators did not agree on the presence of the OSPD. Said otherwise, there were 61 cases in which one of the evaluators diagnosed OSPD and the other evaluator did not. During these instances, about a third (31%) of the time the other evaluator diagnosed UPD (n = 19). Of the remaining 42 cases, a paraphilic disorder of some type was provided 18 times (43%). These paraphilic disorders were typically pedophilia (n = 12) or sexual sadism (n = 5). The majority of remaining cases included a personality disorder (n = 13; 31%).
Aim 3. Specifiers of OSPD
As noted, OSPD was the second most frequent paraphilic diagnosis. Out of the 380 evaluations conducted, this diagnosis was offered 107 times (28%). Per the DSM-5, clinicians are required to provide a specifier to indicate the source of sexual arousal. Of the 107 times this diagnosis was offered, clinicians provided a specifier 87% of the time (n = 93).
The specifiers offered for OSPD were far-ranging (see Appendix 1). The most frequently used label was “non-consent,” which was provided 42 (39%) times. However, there were several more terms (n = 16) that used a variation of the term ‘non-consent’ (e.g., “non-consenting persons”). Further, there were several other specifiers that represented terminology often used interchangeably in the clinical literature referencing the concept of nonconsensual sex (e.g., biastophilia; “paraphilic coercion and courtship disorder”). The label “hebephilia” was the next most commonly used specifier, although it was used far less frequently (n = 5; 5%). Three specifiers combined hebephilia with another term (e.g., “ephebophilia;” “pedo-hebephilia”) and several specifiers appeared to represent the concept of hebephilia (e.g., “sexually attracted to young pubescent females”). Two specifiers combined the distinct terms of non-consent and hebephilia, whereas several more specifiers appeared to represent this combined presentation (e.g., “Non-Consensual Sexual Activity with Adolescent”). Two specifiers referenced remission (e.g., “in a controlled environment”), and some specifiers added the remission status with a specified sexual arousal pattern (e.g., “Non-consent, in a controlled environment”). Other sexually related labels occurred seven times (e.g., “zoophilia;” “sexting”). While UPD does not require a specifier, clinicians provided one 18% of the time (n = 10). These specifiers included “non-consent” on three occasions; “nonconsenting” on two occasions; “force”; “rule out: sexual sadism, pedophilic disorder;” “rule out exhibitionism;” “with elements of exhibitionism;” and “in a controlled environment.”
Exploration of Paraphilia Non-Consent and Hebephilia Specifiers
To better understand the diagnostic practices of the most frequent specifiers, four variables were created as related to paraphilia non-consent and hebephilia. A variable for ‘non-consent’ was created based on cases in which the evaluator provided a label utilizing the word ‘non-consent’ (n = 62). A ‘non-consent combined’ variable was created to include all the ‘non-consent’ specifiers, as well as labels presumed to depict ‘non-consent’ (e.g., biastophilia; paraphilic rape; n = 69). Therefore, there were seven instances in which a label without the word ‘non-consent’ was added to this combined category (see Appendix 1). Similarly, a ‘hebephilia’ variable included only those cases in which the evaluator used the distinct language of ‘hebephilia’ (n = 8). The ‘hebephilia combined’ included nine more labels presumed to address hebephilic preferences (e.g., ‘sexual activity with an adolescent;’ see Appendix 1).
To note, there were nine instances a label was used which did not appear to align with ‘non-consent’ or ‘hebephilia’ (e.g., ‘bestiality;’ ‘sexting’); these labels were not included in subsequent analyses but can be reviewed in Appendix 1.
Reliability of Paraphilia Non-Consent and Hebephilia
Consistency in providing a ‘non-consent’ label was analyzed using both of the ‘non-consent’ variables as identified above. Overall, evaluators showed poor agreement in using a non-consent specifier, no matter how non-consent was defined (see Table 3). Results demonstrated poor agreement in the use of the ‘non-consent’ (only) specifier (kappa = 0.17, p < .05); there was 30% agreement between evaluators in providing this label. The ‘non-consent combined’ variable—which provided more inclusivity of labels—only slightly improved reliability between evaluators (kappa = 0.22, p < .01); proportion of agreement increased to 35% (vs. 30%)
Evaluators also demonstrated poor agreement in the use of the ‘hebephilia’ specifiers (e.g., ‘hebephilia only’ and ‘hebephilia combined’). Poor agreement in the use of the ‘hebephilia’ (only) label was demonstrated (kappa = 0.27, p < .01) as well as the ‘hebephilia combined’ variable (kappa = 0.24, p < .001). Proportion of agreement between evaluators in providing this label decreased from ‘hebephilia only’ to the combined variable (29% vs. 27%).
Utilization of Hebephilia and Paraphilia Non-consent Specifiers
To further explore the trends in utilization of these specifiers, several analyses were conducted. For inclusivity and slight increase in statistical power, the ‘combined’ variable for each specifier was solely utilized.
As noted above, ‘non-consent’ terminology was used 69 times out of the 107 times OSPD was diagnosed. It appears, on average, this specifier was used in half of the OSPD diagnoses (52%); however, this statistic is somewhat misleading. Rather, there were 13 evaluators who provided the non-consent specifier in 50% or more of their evaluations, whereas five evaluators never used this descriptor.
Hebephilia terminology was used 17 times out of the 107 times OSPD was diagnosed (16%). Again, it appears as if ‘hebephilia’ was used, on average, in 16% of the OSPD diagnoses. However, three evaluators used this term half the time they provided an OSPD diagnosis, nine evaluators never used this specifier, and eight evaluators used ‘hebephilia’ as a specifier 8 – 30% of the time they diagnosed OSPD.
Discussion
This study sought to update the literature about the prevalence of diagnoses utilized in SVP evaluations. Specifically, previous research (e.g., Becker et al., 2003; Elwood et al., 2010; Jackson & Richards, 2007; Janus & Walbek, 2000; Jumper et al., 2011; Levenson, 2004; Lu et al., 2015; McLawsen et al., 2012) found paraphilia NOS to be one of the most frequently diagnosed paraphilic disorders within the SVP context. However, prior findings were based on an earlier version of the DSM and not the most current DSM-5, which has since modified the paraphilia NOS category by dividing it into two separate diagnostic categories—OSPD and UPD. Additionally, this revision now requires clinicians to specify an OSPD diagnosis by indicating the source of sexual arousal for the individual being evaluated. Until now, however, no study had systematically explored how the revision to paraphilia NOS has affected diagnostic tendencies within SVP evaluations.
Frequency of Paraphilic Diagnoses
In the current study, pedophilic disorder was the most common (31%) and OSPD was the second most diagnosed paraphilic disorder (29%). UPD was the third most common diagnosis (18%), which was diagnosed in higher frequency than the rest of the paraphilic disorders (e.g., exhibitionistic disorder, sexual sadism disorder). On the one hand, these findings may suggest that OSPD from the DSM-5 has ‘replaced’ the paraphilia NOS diagnosis from the DSM-IV. On the other hand, UPD was offered as the third most frequent diagnosis in this study. This finding, along with the findings reported next, may actually suggest that the revision to paraphilia NOS has led to further diagnostic confusion.
Diagnostic Reliability
Prior research demonstrated poor diagnostic reliability of paraphilia NOS (see, Packard & Levenson, 2006; Perillo et al., 2014) and results from the current study suggested that diagnostic reliability for neither OSPD nor UPD improved above the former – arguably more ambiguous – paraphilia NOS diagnosis. Regarding OSPD, evaluators agreed on the presence of this disorder less than 50% of the time it was diagnosed. Evaluators agreed on the presence of UPD about 30% of the time it was diagnosed. Thus, the attempt to further clarify paraphilia NOS did not improve diagnostic reliability. Interestingly, when a pseudo paraphilia NOS diagnosis was computed for the current study (by combining OSPD and UPD), diagnostic reliability improved in comparison to the reliability for either OSPD or UPD alone; however, the improvement was only modest.
The poor diagnostic reliability of paraphilia NOS has been attributed to its lack of objective or quantifiable diagnostic criteria (Frances & First, 2011a); an argument partly supported by findings of the present study. Specifically, pedophilic disorder—a paraphilic disorder with more explicit diagnostic criteria than OSPD—demonstrated greater agreement between evaluators in the current study (as well as other studies, e.g., Packard & Levenson, 2006; Perillo et al., 2014). That is, evaluators tended to agree on the presence or absence of pedophilic disorder in the individual they were evaluating (PA = 91%). Further, pedophilic disorder was used by all the evaluators (e.g., evaluators offered this diagnosis in 13% to 47% of the evaluations they conducted). To the contrary, the current study demonstrated a wide range in the frequency with which evaluators used the OSPD diagnosis (0% - 67%). Taken together, one may conclude that diagnoses with more explicit diagnostic criteria not only improves inter-rater reliability, but also increases the likelihood an evaluator is willing to proffer such diagnosis. That is, explicitly defined paraphilic disorders with a stronger basis in empirical support may increase their usage, while there may be some hesitancy amongst evaluators on proffering diagnoses with vague diagnostic criteria and debatable level of empirical support.
However, the argument that explicit diagnostic criteria improve diagnostic reliability is not fully supported in the current study. Although the other seven explicitly defined paraphilic disorders (e.g., exhibitionistic disorder, sexual sadism disorder) demonstrated better diagnostic agreement, the consistency in diagnosing these disorders was still statistically considered poor (e.g., kappas ranged from 0.43 – 0.58). Moreover, despite being one of the explicitly defined paraphilic disorders, sexual sadism, frotteuristic, and transvestic disorders were only proffered by one evaluator but never diagnosed by both evaluators in a single case. 5 As such, it appears that pedophilic disorder demonstrates the most consistent and reliable use between evaluators. This may be due to its stronger empirical base compared to the other paraphilic disorders; and/or, its perceived distinct association with criteria for SVP commitment (in comparison to non-contact behaviors; i.e., exhibitionistic disorder).
OSPD Specifiers
Another prominent change with the advent of the DSM-5 revision is that the use of the OSPD diagnosis—but not UPD—requires clinicians to give a specifier or label indicating the source of atypical sexual arousal. Results from this study demonstrated that, for the most part evaluators did provide some type of label when they diagnosed OSPD. Specifically, a specifier was provided 87% of the time OSPD was diagnosed. This means, however, that 13% of OSPD diagnoses did not include any specifier. Additionally, the UPD diagnosis was accompanied with a specifier 18% (ten instances) of the time this disorder was proffered which is an incorrect utilization of UPD. In these instances, the clinician should have diagnosed OSPD and provided the specifier. As a reminder, the DSM-5 specifically states, “The unspecified paraphilic disorder category is used in situations in which the clinician chooses
With respect to the DSM-5’s directive for OSPD to include a specifier, there appears to be vague guidance offered (“Examples of presentations that can be specified using the ‘other specified’ designation include, but are not limited to [formatting added for emphasis] …” APA, 2013; p. 705; see Table 1). The results from the current study demonstrated that there does not appear to be a standard, methodological approach for how SVP evaluators determine an appropriate specifier. As delineated in Appendix 1, the specifiers encompassed varied terminology, some seemingly created by the evaluator rather than guided by a scientific manual or approach. For example, one OSPD specifier remarked upon the “complex” nature of the individual’s behaviors with the specifier, “Complex: nonconsent, force, violence, compulsive use of pornography, and telephone scatologia.” Another specifier did not address the age of the victims, rather noted, “Nonconsensual sexual activity with age-inappropriate individuals.” Upon review of the specifiers one may be particularly concerned that several of the labels appear custom to the facts of the specific case rather than resting on any empirically derived diagnosis.
From a quantitative perspective, analyses revealed that the labels ‘non-consent’ and ‘hebephilia’ were the most frequently used specifiers. The distinct term ‘non-consent’ was the most frequently used specifier as it was provided 39% of the time OSPD was diagnosed. This is in line with previous research that demonstrated that 57-67% of civilly committed individuals diagnosed with paraphilia NOS received a ‘non-consent’ label (Elwood et al., 2010; Jackson & Richards, 2007; McLawsen et al., 2012). Although the distinct term ‘hebephilia’ was used with far less frequency (5%), it was the next most common label. Further, there were several additional specifiers that could be conceived to mean ‘non-consent’ or ‘hebephilia.’ For example, the term ‘biastophilia,’ which is often interchanged for paraphilia non-consent, was used with some frequency in the current sample. The construct of hebephilia appeared to be communicated with labels such as ‘sexually attracted to teenagers.’ Additionally, there were instances in which the labels clinicians used appeared to combine both constructs (e.g., ‘non-consensual sexual activity with an adolescent’). Further, the diagnostic reliability of the most frequent specifiers (non-consent and hebephilia) was poor and use of such specifiers was inconsistent. With respect to labels encompassing non-consent terminology, both evaluators used this term in only 35% of the instances in which this term was provided. Interestingly, one evaluator used this label every time they diagnosed OSPD and over half of the evaluators used this specifier more than 50% of the time they diagnosed OSPD; still, a quarter of the evaluators never used this or similar (e.g., biastophilia) terms when diagnosing OSPD. With regard to a construct of hebephilia, three evaluators used this (or similar) label half the time they provided an OSPD diagnosis, yet nine of the evaluators never once used hebephilia-associated terminology. Moreover, evaluators agreed in providing this label less than a third of the time (27%).
In contrast to this study which explored diagnoses with legal implications, Seto and colleagues (2016) explored the reliability and validity of modified DSM-5 criteria for pedophilia and hebephilia in a clinical sample—with no legal implications. Also different in Seto et al.’s study was that evaluators were provided a summary of diagnostic criteria for these proposed modifications—in contrast to the current study utilizing the DSM-5 OSPD/UPD open-ended criteria. The results of their study demonstrated higher rates of inter-rater reliability for the proposed hebephilia criteria (kappas ranging from 0.43 – 0.56); albeit higher, these rates of consistency between evaluators remain below adequacy.
Of relevance, poor diagnostic reliability has been one factor associated with the rejection of both non-consent [paraphilic coercive disorder] and hebephilia [including, pedohebephilia] from the DSM-5. As noted above, both of these constructs were considered but ultimately rejected from the revised paraphilic disorder section, primarily due to the lack of established reliability and validity of these disorders (see Beech et al., 2016 for a review).
Implications for Practice
One of the main findings from the current study is that the DSM-5’s revision to paraphilia NOS has not improved diagnostic reliability, at least not within this sample of SVP evaluations. In fact, the results of this study suggested that the diagnostic reliability of OSPD is worse than the reliability of the already criticized paraphilia NOS diagnosis. Second, and perhaps of marked prominence, is the concern that these disorders (OSPD and UPD) and clinical specifiers may suggest an idiosyncratic application of diagnostic entities, rather than a nomothetic approach of diagnosing paraphilic disorders based on established scientific principles. Moreover, the most commonly used specifiers were two constructs that were, after much debate, rejected from the DSM-5 due to a lack of empirical support. This is a significant consideration as the lack of empiricism of OSPD and its specifiers could lead to challenging the admissibility (e.g., Daubert criteria) of this diagnosis in SVP hearings. Recently, Holoyda (2020) exemplified challenges paraphilia non-consent would likely face in consideration of an admissibility hearing. His depiction (see Table 1 of Holoyda, 2020) challenging the Daubert criteria could also lend to the construct of hebephilia given its rejection from the DSM-5, poor interrater reliability, limited research support, and lack of established diagnostic criteria.
On the other hand, it is important to address that, from a broader perspective, the results from this study did suggest that evaluators appeared to demonstrate greater consistency than the results may suggest. Despite less than chance agreement on an OSPD diagnosis, evaluators overall seem to agree that there appears to be perceived abnormal sexual interests or behaviors. For example, of the 61 times that evaluators did not agree on the presence of an OSPD diagnosis, the opposing evaluator provided a paraphilic diagnosis of some type more than half the time (61%; n = 37). Typically, this was a diagnosis of UPD (51%, n = 19), but sometimes was a diagnosis of pedophilic disorder (32%, n = 12) or sexual sadism disorder (14%, n = 5). It is also important to note that diagnostic reliability improved when OSPD and UPD were combined into one variable. As such, these findings may imply that evaluators are often still ‘on the same page,' in terms of diagnostic ideology. Nevertheless, recognizing that something is “off” about one’s sexual interests, preferences, or behaviors is a deficient standard for psychological decision making within SVP evaluations.
The use of unreliable diagnostic decision making can have significant implications across the board. Relying upon diagnoses with poor empirical support can perpetuate the use of bad science in the courtroom. While the impetus is on the psychologist to practice in a scientifically-valid manner, it is ultimately up to the decision maker(s) to evaluate the science proffered in the courtroom and determine the weight to give the expert’s testimony (Janus & Prentky, 2003). That said, judges are not necessarily trained in evaluating psychological science and therefore are not always equipped to recognize empirically supported decisions or diagnoses; and the pressure inherent in the SVP process can thereby increase the chance of poor scientific practices being introduced into the courtroom (Prentky et al., 2006). Moreover, this practice can hamper the credibility of psychology within the legal system—a status the field has worked hard to achieve. Similarly at stake is our communities’ safety and financial expenditure. The decisions proffered in SVP cases can lead to profound consequences for the individual by violating his or her civil liberties with an unconstitutional commitment. On the other hand, psychological testimony about empirically-validated mental disorders can have a positive impact within the SVP courtroom by providing further clarity between how science can inform the law; as well as balancing consequential decisions between public safety and impeding on constitutional rights.
Limitations
Although this study is the first of its kind to examine the reliability of DSM-5 paraphilic disorders and specifiers, it is not without limitations. For one, this study only included 21 distinct evaluators and it was solely evaluations conducted in Florida. As such, the findings from this study may not be reflective of the practice of SVP evaluations across the other twenty jurisdictions which apply this legislation and thus highlights the need for additional field studies. Additionally, given that an aim of this study was to explore inter-rater reliability, this sample was limited to those individuals who received two SVP evaluations. This sample may possess distinct characteristics that can limit generalizing the current findings to those individuals who receive only one evaluation. For example, the practice of the Florida SVPP requires defendants to receive an additional evaluation if the first evaluation rendered a civil commitment recommendation. As such, the defendants in the current sample may be considered of a higher risk, or greater diagnostic complexity, in comparison to those who only received one evaluation. This may pose limitations on the generalizability of diagnostic reliability. On the one hand, the diagnostic reliability of the current sample may be lower than the theoretical diagnostic reliability of individuals referred for SVP evaluation. That is, if all referred individuals were to receive two evaluations, and there was high concordance of evaluators not providing a paraphilic disorder, it is possible diagnostic reliability would be higher. On the other hand, it is possible that the diagnostic reliability in the current study may be artificially inflated due to the confounding factor of the second evaluator. That is, a second evaluator is not typically appointed unless the first evaluator recommended civil commitment. While in theory, the second evaluator is not aware of the first evaluation, in practice it is quite possible that an attorney or the offender relays this information. As such, the second evaluator may be subjected to the effects of implicit and/or confirmation bias in which they are motivated to find a diagnosis favoring commitment—which ultimately would require a relevant diagnosis. Additionally, given that the average time between the two evaluations was around one month, there may have been time for the evaluee to change their presentation due to “practicing,” coaching, or responsiveness to treatment.
To this regard, interpretation of diagnostic prevalence rates should be considered within the context of these evaluations. As discussed, the sample of evaluators in the current study were those contracted by Florida DCF. As such, this sample did not include any privately retained SVP evaluators who, due to adversarial allegiance (see Murrie & Boccaccini, 2015), may be more – or less – likely to diagnose a paraphilic disorder. Interestingly, this current study found an association, albeit not of statistical significance, suggesting that those state-contracted evaluators who conducted more evaluations in the study sample were more likely to provide a ‘non-consent’ specifier. Therefore, further research might examine whether there are patterns in how defense or prosecution retained evaluators apply paraphilic diagnoses in the SVP context.
Finally, the current study only explored diagnostic reliability of disorders frequently used within SVP evaluations. Additional research is required to understand the validity of such diagnostic categories and associated clinical specifiers. Specifically, research is needed in understanding the etiology, prevalence, and course of these disorders as well as their nexus with sexually violent behavior, and their potential response to treatment. Importantly, as First (2014) alluded to, at the present moment residual diagnostic categories such as OSPD can have utility in clinical settings to communicate to treatment providers that there is a psychiatric concern, and the specifier can help to identify a potential target for treatment. However, within the forensic setting OSPD has the potential to be significantly misused due to its lack of scientific foundation (First, 2014). In fact, many scholars would argue this lack of accompanying research places diagnoses such as OSPD “outside of what is generally accepted by the field,” (First, 2014, p. 199) and therefore could be contested on its admissibility in the court of law (DeClue, 2006; Frances & First, 2011a; Prentky et al., 2006; Tucker & Brakel, 2012; Wakefield, 2011).
Conclusion
Overall, the findings of this current study should be considered in the context of forensic evaluations and clinical psychology. That is, forensic evaluations are inherently difficult (Guarnera et al., 2017) and SVP evaluations are not ‘alone’ in demonstrating low levels of consistency between evaluators (see for example, Guarnera & Murrie, 2017; see also, Kahn et al., 2022). As such, the inconsistency between evaluators posed in this study is not necessarily at fault of the evaluators; forensic mental health assessment—overall—is not perfect and unlikely ever will be (see, Mossman, 2013). Likewise, these findings are not meant to suggest that paraphilia non-consent or hebephilia will never be established as paraphilic disorders. However, at this present time, both of these constructs lack scientific support to establish valid diagnostic methodology and criteria. While it is certainly true that there are high-risk individuals who are likely to sexually recidivate upon their release from prison, providing makeshift diagnoses to satisfy civil commitment criteria significantly questions the ethical practice of psychological decision making.
Footnotes
Author’s Note
The authors takes responsibility for the integrity of the data, the accuracy of the data analyses, and have made every effort to avoid inflating statistically significant results.
Acknowledgement
The primary author would like to acknowledge Dr. Sandi Lewis, Psychological Services Director of the Florida Department of Children and Families Sexually Violent Predator Program for her significant support in providing access to this data and making this research possible. She would also like to thank Alexandria Imbriale, Sean McKinley, and Therese Todd, for their considerable assistance in data collection.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This research study was funded by a Pre-doctoral Research Award granted by the Association for the Treatment and Prevention of Sexual Abuse.
