Abstract
This prospective study examined the predictive validity of the Sex Offender Treatment Intervention and Progress Scale (SOTIPS; McGrath et al., 2012), a sexual recidivism risk/need tool designed to identify dynamic (changeable) risk factors relevant to supervision and treatment. The SOTIPS risk tool was scored by probation officers at two sites (n = 565) for three time points: near the start of community supervision, at 6 months, and then at 12 months. Given that conventions for analyzing dynamic prediction studies have yet to be established, one of the goals of the current paper was to demonstrate promising statistical approaches for the analysis of longitudinal studies in corrections. In most analyses, static SOTIPS scores predicted all types of recidivism (sexual, violent, and general [any]). Dynamic SOTIPS scores, however, only improved the prediction of general recidivism, and only when the analyses with the greatest statistical power were used (Cox regression with time dependent covariates).
Effective interventions for reducing criminal recidivism require knowing where to intervene. When the targets for intervention are characteristics of individuals, such characteristics have been referred to as criminogenic needs (Andrews et al., 1990), psychologically meaningful risk factors (Mann et al., 2010), and dynamic risk factors (Douglas & Skeem, 2005). Although the working definitions of these constructs vary, the intention of all these constructs is to orient supervision and rehabilitation efforts towards factors that are potentially changeable through deliberate intervention.
The major risk factors for general recidivism have been well summarized by Andrews and Bonta (2010) as the Central Eight: history of antisocial behavior, antisocial personality pattern, antisocial cognition, antisocial associates, poor family/marital relationships, poor adjustment in school/work, aimless use of leisure time, and substance abuse. In addition, crime is more common among young males than any other demographic group. Although there is a certain amount of specialization, the same major risk factors are associated with both violent and nonviolent criminal behavior. They are also associated with sexual crime and sexual recidivism (Brouillette-Alarie et al., 2016; Hanson & Morton-Bourgon, 2005)
There are, however, certain sex crime specific risk factors that are positively related to sexual recidivism but unrelated, or negatively related, to nonsexual recidivism (Brouillette-Alarie et al., 2016; Hanson & Morton-Bourgon, 2005). These sex crime specific risk factors include deviant sexual interests and sexualized coping (Mann et al., 2010). Consequently, risk assessments for individuals with a history of sexual crime need to consider sex crime specific factors as well as the Central Eight. Sexual crime is crime, so the common risk factors for rule violation also predict the onset of sexual crime, as well as sexual recidivism.
More than 60 years ago, Meehl (1954) cogently argued for the benefits of statistical prediction in applied psychology. Statistical prediction was usually more accurate than unstructured professional judgment, and it limited the effects of certain types of personal and professional biases. His call has been largely ignored, except in the fields of correctional and forensic psychology, where empirically derived prediction tools are now ubiquitous (Bourgon et al., 2018; Kelley et al., 2020; Neal & Grisso, 2014). In corrections, the first commonly used prediction tools were developed for general recidivism and included primarily static, historical risk factors (Bonta, 1996; Bourgon et al., 2018). Although static factors, such as age and prior criminal history, are efficient predictors of criminal recidivism, they provide little information concerning what needs to be done to mitigate this risk. Consequently, beginning in the 1990s, there was increasing demand for risk/need assessment tools to identify targets for supervision and rehabilitation efforts. These measures included the Level of Service/Case Management Inventory (LS/CMI; Andrews et al., 2004) for general recidivism, and, for sexual recidivism, the STABLE-2007/ACUTE-2007 (Hanson et al., 2007), the Violence Risk Scale—Sexual Offender Version (VRS-SO; Olver et al., 2007), and the Sex Offender Treatment Intervention and Progress Scale (SOTIPS; McGrath et al., 2012). The SOTIPS is the focus of the current study.
Following Kraemer et al. (1997), there are three empirical criteria that need to be satisfied for a characteristic to be considered an empirically validated dynamic risk factor. First, differences between individuals in the characteristic must be linked to differences in the likelihood of the outcome in the future. Second, the characteristic must show intra-individual change. Third, intra-individual change must be associated with corresponding changes in the likelihood of the outcome.
The early research found consistent support that individual differences in potentially dynamic factors (e.g., attitudes, peer associations) predicted recidivism (e.g., Gendreau et al., 1996; Hanson & Morton-Bourgon, 2005). The early studies, however, provided only weak evidence that reassessment improved prediction (see review by Douglas & Skeem, 2005). The lack of incremental effect of reassessment can at least be partially attributed to how change was studied. Most of the early research included assessments at only two time points (usually pretreatment and posttreatment), which make it difficult to differentiate true change from measurement error (Singer & Willett, 2003, §1.3.1). As well, to the extent that interventions are effective, post-treatment scores will yield smaller correlations with the outcome because of restriction of range. Perhaps the most important limitation of previous research was most previous research suffered from low statistical power. Large sample sizes are required to detect incremental effects of highly correlated measures.
In a recent meta-analytic review, van den Berg et al. (2018) examined the predictive validity of sexual recidivism risk tools that included dynamic risk factors. Their review considered 52 studies of 14 different dynamic risk tools. They concluded that the dynamic risk tools had moderate to strong ability to discriminate recidivists from nonrecidivists, and contributed incrementally to prediction beyond static risk tools. Importantly, change scores contributed incrementally after controlling for static and initial dynamic risk scores. Although these results are encouraging, they need to be interpreted cautiously. The incremental validity analyses, in particular, were based on less than 10 studies using different dynamic risk measures. Consequently, the results do not speak to the predictive properties of any specific measure.
The purpose of the current study was to examine the predictive validity of one of the sexual recidivism risk tools included in van den Berg et al.’s (2018) review: the Sex Offender Treatment Intervention and Progress Scale (SOTIPS; McGrath et al., 2012). The SOTIPS was designed to assess dynamic risk factors for recidivism among adult males serving a sentence for a sexual offense. It has 16 items grouped into three broad categories: (a) sexual deviance (e.g., attitudes tolerant of sexual offending, sexual interest in illegal activities), (b) criminality (e.g., impulsivity, lack of cooperation with supervision, antisocial attitudes), and (c) social stability and supports (e.g., residence stability, emotion management, social influences). In the development study of 759 individuals, repeated SOTIPS assessments significantly predicted sexual, violent, and any recidivism, and was incremental to the Static-99R (Helmus, Thornton, et al., 2012), an established measure of static, historical risk factors. The SOTIPS scores for most individuals improved over the course of a 2-year period of supervision, and, as expected, individuals who eventually reoffended improved less than those who remained offense free (Lasher & McGrath, 2016).
Although these results are encouraging, there are several features of the development study that limit confidence in Lasher and McGrath’s conclusions. First, the 16 items of the SOTIPS were selected from an earlier 22 item version of the same scale (McGrath & Cumming, 2003) on the basis of having statistically significant relationships with sexual recidivism in the same dataset (validation sample was the same as the development sample). All datasets have unique features, which makes it hard to tell how well the SOTIPS would function in other settings. As well, the predictive validity of the SOTIPS has not been tested outside of Vermont, which is a predominantly rural state, almost exclusively populated by White individuals (96.4% of the development sample). Another limitation is that the analytic approach used by McGrath et al. (2012; generalized estimating equations [GEE]; Liang & Zeger, 1986) did not explicitly test whether reassessments predicted recidivism better than the initial, baseline SOTIPS assessment (see comments below).
Statistical Analysis of Dynamic Prediction Data
Given that conventions for analyzing dynamic prediction studies have yet to be established, another goal of the current paper was to demonstrate promising statistical approaches for the analysis of longitudinal studies in corrections. Dynamic prediction studies with a terminal, dichotomous outcome (e.g., death, recidivism) present special challenges for statistical analysis. This was the case of the current data, which shared many features of real world dynamic prediction studies. First, SOTIPS scores and the outcome (recidivism) were expected to be measured with error, and the association between SOTIPS scores and recidivism was expected to be stochastic (imperfect probabilistic prediction). Second, although SOTIPS scores have some intrinsic meaning, they obtain credentials as risk factors because of their relationships to recidivism. Third, the potentially dynamic predictor variable (in our case, SOTIPS total scores) was assessed repeatedly, and recidivism could happen any time after the first assessment. Fourth, data were incomplete. Different individuals had different numbers of assessments, and at ragged intervals (i.e., not everybody was assessed at the same time). In prospective prediction studies, recidivism must be after the assessment; SOTIPS scores reported after recidivism events were not used (except when predicting different types of recidivism). Consequently, there were systematically fewer assessments for the recidivists than for the nonrecidivists. As well, the assessments for the recidivists, compared to those of the nonrecidivists, were earlier in the follow-up period.
Although these features present challenges for many forms of statistical analysis, they are not intractable. The field of medical epidemiology, in particular, has advanced various approaches for answering practical questions of real world, longitudinal data with complex structure (Aalen et al., 2008; Singer & Willet, 2003; Twisk, 2013). The primary approach used in the current study was Cox regression survival analysis with time dependent covariates (Altman & de Stavola, 1994; Singer & Willet, 2003). Given that the field of recidivism risk prediction has yet to develop conventions for the analysis of dynamic prediction studies, the following comments are intended to advance discussion of the strengths and limitations of the available options.
Pearson correlation coefficients are generally not recommended for dynamic prediction studies because of their sensitivity to base rates (Babchishin & Helmus, 2016). Given that the total number of recidivists available for analysis will decrease after each new assessment, so too will the magnitude of the correlations. Although there are ways of correcting correlation coefficients for declining base rates, a simpler option would be to use statistics that are less influenced by base rates, such as logistic regression (Hanson, 2008; Hosmer et al., 2013), or the area under the curve (AUC) from receiver operating characteristic (ROC) curves (Swets et al., 2000). AUCs are preferable to correlation coefficients, but they are still influenced by restriction of range in the predictor variable (Howard, 2017; Humphreys & Swets, 1991). In most dynamic prediction studies, the average risk score for the group declines because a) individuals with high scores reoffend before those with lower scores (if the risk tool works as intended), and b) individuals often improve (with or without intervention).
A more serious problem with the AUC is that it is difficult to test for changes in AUC values across assessments. Examining the overlap in confidence intervals is a very low powered test (Cumming & Finch, 2005), as are statistical tests that treat the observations as uncorrelated (e.g., Hanley & McNeil, 1982). Test of correlated AUCs, such as the Delong method (Delong et al., 1988), have high statistical power; however, these require the same individuals to be assessed at each time period, which is not the case in dynamic prediction studies.
The primary statistical approach used by McGrath and colleagues in their original SOTIPS study was a version of generalized estimating equations (GEEs; Liang & Zeger, 1986). GEE is a form of regression analysis in which the relationship between the variables at different times are considered simultaneously. The GEE regression coefficients indicate the longitudinal relationship between the predictors and the outcome based on all available data. Although GEEs are appropriate for modeling directional effects of one variable on another over time, GEE coefficients are difficult to interpret. Given that they combine between-individual and intra-individual effects (Twisk, 2013, §4.5.3), a large GEE coefficient could indicate that (a) there were only static differences in risk levels, (b) dynamic reassessment improved the prediction of recidivism, or (c) a combination of the two effects.
For evaluating change during institutional treatment, it is common (and often appropriate, e.g., Hogan & Olver, 2019) to enter both pretreatment and posttreatment scores as predictors in some form of regression analysis (e.g., logistic, Cox). Testing the incremental effect of the posttreatment score over the pretreatment score is equivalent to simultaneously entering the pretreatment score with the pre-post difference score (Laird & Weems, 2011 demonstrate the mathematical equivalence). Although incremental effects of posttreatment scores are evidence that change in observed scores improves prediction, this may or may not be attributable to true change in the individuals. Each observation contains error, and it is extremely difficult with only two assessments to separate reduced error from true change (Singer & Willet, 2003, §1.3.1). As well, large samples sizes are required to detect incremental effects among highly correlated variables. For dynamic prediction studies in the community, there is the further complication that subsequent scores are not available for individuals who have already reoffended.
Survival analysis is frequently recommended for analyzing data from studies with incomplete follow-up (e.g., Aalen et al., 2008; Singer & Willet, 2003; Steyerberg, 2019). Follow-up information is incomplete when start and end times vary, and not everybody is followed to an endpoint where the outcome is no longer possible (e.g., death terminates recidivism risk). Survival analysis is based on hazard rates, defined as the likelihood of the outcome in the next interval given that the individual has not yet experienced the outcome. Cox regression survival analysis calculates hazard ratios, which are summaries of the ratios of the hazard rates for individuals with different attributes (covariates). Specifically, Cox regression compares the characteristics of each individual at the time they reoffend to the characteristics of all the other individuals still at risk at that time (i.e., the risk set). Cox regression allows for time-dependent covariates, such that the characteristics of individuals can change over time. It also allows for individuals to exit and re-enter the risk set, such as when individuals are temporarily incarcerated for technical violations (as was the case in the current study).
Because it is impossible to measure characteristics at the time of recidivism, researchers must decide how to attribute characteristics to individuals for all times that they are at risk. The most common methods for attributing characteristics to cases are to either project forward a) the first assessment or b) the most recent assessment. With large numbers of repeated assessments, other methods of attribution can be used, such as rolling averages (see Lloyd et al., 2020), or extreme scores (e.g., worst, best; Babchishin & Hanson, 2020). Each method of attribution implies a theoretical model about the individuals in the study (e.g., they change, they don’t change, they change in such-and-such a fashion). One complication is that comparing models requires fit indices; traditional null hypothesis approaches are invalid (i.e., there is no null hypothesis to test). The most commonly used fit indices are the Akaike Information Criteria (AIC, Burnham & Anderson, 2004) and the Bayesian Information Criteria (BIC, Raftery, 1995). Both fit indices start with the difference in the observed and predicted values and then add a penalty proportional to the number of predictor variables. The model with the best fit to the data is the most plausible.
The number of recidivists is usually the limiting factor for statistical power, regardless of the prediction statistic used. When there are 1,000 individuals in the study but only 10 recidivists, do not expect stable results. In general, 100 recidivists and 100 nonrecidivists are required for stable logistic regression estimates (Vergouwe et al., 2005). Minimally, there should 10 events per predictor variable when using either logistic regression (Hosmer et al., 2013) or Cox regression (Moons et al., 2009).
Both time dependent Cox regression and pre-post change regressions directly test whether including change improves prediction (the third of Kraemer et al., 1997 criteria for a dynamic risk factor); however, each frames the question differently. In time dependent Cox regression, the question is whether models that allow for individual scores to change fit the data better than models that do not. In time dependent analyses, individuals who are only assessed once are considered equivalent to individuals who have identical scores on repeated assessments. In contrast, individuals who are only assessed once do not contribute to pre-post change studies. Time dependent Cox regression and pre-post change analyses provide similar results when there are only two assessments, and the proportion who reoffended between assessments is small. Pre-post change analyses, however, are poorly suited for studies in which some individuals have more than two assessments, or when the time between assessments is not the same for all individuals. In contrast, large numbers of repeated assessments at unequal intervals presents no conceptual or analytic challenges for time dependent Cox regression. Consequently, it was the approach privileged in the current study. For comparison purposes, however, we also present Cox regression analyses where the SOTIPS scores at each wave are treated as static variables (not time dependent). These static Cox regression analyses are conceptually equivalent to using changes scores to predict a dichotomous outcome.
The Current Study
The current study examined the predictive validity of SOTIPS as a dynamic risk tool in two new settings. The individuals in this prospective study were supervised in the community for a sexual offense in either Maricopa County (metropolitan area of Phoenix, Arizona) or New York City. We expected (a) the first (static) SOTIPS to predict recidivism, (b) the first SOTIPS to predict recidivism incrementally to another measure of static risk factors (Static-99R), and (c) for the dynamic (changing) SOTIPS scores to have a stronger relationship to recidivism than the first (static) SOTIPS scores. The data were collected as part of a National Institute of Justice study examining the influence of SOTIPS on community supervision practices and public safety (Miner et al., 2018; Newstrom et al., 2018).
Methods
Participants
All participants were adult (18+) males convicted of a contact or non-contact sex offense with an identifiable victim (i.e., Static-99R eligible), mentally cognizant, and supervised in the community at or after the start of data collection (April, 2013). Although new releases were prioritized, the men could have been already released in the community for up to 2 years following their index sexual offense. Release dates were between August 12, 2011, and November 23, 2015, with the first SOTIPS assessment received by the research team on June 1, 2013. Three individuals died during follow-up. They were included in the analyses, with their follow-up ending at their time of death.
Data collection procedures differed at the two sites. The Maricopa County Probation Department mailed paper copies of the Static-99R, the three SOTIPS assessments completed by the probation officers, and the treatment progress reports completed by therapists. Maricopa County also provided data extractions from their internal probation tracking system (APETS) every 3 months, which identified individuals’ current status in the system (e.g., on probation, probation expired, re-incarcerated). In New York City, data acquisition used REDCap, a secure data management system designed for multi-site studies (Harris et al., 2009). Headquarters staff entered demographic and background data when the offender was first assigned to adult probation. After their first meeting with the probationer, probation officers took over data collection. Officers in New York had unique log-in credentials and completed Static-99R, SOTIPS, 6-month progress reports, and reported any changes in probation status on-line. In this site, Static-99R and SOTIPS were scored by the officers and then entered into their data management system; officers could download and print copies of the completed instruments in .pdf format, as needed.
Probation officers at both sites were directed to complete the Static-99R and the first SOTIPS assessment at enrollment. About 6 months after the initial SOTIPS assessment, they were to complete their second SOTIPS assessment and a 6-month progress report. The third (and final) SOTIPS assessment and 6-month progress report were to be completed 1 year after enrollment.
Of the total sample with any information (n = 735), this study included individuals who had at least one SOTIPS total score, recidivism information, and a valid Static-99R score (n = 565). Static-99R scores were considered valid if the individual had ever been charged or convicted of a sexually motivated offense against an identifiable victim (child exploitation image offenses were excluded, as were certain other indecency offenses; i.e., all cases must have at least one Category A offense, see Phenix et al., 2017). Because risk declines the longer individuals remain offense free in the community, Static-99R scores were also considered invalid if the individual had been 2 or more years in the community following release from the index sexual offence prior to the first assessment date (Phenix et al., 2017, p. 18). Individuals with a history of sexual offending who were returned to corrections within 2 years for a nonsexual offence were included. Officers were instructed to select consecutive cases that met the selection criteria (i.e., currently supervised as sexual offender). It was not necessary that the sexual offense supervision order was their only recent criminal justice sanction. As well, their index sexual offense could include a custodial sentence along with the supervision order that brought them to the attention of this study.
Descriptive information of the sample is presented in Table 1. Overall, the New York sample was predominantly African American (36.7%) or Latino (35.6%), whereas the Maricopa County sample was predominantly White (55.9%).
Descriptive Information for Participants From New York City and Maricopa County, Arizona.
Note. For effect sizes, New York = 1; Maricopa = 0. Confidence intervals that do not include 0.50 (for AUC) and 1 (for odds ratios) are in
Risk Assessment Measures
Static-99R
Static-99R (Hanson & Thornton, 2000; Helmus, Thornton, et al., 2012) was used as a static measure of sexual recidivism risk. Static-99R contains 10 items based on commonly available demographic (age, relationship history) and criminal history information (e.g., prior sexual offenses, any unrelated victims, total number of prior sentencing occasions for any offense). Static-99R (and its previous version, Static-99) are the sexual recidivism risk assessment tools most commonly used in corrections and forensic mental health (Bourgon et al., 2018; Kelley et al., 2020; Neal & Grisso, 2014). It can be scored with high rater reliability (for a review, see Phenix & Epperson, 2016) and has moderate ability to discriminate recidivists from nonrecidivists (Helmus, Hanson, et al., 2012).
Static-99R total scores range from −3 to 12, and correspond to the following risk levels: I—very low risk (scores of −3 and −2), II—below average risk (scores of −1 and 0), III—average risk (scores of 1, 2, and 3), IVa—above average risk (scores of 4 and 5), and IVb—well above average risk (scores of 6 and higher; Hanson, Babchishin, et al., 2017). Static-99R risk levels parallel the standardized risk levels developed for general correctional populations by the Justice Center of the Council of State Governments (Hanson, Bourgon, et al., 2017). These standardized risk levels address the crime relevant characteristics of individuals in the criminal justice system, the intensity of correctional supervision and rehabilitation programming needed to manage their risk, their personal strengths, and expected prognosis.
In both jurisdictions, Static-99R was already being used to estimate sexual recidivism risk. If a Static-99R score was not already on file, the supervising officer was instructed to score Static-99R for the purposes of this research study. Static-99R training was provided at both sites by a certified trainer. Rater reliability information was not collected.
Sex Offender Treatment Intervention and Progress Scale
Sex Offender Treatment Intervention and Progress Scale (SOTIPS) is a rating scale composed of 16 dynamic risk items designed to be completed by trained treatment and/or supervision professionals with knowledge of the individual being assessed (McGrath et al., 2013). It was designed to aid clinicians and community supervision officers in identifying and monitoring the supervision and treatment needs of adult males who have committed sex offenses. Items are scored on a 4-point scale; minimal to no need for improvement (0), some need for improvement (1), considerable need for improvement (2), and very considerable need for improvement (3) (McGrath et al., 2012). SOTIPS total scores range from 0 to 48 and are organized into three risk/need groups: low (0–10), moderate (11–20), and high (21–48).
Inter-rater reliability
Each site was asked to double code a 10% sample of SOTIPS. This was done differently in the two sites. In Maricopa County, SOTIPS was double scored for 57 individuals by their probation officer and treatment provider. The demographic characteristics of this subsample mirrored those of the entire Maricopa County sample. Most of the paired assessments were scored within 1 month of each other, although a few were up to 6 months apart. The intraclass correlations (ICC, one-way random, single rater; ICC [1,1] from Shrout & Fleiss, 1979) found in Maricopa County were 0.65 for all 57 participants, 0.78 for those scored within 2 months of each other (n = 37), and 0.82 for those scored within 1 month of each other (n = 26). Although there are no absolute standards for the interpretation of ICCs, Cicchetti’s (1994) standards are widely cited in psychology: ICC’s ≥ 0.60 being good, and ≥0.75 being excellent.
In New York City, 20 cases were double coded by probation officers and their supervisor. The demographic characteristics of the reliability subsample mirrored those of the entire New York sample. The probation officers’ SOTIPS scores were coded in the usual ways based on interview and file review. The supervisors’ ratings were only based on case file notes. Coding lag time ranged from less than a month to 8 months, with the second rating being completed an average of 3.9 months (SD = 2.2; median = 5 months) later. The ICC for these 20 cases was 0.54. Given the gap between initial coding and recoding in New York City, and the non-standard manner in which the re-coding was done, it is difficult to tell how much of the between-rater variability should be attributed to raters using different information.
The average Static-99R scores for both sites were in the range expected for routine/complete samples (Risk Level III; see Table 1). The initial (time 1) SOTIPS scores were also in the moderate range; however, the New York sample was rated as slightly higher risk than the Maricopa County sample (a difference of 2.4 SOTIPS points).
Recidivism Information
In Maricopa County, recidivism data were provided by the Adult Probation Department. An automated search was conducted on the Arizona state criminal history database, which records arrests and dispositions for crimes committed in Arizona. The database was searched using the unique identifier assigned to each participant and by name, race, and date of birth. In New York, data were obtained from the New York State Office of Justice Research and Performance. The database was searched using the unique identifier assigned to each participant. Recidivism information included “flat files” that contained arresting charges, dispositions, and sentencing information. Four raters participated in coding recidivism information, which took place from August through December, 2017. Raters coded for noncontact sex offenses, contact sex offenses, nonsexual violent offenses, and nonviolent offenses as recidivism events. When the only information available was that the individual was returned to custody, the readmission was coded as a new offense (not pseudo-recidivism; see Phenix et al., 2017) and the offense type was coded as unknown. At least one criminal history record was obtained for all individuals. Those without follow-up criminal history records were considered non-recidivists. The date of recidivism was most commonly coded as the arrest date; however, the incident date was used in the minority of cases (<20%) when that information was available.
At-risk time excluded months spent in jail or prison (rounded to a complete month); any lockup time greater than 14 days was considered a month. Overall, 43 individuals were removed from the community for 1 month or more (range of 1–15 months) prior to a subsequent assessment, with the average time of just over 4 months. We calculated the survival end date as the last possible date a probationer could have committed a crime while in the study. If a probationer was incarcerated at the time follow-up ended, we coded his last day in the community prior to sentencing, the day prior to his recorded death, or the day prior to deportation as his survival end date. For supervision violations, raters identified the amount of time (in months) probationers were sentenced. Death (n = 3) and deportation (n = 3) ended the follow-up time.
The follow-up time varied from 1 to 50 months, with an overall average of 39.6 months (SD = 7.3, median = 40). Excluding the 179 months that 43 of the 565 individuals were detained, the average at-risk time was 39.3 months. The Maricopa County sample was followed for about 7 months longer than the New York sample (42 vs. 35 months). Despite the shorter follow-up period, the sexual recidivism rates were higher in the New York sample than the Maricopa County sample (5.1% vs. 1.0%, charges). In contrast, the rate of any new charges was higher in the Maricopa County than New York (34.8% vs. 13.6%). In the combined sample, the overall recidivism rates were 1.9% (11/565) for a sexual offense conviction, 2.3% (13/565) for a sexual offense charge, 5.3% (30/565) for a charge for nonsexual violent or sexual offense, and 28.1% (159/565) for any new charge. The sexual offenses included noncontact offenses. Our intent was to include only criminal recidivism, not technical violations, in the any recidivism category; however, the nature of the charges was unknown in most cases (104 out of 159 offenses were “unknown”). Based on other, indirect sources, a significant proportion of the any recidivism cases in Maricopa County may be violations of the conditions of probation. Each of the recidivism types were hierarchical, such that all sexual charges were included in sexual or violent offenses, and all sexual and violent offenses were included in the category of all criminal recidivism. Consequently, a sexual recidivism event would be included in all three recidivism types whereas a nonviolent recidivism event would only be included in the any recidivism category.
Plan of Analysis
The small numbers of sexual and violent recidivism events prevented meaningful comparisons of the predictive accuracy of the SOTIPS across sites. Consequently, the AUC discrimination analyses used an aggregated sample that combined participants from both sites. For the Cox regression analyses, the hazard rates were calculated separately for each site, then averaged (see below). With the exception of racial composition, the participants at both sites were similar (see Table 1). There was insufficient statistical power to examine racial differences.
AUC
The AUC was used as a measure of the accuracy (discrimination) of the Static-99R and SOTIPS prediction tools. Specifically, the value of the Area Under the Curve (AUC) is the likelihood that an individual with the outcome will have a higher score than an individual without the outcome. The curve in question is the Receiver Operating Characteristic curve (ROC; Swets et al., 2000). AUC values can vary between 0 and 1, with 0.50 indicating no difference between the groups. AUC values above 0.50 indicate that the individuals with the outcome have higher scores than individuals without the outcome. AUC values below 0.50 indicate that individuals with the outcome have lower scores than individuals without the outcome. Although originally designed for signal detection tasks, AUC are widely used to describe the accuracy of medical diagnostic procedures (e.g., scans for brain cancer) and statistical prediction tools. AUC values are expected to be smaller in prognostic studies than in diagnostic studies because the outcome of interest in prognostic studies does not exist at the time of assessment, and may never happen (Helmus & Babchishin, 2017; Royston et al., 2009). Although there are no universal standards for describing effects sizes (for anything), Cohen’s (1988) heuristics are widely used in psychology. For AUC values, 0.56 indicates a small effect, 0.64 indicates a moderate effect, and 0.71 indicates a large effect (Rice & Harris, 2005). The AUC values and their confidence intervals were calculated using the trapezoidal method in IBM/SPSS version 25.
Comparison of sites on descriptive information
The odds ratio was used as the effect size indicator to compare individuals across sites on dichotomous variables (Table 1). Following Fleiss et al. (2003, Equations 6.20, 6.32, and 6.33), the standard errors used the log of the odds ratio with 0.5 added to each cell to stabilize the estimates. The values in Table 1 are transformed back into the original odds metric. When there is no difference between the groups, the expected value of the odds ratio is 1.0.
Following Ruscio (2008), AUC values were used as the effect size indicator when comparing the characteristics of the individuals at the two sites (for the non-dichotomous variables, Table 1). AUC values have the advantage of providing both a significance test (95% confidence interval does not include 0.5) and an indicator of the size of the difference between the groups: 0.56 is a small, 0.64 is a moderate, and 0.71 is a large difference (Rice & Harris, 2005)
Survival analysis
Cox regression survival analysis (Singer & Willet, 2003) was used to compare the predictive accuracy (discrimination) of Static-99R, the first (static) SOTIPS, and the dynamic SOTIPS scores. Cox regression requires that researchers specify the values of the predictor variable for each case at the time of recidivism. The default is the static model in which the case’s first assessment describes the case throughout the full follow-up period. This contrast with dynamic models that allow for the values assigned to a specific case to change when updated assessment information is available. The current study used the default, static model as well as a dynamic model in which scores were updated with each new SOTIPS assessment and stayed until the next assessment, recidivism, or the end of follow-up.
Static and dynamic models were compared using the Akaike Information Criterion (AIC; Burnham & Anderson, 2004) for testing non-nested models. The AIC is computed based on the deviance (-2 log likelihood; -2LL) plus a penalty proportional to the number of parameters (K) used in the model. For the AIC, the penalty is twice the number of parameters (AIC = − 2LL + 2K). Absolute AIC values are not interpretable. The difference between models, however, identifies the model that best fits the data, with low values indicating better fit. Although there are no absolute standards for evaluating differences, Burnham and Anderson (2004) interpret the difference between the minimum AIC and a model’s AIC as indicating the degree of support for the model, with differences of less than 2 indicating no difference between the model fits, 4–7 indicating modest differences, and more than 10 indicating the model with the lower AIC fits the data better than the other.
The effect size indicator for the Cox regression analyses was Harrell’s C (Harrell et al., 1982). Harrell’s C is analogous to the AUC and can be interpreted as the probability that, given two randomly selected individuals, the one with the higher score will reoffend before the individual with the lower score.
Survival analyses were run using the Coxph program (Therneau, 2015) in R statistics version 3.6.1 (R Core Team, 2019). Given the difference patterns of recidivism across sites, sites were considered strata in the Cox regression analyses, that is, the hazard ratios were calculated independently for each site, then averaged. All analyses were verified (double run) independently by two different authors.
Results
The initial SOTIPS assessments were conducted between June 1, 2013, and February 10, 2016, with the last follow-up ending on August 2, 2017. Of the 565 cases with an initial assessment, 87% (495) had at least two assessments, and 79% (448) had three assessments. A small percentage had a fourth (8%, 45) or a fifth assessment (1.2%, 7). The fourth and fifth assessments were included in the analyses of the dynamic SOTIPS scores, but not considered as predictors on their own. The average time between the first and last SOTIPS assessment was 12 months (SD = 7.1, median = 12, range 0–48 months). Static-99R had small correlations with the SOTIPS at Time 1 (r = 0.225, n = 565), Time 2 (r = 0.281, n = 495), and Time 3 (r = 0.228, n = 448). The correlations between Time 1 SOTIPS and Time 2, Time 3, Time 4, and Time 5 SOTIPS were 0.547, 0.411, 0.430, and 0.970, respectively. The correlations between Time 2 SOTIPS and Time 3, Time 4, and Time 5 SOTIPS were 0.637, 0.431, and 0.824, respectively. The correlations between Time 3 SOTIPS and Time 4 and Time 5 SOTIPS were 0.679 and 0.980, respectively. Finally, the correlation between Time 4 and Time 5 SOTIPS was 0.988. The very high correlations with the fifth assessment are unlikely to have substantive meaning because they were based on only 7 individuals. The remaining correlations ranged from 0.411 to 0.679.
Table 2 presents the average SOTIPS score for each of the five assessment periods. The general pattern was a gradual decline (improvement) over time, with the largest decline between the first and second assessment. When the change scores were averaged, the overall change for the sample was small; however, this does not mean that there was little change for individuals. A change of 8 points or more was considered noteworthy. The value of 8 was selected because it was larger than the standard deviation of the difference scores (7.7, 7.5, and 7.1) and would represent a large amount of change when expressed as Cohen’s d (>1.0; Rice & Harris, 2005). Of the 495 individuals who survived to the second assessment, 80 (16.2%) were rated as less problematic (SOTIPS scores declined by 8 or more points) and 50 (10.1%) were rated as more problematic than before (scores increased by 8 or more). Of the 448 individuals assessed at Time 3, equal proportions were rated as improved (12.3%) or deteriorated (12.3%). Overall, between 20% and 25% of the individuals had meaningfully different scores between consecutive assessments.
Average SOTIPS Total Scores During the Follow-Up Period.
Table 3 presents the relationship of SOTIPS scores to recidivism at three time points. Based on AUC values, the relationship between SOTIPS was significant for 8 of the 9 comparisons, with values ranging from 0.590 to 0.695 (median of 0.670). In comparison, the AUC values for Static-99R were significant for only one of the three comparisons: 0.596 for sexual recidivism (ns), 0.608 for sexual or violent recidivism (ns), and 0.632 for any recidivism (95% confidence interval of 0.581–0.683). Readers should note the wide confidence intervals for most of the AUC analyses (with the exception of any criminal recidivism). Also note the declining numbers of recidivists for each successive wave of assessment.
The Relationship of Static-99R and SOTIPS Assessments at Three Time Points to Subsequent Recidivism.
Note. Confidence intervals that do not include 0.50 (no difference) are in
SOTIPS scores range from 0 to 48 with the following risk levels: low (0–10), moderate (11–20), and high (21–28).
Static-99R scores range from −3 to 12 with the following risk levels: I—very low risk (−3, −2), II—below average risk (−1, 0), III—average risk (1, 2, and 3), IVa—above average risk (4, 5), and IVb—well above average risk (6–12).
The comparison between the static, initial assessments and the dynamic SOTIPS assessments is presented in Table 4. For all three outcomes, the dynamic SOTIPS was the best predictor in the univariate analyses, with Harrell’s C’s of 0.715 (sexual), 0.693 (violent), and 0.740 (any). These values are in the moderate to large range, and similar in magnitude to those found for other established risk assessment measures. For any violent recidivism and any recidivism, the best fitting models (lowest AIC) were the multivariate analyses that included the latest SOTIPS scores along with Static-99R. Neither Static-99R nor the first SOTIPS showed significant univariate relationships with sexual recidivism. For sexual recidivism and for sex/violent recidivism, the differences in predictive accuracy between measures were small, with most AIC differences being less than 2. For any recidivism, however, there were strong differences between the multivariate model (the best fit) and the dynamic SOTIPS (the next best model), as well as between the dynamic SOTIPS and initial SOTIPS (AIC difference of 15.85).
Time Dependent Cox Regression Analysis of Static-99R, First SOTIPS, and Dynamic SOTIPS Scores for 565 Individuals.
Note. All analyses include site (Maricopa County, New York) as strata. The dynamic SOTIPS was based on up to five assessments. Confidence intervals that do not include 1 (no difference) are in bold.
Table 5 presents another way of analyzing the data. In these analyses, each individual could have up to 3 SOTIPS scores: Time 1, Time 2, and Time 3. In contrast to the dynamic scores in the previous set of analyses, these SOTIPS scores were static. Subsequent scores did not replace previous scores; instead, subsequent scores were added as a new predictor in Cox regression. These analyses only considered the outcome of any recidivism because there were insufficient sexual or violent recidivism events (Cox regression should have 10 events per predictor variable). Static-99R was retained as a control variable in all analyses.
Cox Regression Analysis of Static-99R and Unchanging (Static) SOTIPS Scores at Three Time Points For Any Recidivism.
Note. All analyses include site (Maricopa County, New York) as strata. Confidence intervals that do not include 1 (no difference) are in bold, as are significant (p < .001) C values.
All variables predicted any recidivism in the univariate analyses. The first SOTIPS was incremental to Static-99R (Model 1). The second and third SOTIPS were incremental to the first SOTIPS (Model 2 and Model 3). The model that included Static-99R and all three SOTIPS assessments (Model 5) was statistically significant (Wald = 17.91, df = 4, p = .001); however, none of the variables made a significant incremental contribution.
Discussion
Consistent with the SOTIPS development study (McGrath et al., 2012), the current study found that repeated SOTIPS assessments predicted recidivism. Importantly, the current study found that the dynamic version of the SOTIPS predicted general (any) recidivism better than the initial SOTIPS assessment. No differences were detected in the predictive accuracies of the initial and the dynamic SOTIPS assessments for sexual and for violent recidivism; however, the low statistical power limits any conclusions concerning the relative predictive accuracy of assessments for these outcomes. For any recidivism, the most accurate model included the Static-99R risk tool along with the dynamic version of the SOTIPS assessment.
These findings contribute to a growing body of research demonstrating that reassessment can improve the prediction of general criminal recidivism (Cohen et al., 2016; de Vries Robbé et al., 2015; Greiner et al., 2015; Howard & Dixon, 2013; Lloyd et al., 2020) as well as sexual recidivism (Babchishin, & Hanson, 2020; Olver et al., 2014). Although incremental effects of reassessment are not always observed (e.g., Kroner & Yessine, 2013; Viljoen et al., 2017), the weight of evidence suggests that assessments proximal to the recidivism event are more informative than more distal assessments.
The current study also demonstrated the value of using Cox regression with time dependent covariates to analyze dynamic prediction recidivism studies (Singer & Willet, 2003). This approach involves specifying a small set of models of theoretical interest and then examining which one(s) provide the best fit to the data. Statistical analyses that treat subsequent assessments as static covariates necessarily lose statistical power because only individuals who have not yet failed are available for the subsequent assessments. Furthermore, the number of possible comparisons increases exponentially with the number of assessments. With three assessments and one control variable (as was the case in the current study), there are 15 possible combinations (4 single variables + 6 pairs + 4 trios + one analysis with all 4 variables = 15, of which only 9 are presented in Table 4). The large number of possible comparisons make it hard to separate meaningful findings from Type 1 errors (falsely perceiving signals in random features of the specific dataset).
Compared to the initial individual differences in recidivism risk at the start of supervision, the increase in predictive accuracy that resulted from considering intra-individual changes was not large. For example, the Harrell’s C for the dynamic SOTIPS predicting general recidivism was 0.740 compared to 0.692 for the first SOTIPS. Large intra-individual change, however, should not be expected. Given all the conditions and choices that shaped the individuals up to the point of their first SOTIPS assessment, the next year on probation would be expected to have a much smaller influence, even if they received effective counseling. Furthermore, the increased predictive accuracy for reassessment may not be due to individual change. Subsequent assessments may simply be better measures of pre-existing, stable propensities. The probation officers making the second and third SOTIPS ratings should have learned more about the individuals on their case loads than they knew at the start of supervision. They may also have become more skilled at scoring the SOTIPS. Experience with other sexual recidivism risk tools has found that it takes about 20 real cases before scorers are proficient (Hanson et al., 2014). Evidence from other studies, however, suggest that some intra-individual change was likely. Dynamic variables, like those measured in SOTIPS, change in predictable ways for individuals during the course of psychological treatment (Olver et al., 2014) and the sexual recidivism risk gradually declines over time (Hanson et al., 2018).
The purpose of dynamic risk tools is to identify targets for intervention and to monitor progress towards orderly reintegration. Much more needs to be known about what actually changes in response to effective intervention. Given that many risk factors change at the same time (Cording, 2018), the actual change process may not require addressing each of the separate items listed on the SOTIPS. Instead, successful reintegration may be better described as a global reorientation of life goals (Maruna, 2001). The SOTIPS items provide valuable, if fallible, clues about the direction an individual is headed.
One issue raised by the current findings is the difficulty of associating recidivism rates with specific risk scores. In New York, the sexual recidivism rates were 5 times higher than in Maricopa County, whereas the overall recidivism rate was 3 times higher in Maricopa County than in New York. It is unlikely that these represent true differences in the rates of real or observed reoffending. Instead, these differences are most likely the result of difference in how new offending is policed, legally processed, and made available for this research study (e.g., the large proportion of unknown offenses). There was an insufficient number of sexual recidivists to justify statistical tests of calibration (Helmus & Babchishin, 2017) with the recidivism rate norms asserted by the test developers (McGrath et al., 2012).
Limitations
One of the significant limitations of the current study was the low number of sexual recidivists (9 in New York; 4 in Maricopa County). With less than 15 sexual recidivists in total, tests of discrimination were severely underpowered and test of calibration were impossible. We were also unable to address more nuanced questions concerning the validity of SOTIPS scores for different subgroups across the sites (e.g., race, offense type). Although the overall sample size was reasonably large (565) and the attrition rate was low, the follow-up period was short. A 1 year follow-up period is adequate, if not ideal, when the outcome of interest is any recidivism. For sexual recidivism, 5 calendar years is recommended (Collaborative Data Outcome Committee, 2007). Given that it often takes considerable time for detected sexual recidivism to be recorded in official records, an additional buffer period of a year (6 years total) is recommended if the research questions concern official sexual recidivism rates at 5 years. The length of the required buffer period will depend, of course, on the efficiency with which police records are made available to researchers. It should also be possible to increase the observed recidivism rates by considering recidivism events in adjacent jurisdictions, which was not done in the current study. Another caution is that all observed recidivism rates underestimate the actual rate by some unknown amount (commonly known as the dark figure of crime).
Additionally, there were a considerable number of recidivism incidents in Maricopa County where we could ascertain that the individual was incarerated, but could not determine the reason. These unknown cases were included in the any offense recidivism category, but many of them are likely to have been probation violations that did not appear in our other sources. It is possible, however, that a few could also have been sexual or violent in nature, which would have decreased the observed recidivism rates for these outcomes.
Another limitation is that the recidivism outcome may not be fully distinct from the SOTIPS ratings. It is possible that the apparent superiority of proximal SOTIPS scores over initial SOTIPS scores was the result of probation officers knowing that charges were coming. In particular, confounding of the predictor and the outcome is most likely when the same probation officer rated the SOTIPS and then soon after initiates charges for non-compliance. Consequently, the general recidivism outcome is the outcome most vulnerable to this threat to validity. It is unlikely that probation officers would have knowledge of new sexual or violent offenses well before these events were known to police. Future research should consider designs that increase the separation between the SOTIPS ratings and the indicators of recidivism.
The absence of rater reliability information for Static-99R made it difficult to tell how much of the relatively weak discrimination of the Static-99R in this study should be attributed to poor implementation, or to other features of this study. As well, the method used to assess rater reliability in New York was far from optimal: the reliability raters did not conduct interviews with the individuals, and, instead, used case files that often contained 5 months worth of new information. The time lag would be expected to introduce substantial variation given that the 6 months stability of SOTIPS scores completed by the same rater was in the r = 0.50 range. Nevertheless, the intraclass correlation obtained in Maricopa County for SOTIPS scores within 1 month of each other provided some confidence in the reliability of SOTIPS ratings obtained in this study.
Directions for Future Research
An outstanding research question concerns how much of the longterm decline in recidivism risk (Hanson, 2018; Hanson et al., 2018) can be attributed to corresponding declines in observable dynamic risk factors (e.g., SOTIPS scores). On average, sexual recidivism risk is reduced by 50% for every 5 years individuals remain sexual offense free (Hanson et al., 2014; Moore, 2018). Although a correlation between the longterm declines in sexual recidivism risk and improved SOTIPS scores is expected, there are no longterm (10+ year) studies that simultaneously measure changes in dynamic risk factors and sexual recidivism risk. All the studies of dynamic risk for sexual offending have either conducted reassessments during incarceration (Olver et al., 2014) or during the first few years following release from the index sexual offense (Babchishin & Hanson, 2020; Lasher & McGrath, 2016; Viljoen et al., 2017). It is possible that individuals could retain certain risk relevant characteristics well after their risk for sexual recidivism has been extinguished. Consider, for example, a 60-year-old man with some risk factors for general crime (hostility, substance abuse, no employment, antisocial friends) but whose last and only sexual crime conviction was when he was 35. Such an individual may be genuinely low risk for sexual recidivism, even if he is not a fully functioning member of society. If, however, future research identifies a substantial portion of individuals who remain sexual offense free for decades while continuing to display high SOTIPS scores, such a finding would be a serious challenge to our current models of offender reintegration.
The current study only considered one dynamic version of SOTIPS scores, namely, new scores replaced all previous scores. Given that all assessments contain error, it may be possible to improve prediction by considering both current and previous assessments. For example, imagine two individuals who are both rated as low risk on the current assessment; however, one of which was rated as high risk on six previous assessments and the other was consistently rated as low risk. Intuitively, the first individual would appear to be the riskier of the two. Consequently, it is worth examining dynamic prediction models in which the dynamic values combine current and previous assessments of functioning, such as rolling averages (Lloyd et al., 2020) and extreme scores (best, worst; Babchishin & Hanson, 2020). Such alternative model testing requires more statistical power than available in this study.
Future research should also explore patterns of changes on clusters of risk factors. It is likely, for example, that individuals will show positive changes on variables related to motivation for change and treatment engagement prior to showing change on longterm propensities. Research is also needed on the extent to which the SOTIPS items measure the same latent constructs over time (i.e., measurement invariance (Liu et al., 2017). Changes in predictive accuracy could be related to changes in the meaning of SOTIPS items over time. For example, evaluators may overestimate risk in their first assessment and then lower their risk ratings once they are familiar with the individual, even if the individual has not actually changed. This measurement invariance testing is complex and beyond the scope of this paper.
Implications for practice
The current study supports the use of the SOTIPS for assessing the quality of psychological and community adjustment for individuals with a history of sexual offending. In particular, SOTIPS scores appear to be indicators of dynamic (changeable) risk such that more proximal assessments provide a better indicator of risk than initial assessments. Consequently, routine reassessments are recommend. The study does not, however, provide any direct evidence concerning the optimal frequency of reassessment.
Conclusion
The assessment of individuals with a history of sexual offending has progressed considerably during the past few decades. Major risk factors for sexual recidivism have been identified, and combined into structured risk tools that are now routinely used in the US, Canada, and internationally. In particular, a number of risk tools have been developed to identify the risk relevant problems that could be addressed in treatment and supervision, such as the SOTIPS. Previous research has found that these dynamic risk tools predict sexual recidivism (van den Berg et al., 2018). The current study joins a growing body of research indicating that the risk they assess is dynamic, such that reassessment improves prediction (e.g., Babchishin & Hanson, 2020; Olver et al., 2014). The next step is identifying the most effective methods of intervening in the risk and protective factors identified in the SOTIPS and similar risk tools, and then demonstrating that the resulting improvements in functioning signal a decreased likelihood of recidivism.
Footnotes
Acknowledgements
The authors wish to thank Chris Hoefer, our project coordinator, who was responsible for all administrative aspects of this project and facilitated data collection in Maricopa County. We would also like to thank Cathy Strobel-Ayres who facilitated data collection in New York City and DCJS. The authors would also like to thank the probation officers and officials of the Adult Probation Departments in Maricopa County, Arizona and New York City.
Declaration of Conflicting Interests
The author(s) declared the following potential conflicts of interest with respect to the research, authorship, and/or publication of this article: R. K. Hanson and David Thornton are co-authors and certified trainers of the Static-99R. The copyright for this instrument is held by the Government of Canada and the authors do not receive royalties for its use.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, or publication of this article: This project was supported by Award No. 2012-AW-BX-0153, awarded by the National Institute of Justice, Office of Justice Programs, U.S. Department of Justice. The opinions, findings, and conclusions or recommendations expressed in this publication are those of the authors and do not necessarily reflect those of the Department of Justice. Reoffense data for the New York City cohort was provided by the New York State Division of Criminal Justice Services, Office of Justice Research and Performance (DCJS). The opinions, findings, and conclusions expressed in this publication are those of the authors and not those of DCJC. Neither New York State nor DCJS assumes liability for its contents or use thereof.
