Abstract
Pretrial assessments are criticized for inherent biases. We conduct research in Kentucky to assess the predictive validity and differential prediction by sex of one pretrial assessment, the Public Safety Assessment (PSA). Our research is unique because we find equal base rates by sex for missing a court date, which allows us to assess for error rate balance by sex. We find the PSA to have predictive validity within acceptable ranges for the criminal legal field. The analyses show a lack of evidence of predictive bias for any arrest or missing a court date, and we find equal error rates for five different error measures. The analyses contribute to methodological debates about how to measure predictive bias with assessments.
There is an ongoing debate about “fairness” related to the development and use of actuarial pretrial release assessments 1 on the grounds of potential disparate impact by race and sex. Fairness, however, is a normative concept that cannot be measured empirically (Chouldechova, 2017; Skeem et al., 2016). Indeed, fairness is a social quality that is deeply embedded in local community moral bonds and provides the basis for social cohesion and solidarity. Data scientists have pointed out that “prevailing definitions of fairness typically do not map on to traditional social, economic or legal understandings” of fairness (Corbett-Davies & Goel, 2018, p. 9), while artificial intelligence researcher Kalluri (2020) has characterized “fair” and “good” as “infinitely spacious words” that sidestep the power dynamics at the heart of their use. Variations in the ways jurisdictions define and record being arrested while on pretrial release or not appearing at a court date can result in the distributions of these “risks” looking different when their occurrence is fairly similar, or conversely, looking erroneously comparable. The differences in policies and practices across jurisdictions affect not only the outcome base rates, but also shape the accumulation of factors purported to lead to those outcomes (Koepke & Robinson, 2018).
In this paper, we contribute to debates about pretrial release assessments by testing multiple definitions of predictive bias between people categorized as male and as female 2 using the Public Safety Assessment (PSA). Although there are several definitions of predictive bias (Berk et al., 2021), a main concern is the effectiveness of an assessment to classify individuals into relatively homogenous groups. In the pretrial context, assessments are used to inform release decisions and set conditions of release based upon an individual’s likelihood to miss court or be rearrested. This is not dissimilar from what one might consider an effective medical screener that can group people according to the presence or absence of a disease—or likely susceptibility to a disease—and contribute to identification of appropriate medical interventions, including preventative strategies (Maxim et al., 2014). Assessments provide estimates of the probability that an event will happen (e.g., illness, rearrest), and, as such, not all the resulting classifications are correct. Simply, people expected to be arrested (or sick) are not (i.e., false positive) and people expected to avoid arrest (or illness) are (i.e., false negative). Error is an inherent component of any probabilistic decision-making tool—including pretrial assessments.
Problems arise when developing a pretrial release assessment that produces equal probabilities (i.e., calibration) and equal error rates (i.e., error rate balance) across subgroups (Berk et al., 2021). Structural racism, gender socialization, and other cultural forces create conditions under which people of color and males tend to be arrested more frequently than other people, which results in higher criminal history scores and as such they score higher on pretrial risk assessments (Cohen & Lowenkamp, 2019; Skeem & Lowenkamp, 2016). The differences in base rates and mean scores across groups makes it mathematically impossible for assessments to achieve calibration and have equal error rates for each group (Chouldechova, 2017; Kleinberg et al., 2018). When average assessment scores vary by subgroups there is an underlying data generating process that can only be understood by looking at how groups score on the presence and absence of the defined risk factors. Corbett-Davies and Goel (2018) showed the importance of the infra-marginality problem with criminal justice release assessments in which group-level differences in underlying risk distributions result in variation in group level error rates. Regardless of the accuracy of the assessment, individuals within the subclass with the higher average score will have a higher false positive rate relative to the other subclass and individuals within the subclass with lower mean scores will have higher false negatives.
Although their study has been refuted for methodological problems, ProPublica showed error rates differed by race when assessing 2-year recidivism rates (Angwin et al., 2016). The ProPublica article contributed to a false sense that pretrial release assessments are inherently biased even though their analyses were not focused on a pretrial sample or pretrial outcomes (i.e., 2-year recidivism study) (Flores et al., 2016). A body of research validating pretrial release assessments has emerged, with the consensus being that these instruments do not inherently exacerbate bias by race or sex (Desmarais et al., 2021). We extend recent research on predictive bias by race showing that pretrial outcomes may not always have base rate or mean score differences by subgroups and that it is possible to assess bias by calibration (DeMichele et al., 2020) and error rate balance (DeMichele & Baumgartner, 2021).
We contribute to debates about predictive bias with pretrial release assessments by assessing predictive bias by sex using the PSA. First, we present perspectives on the relationship between being arrested and sex. Males consistently have higher rates of criminalized behavior and violence, yet release assessment studies rarely investigate the sex specific fit of their scales. Second, we review literature on predictive bias and build on prior methodological recommendations by others examining predictive bias by race (DeMichele & Baumgartner, 2021). Although there are differences in pretrial arrests by sex, we do not find differences in rates of not appearing in court. The “failure to appear” (FTA) base rate is the same for males and females (14.8%), which allows us to assess five metrics assessing error rate balance. The conclusion emphasizes that researchers need to apply methods that demonstrate an understanding of their data and arguments for and against pretrial release assessments should focus on intended policy goals.
Criminalized Behaviors and Sex
Most of the opposition to pretrial release assessments has focused on predictive bias by race, but there is potential for these assessments to overclassify females. The main concern is that pretrial release assessments may place females into higher risk categories than is warranted given the nature of their propensity to engage in these behaviors (Skeem et al., 2016). This could produce situations where judges are basing their decisions on inflated risk profiles for females that result in setting higher bond recommendations, applying overly strict supervision conditions, and even recommending pretrial detention. Studies analyzing data on arrests, convictions and self-reported crimes and victimizations consistently find that men commit more crimes, especially more serious and violent crimes (Denno, 1997; Fox & Fridel, 2017; Lauritsen et al., 2009; Rennison, 2009). Evidence shows that sex differences in crime patterns include recidivism and gendered pathways to crime (Belknap, 2006; Salisbury & Van Voorhis, 2009; Webb, 2017). Controlling for typical factors such as criminal record history, age, and drug use, females are less likely to re-engage in criminalized behaviors than males.
Sex differences in law-breaking frequency has critical implications for release assessments. Smith et al. (2009) conducted a meta-analysis to assess how well the LSI-R assessment predicted recidivism for females. They concluded that the effect sizes for males and females are statistically similar, though Smith et al. (2009) noted a high degree of variation across studies. Skeem et al. (2016) assessed predictive bias by sex of a post-conviction assessment tool that omits sex, the Post-Conviction Risk Assessment (PCRA), and found that the distribution for risk of arrest differed between males and females. Females had lower mean scores on the PCRA due to higher criminal record history scores for males. The PCRA strongly predicted any arrests and arrests for violent crimes for both sexes, but it overestimated both rates for females compared to males (Skeem et al., 2016).
The studies cited above are focused on post-conviction samples that show large differences in recidivism rates between males and females. Although the post-conviction literature is helpful to assess pretrial assessments, but more recent research has demonstrated that some post-conviction recidivism patterns may not hold for pretrial. DeMichele and Baumgartner (2021) noted that pretrial follow-up periods are shorter than traditional post-conviction studies, and that many crime patterns found in longer recidivism studies do not hold in pretrial data (e.g., lower base rates, smaller differences by race and sex). For these reasons, it is essential to consider assessment research focused on the pretrial phase using pretrial outcomes (e.g., missed court appearances, new arrests).
Pretrial researchers found that recidivism rates for males and females awaiting trial with the same COMPAS risk scores differed such that a two-point differential persisted across scores (Corbett-Davies & Goel, 2018). Females with a COMPAS score of 6 (i.e., medium risk) were rearrested at about the same rate as males with a score of 4 (i.e., low risk) and females with a score of 7 were rearrested at around the same rate as males with a score of 5. Corbett-Davies and Goel’s (2018) findings are essential to consider when introducing assessments to real-world settings because ensuring equal probabilities of negative outcomes conditioned on a score is a fundamental form of predictive bias. Finding different probabilities of negative outcomes within scores across sex presents ethical and practical challenges related to detention, supervision conditions, and public safety. Criminological literature demonstrates that males are undoubtedly more involved in violent and serious crime, whereas females tend to be involved with more minor criminalized behavior and suffer from ongoing and prior abuse and trauma (Richie & Eife, 2021).
Cohen and Lowenkamp (2019) found that the pretrial risk assessment instrument (PTRA) used in the U.S. federal pretrial system was a valid predictor for males and females. They found that males and females had similar rearrest probabilities for any reason, but, as was expected, males had “significantly higher likelihoods of being arrested for violent offenses than females.” These findings were somewhat surprising as Cohen and Lowenkamp (2019) “expected to see males uniformly failing at different rates than females across the rearrest categories of interest.” They found inconsistent differences in the outcomes by sex and hence were somewhat equivocal in their interpretation by suggesting that “future research should monitor these differences” (p. 256) and that there is “some evidence of overprediction depending upon the outcome being examined” (p. 234).
There are few studies assessing the sex specific performance of pretrial assessments, yet pretrial decisions are high stakes for individuals (and their families). Pretrial detention is associated with loss of employment, higher conviction rates, and higher rates of incarceration (Dobbie et al., 2018; Mayson, 2019). This is the first peer-reviewed study assessing the validity and predictive bias by sex for the PSA testing for calibration and error rate balance, whereas other researchers have assed PSA validation and tested for differential prediction by race (Brittain et al., 2021; DeMichele et al., 2020).
Definitions of Predictive Bias
There are several competing definitions of predictive bias that cannot always be achieved simultaneously (Chouldechova, 2017; Kleinberg et al., 2018). ProPublica assessed error rate balance (i.e., classification parity) and found that a greater proportion of black defendants had false positives (i.e., scored higher risk but did not recidivate) and a greater proportion of white defendants had false negatives (i.e., scored lower risk but recidivated). Flores et al. (2016) tested for moderation to assess calibration (i.e., showing that a score X has the same meaning regardless of race). Testing for moderating effects is a technique that has been used by the Standards for Educational and Psychological Testing (American Educational Research Association, American Psychological Association, & National Council on Measurement in Education, 2014) and the Principles for the Validation and Use of Personnel Selection Procedures (Society for Industrial and Organizational Psychology, 2018) since the late 1960s, it was recently supported in criminology (Skeem & Lowenkamp, 2016), and was identified as an essential test for predictive bias by data scientists (Kleinberg et al., 2018).
In this paper, we are responding to calls for more research on predictive bias by sex (Cohen & Lowenkamp, 2019) and contribute to methodological approaches to assess predictive bias with pretrial assessments (DeMichele & Baumgartner, 2021). The ProPublica v. Flores et al. (2016) debate is a methodological one in which ProPublica used a simplified definition of bias (i.e., error rate balance). ProPublica ignored the mathematical reality that when base rates differ, error rates will vary because they overlooked the common issue of infra-marginality that exits in many criminal legal system datasets (Corbett-Davies & Goel, 2018). The group with higher rates of negative outcomes will score higher (on average), and, as such, have higher false positive rates, and the group with lower rates of negative outcomes will score lower (on average), and, similarly have higher false negative rates (Berk et al., 2021; Chouldechova, 2017). These are realities based on the underlying distributions of the data that need to be considered before assessing predictive bias. Although there is no silver bullet or single measure that definitively answers the bias question, what is well-established is that the central purpose for assessments is to classify individuals such that there are equal probabilities of outcomes by scores regardless of group membership. This is not to suggest that error rates are unimportant, but rather to emphasize that assessments are used to provide criminal legal system stakeholders with information about the likelihood of whether a person awaiting trial will miss court, commit a new crime, or commit a violent crime.
We extend an approach to the predictive bias debate that accounts for the variability in pretrial populations. DeMichele and Baumgartner (2021) contributed to the methodological debate by highlighting that pretrial data used to study assessments are generated in highly localized institutional frameworks that are characterized by varying distributions of potential negative outcomes and heterogeneous arrest practices. Simply, they recognized that there are several pretrial outcomes of interest, missed court appearance and new arrests, and researchers need to determine if there are equal base rates and similar mean scores by subgroups that allows for multiple tests of predictive bias (DeMichele & Baumgartner, 2021). We build on this perspective to assess calibration across three pretrial outcomes of “failure to appear” (FTA), 3 new criminal arrests (NCA), and new violent criminal arrests (NVCA). Next, we test for error rate balance with FTAs because mean scores and base rates do not differ.
The Public Safety Assessment
The PSA 4 is completed with a criminal history review using administrative records collected from the pretrial services agencies, jail booking information, and criminal history repositories. This process avoids the need for an interview to collect self-reported data and is expected to reduce the time needed to complete an assessment, with the intention to provide information more quickly to require less time for release and supervision decisions. The PSA is available to the public and jurisdictions can use a web-based application to implement the PSA on their own (https://advancingpretrial.org).
Study Methods
The analysis is a test of predictive bias of the PSA using an historical dataset of all individuals booked into jail in Kentucky between July 1, 2013and December 30, 2014. The dataset was made available to the authors by Arnold Ventures (formerly the Laura and John Arnold Foundation) through a cloud-based repository as part of a larger study of the PSA and legal actor decision making. The instrument development team (Luminosity) collected the data from Kentucky and processed the datasets to develop deidentified analytic files with the binary outcome variables (i.e., failure to appear, new criminal arrest, and new violent criminal arrest), assessment factors (see details below), age, race, sex, and booking and release dates. The dataset for the current study was collected as part of post-development validation of the PSA. 5
Scoring the PSA
The PSA is scored by applying the predefined weights to each factor, summing the weights, and converting the weighted scores to scale scores that range from 1 to 6 for each of the outcomes. Three of the FTA factors are binary indicators (i.e., pending charge, any prior conviction, any FTA older than 2 years) scored as 0,1, and prior FTA within past 2 years (0 = 0, 1 = 2, and 2+ = 4). These risk scores range from 0 to 7 and were converted into an FTA scale score according to the PSA instructions. The weighted risk scores are converted to FTA scale scores as follows: 0 = 1, 1 = 2, 2 = 3, 3 and 4 = 4, 5 and 6 = 5, and 7 = 6.
The NCA risk score includes seven factors that are weighted as follows: prior misdemeanor (No = 0, Yes = 1), prior felony conviction (No = 0, Yes = 1), pending charge (No = 0, Yes = 3), prior incarceration sentence (No = 0, Yes = 2), prior violent convictions (0 = 0, 1 or 2 = 1, 3+ = 2), prior FTA in past 2 years (0 = 0, 1 = 1, 2+ = 2), and age at current arrest (23+ = 0, 21 and 22 = 2, 20 or younger = 2). These weights are converted into an NCA scale score as follows: 0 = 1, 1 and 2 = 1, 3 and 4 = 3, 5 and 6 = 4, 7 and 8 = 5, 9 to 13 = 6.
The NVCA risk scores range from 0 to 7. Three of the factors are binary indicators (i.e., any pending charges, any prior convictions, and current offense is violent and ≤20 years old) measured as 0,1; current violent offense is a binary factor measured 0,2; and prior violent conviction is measured as follows: 0 = 0, 1 and 2 = 1, and 3+ = 2. These scores are converted into a scale score ranging from 1 to 6 as follows: 0 = 1, 1 = 2, 2 = 3, 3 = 4, 4 = 5, 5 or above = 6. The NVCA scale scores are used to create a binary indicator with defendants with an NVCA scale score of 5 and 6 receiving a violent flag.
Dataset and Research Questions
The analyses include adults released prior to their trial in Kentucky during the 18-month study period. Cases were removed if they were under 18 years of age at the time of booking (n = 679) or had booking dates after December 2014 (n = 45,299). Cases were defined as released if they had release dates and did not have disposition dates. The validation released sample included n = 164,597 (68.5%) with n = 75,662 (31.5%) detained.
Cases were scored using the scoring criteria for the PSA that provides separate scale scores for each of the three outcomes. FTA is a variable measuring whether a released individual missed any court date before the disposition of their case. New criminal arrest (NCA) is a variable measuring whether a released individual was arrested for any reason prior to the disposition of their case. New violent criminal arrest (NVCA) is a variable measuring whether a released individual was arrested for a violent charge based on the PSA definitions prior to the disposition of their case. The scoring and weighting rules were applied for each of the three outcomes (i.e., FTA, NCA, NVCA) for each case to address the following research objectives: 6
Assess Predictive Validity by Sex: How accurately does the PSA predict each of the three outcomes of interest by sex?
Assess Differential Prediction by Sex: Does the PSA provide different results based on sex? We expect to find that the PSA predicts equally well across sex (i.e., sex will not moderate the relationship between the PSA and failures).
Assess Error Rate Balance for FTAs by Sex: Do males and females have similar error rates for positive and negative classifications for FTAs?
Analyses
We assess predictive validity using the Area Under the Curve (AUC) Receiver Operator Characteristics (ROC) estimates. AUCs are commonly used to evaluate assessment tools because they are not influenced by base rate differences and allow for making comparisons across models and groups. The AUCs range from 0 to 1.0 with 0.5 referring to random chance and 1.0 referring to perfect prediction. The AUC provides a rather intuitive interpretation as it reports the likelihood that when randomly selecting a case that had one of the outcomes, that case would have a higher score on the PSA than a randomly selected case that did not have one of the outcomes. Desmarais et al. (2016) have conducted meta-analyses and evaluations of criminal legal system assessment instruments. They concluded that AUC values of 0.54 and below are poor, 0.55 to 0.63 are fair, and 0.64 to 0.7 are good, with values higher than 0.71 being excellent. Using these ranges, the ROC values for PSA for the three outcomes are in the good range. 7
Second, we assess the presence of predictive bias by sex with the PSA. We use a moderator regression technique commonly cited in psychological studies and testing literature for the three outcomes (e.g., Sackett et al., 2008). This approach estimates four regression models to assess the extent to which “. . .a given score will have the same meaning regardless of group membership (e.g., an average risk score of X will relate to an average recidivism rate of Y for all relevant [sub] groups)” (Monahan et al., 2017, p. 193). Predictive bias is tested by assessing the extent to which subgroups have similar (i.e., not significantly different) intercepts and slopes (i.e., they possess similar regression lines) (Skeem & Lowenkamp, 2016). The approach is designed to test for whether the PSA scores are moderated or conditioned by sex as they predict the outcomes (i.e., a given score on the PSA does not have the same meaning for males and females).
Lastly, we analyze error rates using a contingency table for FTAs. Males and females have equal base rates of FTAs, which allows us to assess for error rate balance (DeMichele & Baumgartner, 2021). Contingency tables allow for calculating several metrics to assess the quality of a classification instrument (Powers, 2011). We review five metrics to compare error rates by sex for positive (i.e., FTAs) and negative (i.e., no FTA) cases and an overall error rate.
Findings
In Kentucky, 68% (n = 164,597) of those awaiting trial were released, with 34 years old being the average age, and 81% (n = 133,517) being white. Nearly 70% of the individuals are male (n = 113,376). The overall base rates for the three outcomes are: FTA rate of 15%, an NCA rate of 11%, and an NVCA rate of 1.1%. FTA rates are equal for males and females (14.8), but they are significantly different (p > .001) for NCA (11.1% vs. 9.7%), and NVCA (1.3% vs. 0.7%).
Risk Factors by Sex
Table 1 shows how the risk factors are distributed across the entire released sample and by sex. The PSA includes a total of 11 factors in which each of the three outcomes are modeled with 4 (FTA), 5 (NVCA), and 7 (NCA) factors. Reviewing the distribution of risk factors by sex allows us to understand what factors contribute to the risk scores. Males and females do not vary on pending charges (19.2% vs. 18.7%) or FTAs within the past 2 years (30.2% vs. 31.4%). Males have more extensive criminal histories with a greater proportion of men have any prior convictions (77.1% vs. 68.8%), misdemeanor convictions (75.4% vs. 67.5%), and felony convictions (33.4% vs. 19.9%). Although violent arrests during the pretrial period are rare in Kentucky, the differences in criminal history were starkest with violent convictions. About four times as many males (5.3%) had three or more violent convictions than females (1.3%), and nearly double the proportion of males (21.0%) had one or two prior violent convictions compared to women (10.9%). Although there is a higher proportion of males that are 22 years old or younger, there were the same proportion by sex of young people booked on a current violent offense. Males and females have similar mean FTAs scores (2.9 vs. 2.8), and males have higher NCA mean scores (3.0 vs. 2.6) and NVCA mean scores (1.6 vs. 1.2).
Public Safety Assessment (PSA) Distribution of Factors for Released Individuals and by Sex.
Note. Male and female individuals will not add to the total N released due to the exclusion of records with “Unknown” sex. Percentages may not add up to 100% due to rounding.
Predictive Validity
Table 2 presents the FTA rates by PSA scale score for men and women. In the case of FTAs, males and females have similar rates across the PSA scores with rates increasing as the scale increases. For FTA scores of 1 to 3, rates range between 7% and 14%, and for scores of three or higher the rates range between 20% and 34%, with females with scores of 6 having slightly higher rates. In total, the FTA rates are nearly identical by sex when conditioned on score, and the overall base rates are identical. There is a surprisingly similar pattern with NCA rates, with the exception that males have a higher base rate (11.1% vs. 9.7%, p < .001). There are no significant differences between males and females for any of the scales scores for FTA or NCA. Rather, there is a high degree of parity in the score specific rates between males and females.
Public Safety Assessment (PSA) Failure to Appear, New Criminal Activity and New Violent Criminal Activity Between Males and Females.
Note. PSA = Public Safety Assessment; AUC = area under the ROC curve; CI = confidence interval.
p < .001.
Significant Difference by sex. p < .001, H0: AUCmale − AUCfemale = 0.
As expected, there are significant differences between males and females for the NVCA scale, with a significantly higher proportion of males arrested for a violent crime among scores of 2 and 3. The base rate for males, although very low, is nearly double the base rate for new violent arrests for females. Males with low NVCA scores (1–4) have a significantly higher rate of NVCAs relative to females with similar scores. There are 1,825 arrests for a new violent crime during pretrial and males account for 81% (n = 1,493) of the violent arrests and 70% of the release pretrial population. There are few females in the higher scale scores for NVCA, which prevents us from making too much of the parity with this scale other than to recognize that arrests for violent crimes during the pretrial period are exceptionally rare in Kentucky, and that males have greater involvement in violent crime.
Table 2 includes tests of differential validity by comparing the AUCs for males and females. The AUCs in Table 2 show that there are no statistically significant differences in the predictive validity between males (AUC = 0.642) and females (AUC = 0.655) (p = .016) for FTAs. Surprisingly, the AUCs for NVCAs do not differ between men (AUC = 0.654) and women (AUC = 0.657) (p = .898). There are significant differences (p < .001) in validity between male (AUC = 0.653) and female (AUC = 0.637) defendants for the NCA—other than the AUC for NCAs for female defendants, all the AUCs are in the good range.
Predictive Bias: Testing for Moderation
We estimate four logistic regression models to test for moderation by sex. In model 1, we estimate the effect of sex on each outcome, and the second and third models are fit with only the PSA score and both sex and the PSA score for each outcome, respectively. The final model includes the sex indicator, the PSA score, and an interaction term of sex by the PSA score. The interaction term tests to what extent the likelihood of a negative outcome during the pretrial period is a matter of whether sex moderates the PSA scores such that a given score has a different meaning for each sex.
Table 3 presents the odds ratios and confidence intervals for the four regression models for FTA, NCA, and NVCA. There are consistent significant effects for sex for NVCA and insignificant interaction terms for the three outcomes. Figure 1 include the predicted probabilities for each of the outcomes by the scores for males and females using the regression equations from model 4 in Table 3.
Logistic Regressions Models Testing the Predictive Fairness of the Public Safety Assessment (PSA) by Sex for FTA, NCA, and NVCA.
p < .001.

Predicted probabilities of pretrial new criminal activity, new violent criminal activity, and failure to appear by Public Safety Assessment (PSA) NCA, NVCA, and FTA scores between males and females.
The relationship between the FTA scores and FTAs are not moderated by sex, but rather have nearly identical predicted probabilities for each of the scale scores. In Table 3, none of the models including sex are significant nor are the odds ratios large. There are consistently significant odds ratios for the FTA score such that a one-point increase in the FTA score is related to a 46% to 48% increase in the odds of an FTA and including sex (model 3) does not diminish the relationship of the PSA score with FTA (model 2). Figure 1 shows overlapping lines (i.e., equal predicted probabilities) of an FTA by sex for each PSA score.
Table 3 includes the results for similar regression models assessing differences by sex for NCAs and NVCAs. For NCAs, any main effects for sex (OR = 1.16, p < .001) in model 1 were diminished once NCA score was included in models 3 and 4. Across the three models, a one-point increase in the NCA score was associated with a 50% increase in the odds of an NCA, and an insignificant interaction term. In Figure 1, we plot the predicted probabilities of an NCA or NVCA by sex (using model 4). The NCA predicted probabilities demonstrate a similar trend found with FTAs of equivalent probabilities by score.
In Table 3, we report the four logistic regression models testing for sex differences on the NVCA scale. 8 Although only 1% of the overall sample had a new violent arrest during their pretrial release, we find that males have twice the probability (OR = 2.02, p < .001) of a new violent offense (models 1 and 4). The regression models indicate that there is a difference in the intercepts (i.e., males have higher intercepts than females), which would result in overclassifying women as high risk. Females have significantly fewer arrests for crimes categorized as violent during pretrial than males, and, although we did not find differences in the slopes, we did find that sex is a meaningful predictor for violent arrests. A one-point increase in the NVCA scale is associated with a nearly 60% increase in the odds of a new violent crime during pretrial release.
Error Rates by Sex
There are equal base rates and similar mean scores for FTAs, which allows us to assess five error rates using a contingency table. We collapse cases with scores from 1 to 4 and 5 to 6 into two groups. Contingency tables are two-by-two tables that allow for assessing the number and proportion of correct (i.e., true) and incorrect (i.e., false) classifications. Berk et al. (2021) provided a thorough treatment of a contingency table in criminology; we summarize key terms here to ease interpretation.
Binary classifiers are used to sort cases into one of two groups according to probabilities of an outcome. An individual with a high probability of the outcome (i.e., FTA) is referred to as a predicted “positive” case and an individual with a low probability of the outcome is referred to as a predicted “negative” case. These labels are not indicative of a qualitative state, they indicate whether a case is predicted to have an FTA (i.e., positive) or absence (i.e., negative) of an FTA. It is important to understand that assessments are not only classifying individuals to the positive class (i.e., those with probabilities above a certain threshold) of an FTA, but assessments also classify individuals into the negative class to ensure that those on the lower end of risk receive limited supervision or few conditions.
A contingency table includes information about predicted and observed instances of positive and negative cases – that is, those that had an FTA and those that did not have an FTA. The observed and predicted information is combined to create four categories: True Positive (TP): Correctly classifying someone high risk that has an FTA; True Negative (TN): Correctly classifying someone as low risk that does not have an FTA; False Positive (FP): Incorrectly classifying someone high risk that does not have an FTA; and False Negative (FN): Incorrectly classifying someone as low risk that has an FTA
Before turning to the findings, it is necessary to introduce some key terms. Sensitivity and specificity are the foundation for many performance metrics (e.g., AUC = sensitivity/1-specificity; likelihood ratio tests). Sensitivity is referred to as the true positive rate (TPR) and is the proportion of cases that have an FTA that are classified as high risk (TP/(TP+FN)). Specificity is the proportion of cases that did not have an FTA and are classified as low risk (TN/(TN+FP)). Sensitivity and specificity are the proportion of positive (i.e., FTA) and negative (i.e., no FTA) assessments out of all individuals that have or do not have an FTA, respectively. Sensitivity and specificity are inversely related such that as one increases, the other decrease. The FTA scale has low sensitivity for (29%) and high specificity (87%) for males and females. Sensitivity and specificity of an assessment tell us how good the assessment is for identifying people with or without an FTA when only looking at those with or without an FTA. Sensitivity and specificity tell us nothing about whether or not some people without (negative) or with (positive) an FTA would also test positive or negative, respectively.
Tables 4 and 5 display five sex specific error rates. What stands out from these tables is the similarity in errors rates between males and females. The false negative rates are 71% and 72% for females and males, respectively, and the false positive rates are 12.6%. The false positive and negative rates demonstrate the tradeoffs inherent in risk classification. Given the rare occurrence of FTAs, the PSA has low sensitivity for FTAs (e.g., high false negatives) but high specificity (i.e., good at identifying no FTA, low false positive rate). The FTA scale trades sensitivity (i.e., higher false negatives, lower hit rate) for specificity to prevent false positives. Berk et al. (2021) pointed out that in most criminal justice applications false negative and positive rates are not equal in their importance due to different costs of an outcome (e.g., comparing a homicide to an FTA), and we extend this argument to suggest other errors measures.
PSA FTA Scale Score Error Rates for Females.
PSA FTA Scale Score Error Rates for Males.
We assess the failure prediction and success prediction error rates. Failure prediction error rate is the proportion of cases predicted high risk that do not have an FTA (FP/FP+TP), and the success prediction error rate is the proportion of cases predicted low risk that have an FTA (FN/FN+TN). The failure prediction error rates are 71.7% and 72.6% for females and males, respectively, and the success prediction error rates are 12.4% and 12.6% for females and males, respectively. These error rates further illuminate the importance of the data generating process and the need for a deeper understanding of assessment performance because we need to understand likelihood of a positive (or negative) prediction being wrong (e.g., was the assessment wrong about a high-risk prediction). The false positive rate only tells us the proportion of those with no FTA that have a positive assessment (Tables 4 and 5, row 1), whereas false prediction error rate tells us the proportion of those predicted to have an FTA that did not have an FTA (Tables 4 and 5, column 1). Similarly, success prediction error is the proportion of success predictions (i.e., negative classifications) that are incorrect (Tables 4 and 5, column 2). The failure and success prediction errors assesses use error (Berk et al., 2021) compared to the true classes (positive and negative, respectively) and these metrics are more direct measures of concern for bias because they provide the probability of error when making positive (failure prediction error) or negative (success prediction error) predictions. Simply, what is the probability of a high (or low) risk assessment being wrong (i.e., someone does not have an FTA)?
The four error measures show two things. First, the FTA scale has similar error rates for males and females. Second, identifying positive cases for FTAs is difficult given the nature of its rare occurrence in the sample. 9 The last error measure provides an assessment of the overall error for negative and positive classifications ((FN+FP)/N) and demonstrates that about 21% of the classifications using the FTA scale are incorrect. The overall prediction error is the inverse of accuracy ((TP+TN)/N) which demonstrates that about 79% of the classifications were correct – mostly because it is less difficult to identify those that do not have an FTA (i.e., low base rate). Overall error rates (and accuracy) do not vary by sex.
Discussion
This paper is the first assessment of differential validity and prediction by sex of one pretrial risk assessment (i.e., PSA) using data from a statewide pretrial agency. The PSA has become one of the most popular actuarial pretrial assessments and is used in dozens of jurisdictions and involved in thousands of pretrial decisions daily. The susceptibility of humans to make systematic errors in judgment necessitates the need to study the use of actuarial tools for pretrial release (Viljoen et al., 2019). Decision makers rely on heuristics such that irrelevant information (i.e., anchoring), presentation style (i.e., framing), misremembering the past (i.e., hindsight bias), poor understanding of statistical relationships (i.e., representativeness heuristic), and overconfidence (i.e., egocentric bias) become the driving mechanisms for important decisions (Guthrie et al., 2007). Behavioral economists have routinely shown that even high-stakes decisions are predicated on intuitive decision-making faculties (Thaler, 2016).
The current study demonstrates (at least in Kentucky) that rates of three negative outcomes are low and similar across sexes, but males have significantly more involvement in violent crimes than females. However, arrests for violent crimes during pretrial were exceptionally rare. The PSA achieves validity measures that meet accuracy standards identified for criminal legal system instruments and there are no differences in AUC scores for two of the three scales between males and females. The AUCs for the NCA scale is.016 higher for males than females, which is significant and means that the PSA is less accurate at classifying females for a new crime, but this is a small difference.
We tested calibration estimating four regression models per outcome to estimate the main effects for the PSA score and sex, and to assess whether sex moderates any association between the PSA score and the outcomes. Table 3 shows there is little effect on FTAs or NCAs by sex and significant effects between sex and NVCAs (models 1 and 3). Concerns of new arrests for violence during pretrial likely influence judicial decisions as judges are intuitively working to identify individuals that they think are more likely to commit a violent act if released. Arrests for new violent crimes are extremely rare, and they do not warrant being a central concern for pretrial release, which is not to say that probability of a violent arrest during pretrial is unimportant.
We assessed equality of five error rates for FTAs because the FTA base rates are the same by sex. The analyses did not reveal differences between males and females in the PSA’s error rates. The contingency tables, hopefully, demonstrate the challenges to validating risk classification instruments for rare events (i.e., few positives). The contingency tables demonstrate that about 75% of the male and female samples were identified as true negatives, whereas about 4% were identified as true positives. This distribution contrasts with the ProPublica (Angwin et al., 2016) data including 37% and 21% true positives and 27% and 46% true negatives for black and white defendants, respectively. The ProPublica study is not a true study of pretrial assessment performance as they used post-conviction recidivism, not pretrial outcomes. We provide these comparisons because the ProPublica article served as a lightning rod for the assessment controversy, and we hope to highlight the inadequacy of this study for pretrial assessments. It is crucial to consider the underlying distributions within the data being used to develop, validate, or test for predictive bias.
Conclusion
In closing, we suggest five areas of future research and policy development. First, we suggest that criminologists should recognize that the pretrial phase is different from post-conviction periods. Several previous studies relied on 2- and 3-year recidivism patterns to assess the effectiveness of pretrial assessments. The pretrial period is much shorter and necessitates a deeper understanding of the dynamics happening between arrest and case resolution, during which people are considered legally innocent (Scott-Hayward & Fradella, 2019). Clearly, pretrial base rates are much lower than sentenced populations with only 10% of the sample arrested for a new crime and 1% arrested for a new violent crime (DeMichele et al., 2020), and the longer term follow up studies (e.g., Angwin et al., 2016) do not even consider FTAs. Moreover, the bulk of the legal scholarship critiquing criminal legal system assessments focus on sentencing concerns (e.g., Harcourt, 2015; Starr, 2014), which also overlooks the unique characteristics of pretrial decisions.
Second, a limitation of this study (and pretrial studies in general) is that we only analyzed the individuals who are released. In Kentucky, nearly 70% of bookings are released pretrial, but in other jurisdictions this proportion could be much lower such that pretrial assessments are designed on a smaller proportion of the population. Kleinberg et al. (2018) demonstrated a promising simulation approach to estimating failures for the detained population using information (about the outcomes) from those released and studying judge specific release practices. Their approach emphasizes some of the tradeoffs that judges make related to detention and success rates. Simply, in their simulation, Kleinberg et al. (2018) found that if judges used an assessment there would be a 42% reduction in jail populations without changing rates of negative outcomes; alternatively, detention rates could be held constant and rates of negative outcomes could be reduced by 25% by releasing individuals following the assessment recommendations.
Third, pretrial assessment development and validation needs to be an ongoing process that is incorporated with other changes to pretrial practices. Koepke and Robinson (2018) demonstrated the impact of developing the Colorado Pretrial Assessment Tool using 2012 data that produced large errors when it was implemented because Colorado had made several changes to their pretrial system. These changes were effective at mitigating individual risk and hence increased success rates and essentially shifted the distribution of risk such that the CPAT was making more errors than when validated. Criminologists need to work with agencies to develop data collection procedures that enable ongoing validation and improvements to classifications that map onto the reality of pretrial dynamics. Moreover, assessments are only one tool among several that stakeholders may consider to improve the treatment of people during pretrial. Regardless of implementation decisions, it is essential that stakeholders consider their socioeconomic context so they do not just push disadvantage farther into the system
Fourth, the criminological field needs standards. The AERA/APA developed test standards over 50 years ago, but the criminological field lacks standards for release assessment development, use, validation, and bias testing. The lack of standards is surprising since assessments have been used since 1928, when Burgess implemented a release checklist in Illinois (Burgess, 1928). The lack of standards for actuarial tools creates a scenario in which for profit companies can create opaque tools with little consideration of liberty, constitutional protections, and fairness. We encourage the criminological community to consider the National Association of Pretrial Services (NAPSA, 2020) standards for actuarial pretrial assessments as a starting place to for researchers and practitioners to consider when developing, validating, and using assessments.
Fifth, more research is needed on the differences (or lack thereof) between males and females in pretrial outcomes. Our findings contribute to the scant research on predictive bias by sex by showing that males and females have similar FTA and new arrests rates. Although new arrest base rates were significantly different (hence, we did not compare error rates), these rates differed by less than 1.5% (11.1% vs. 9.7%). The predicted probabilities (Figure 1) demonstrate similar recidivism patters. Given evidence of overpredictions in other studies, we suggest taking these results with some caution as additional research findings accumulate. We also underscore the need for the inclusion of gender categories that represent people who do not identify as male or female, and the ability of all people processed by the criminal legal system to select their own gender identity and not be assigned to a category based on biology. The recommendation for research on differential prediction by sex is not meant to displace the need for research on race. Rather, we would go further to suggest that researchers should start to look at the intersectionality of race and sex with further research.
Legal system decisions are human endeavors in which people are making decisions about other people. Individuals that pass through legal systems typically are among the most vulnerable in society and these individuals need protections such that judicial decisions that are informed by actuarial assessments cannot ignore issues of liberty, fairness, and inherent rights to freedom. Moreover, assessments are one tool among many that stakeholders may consider to avoid unnecessary pretrial detention. The groundwork required to thoughtfully adopt and implement an actuarial assessment may help system actors identify other ways the pretrial process could be changed to decrease custodial arrests and increase connection to services, such mental health and substance use treatment, as well as court date reminders, transportation assistance, options for virtual hearings, and other means of facilitating court appearances. Given the clear links between structural disadvantage and criminal legal system involvement, all policies and interventions – including pretrial assessments – must be designed and implemented in ways that aim to strengthen equity.
Footnotes
Authors’ Note
Peter Baumgartner is also affiliated to Machine Learning Engineer, Explosion AI, Durham, NC, USA.
Michael Wenger is also affiliated to Data Scientist, Division of Statistical and Data Science, RTI International, Research Triangle Park, NC, USA.
Megan Comfort is also affiliated to Senior Fellow, Transformative Research Unit for Equity, RTI International, San Francisco, CA, USA.
Amanda Witwer is also affiliated to Graduate Student, School of Criminal Justice, Michigan State University, East Lansing, MI, USA.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This study was financially supported by The Laura and John Arnold Foundation.
