Abstract
Introduction
Young adults aged 18–24 comprised approximately 7% (n = 47,436) of the unhoused population identified in the United States of America’s 2023 annual point-in-time count (US Department of Housing and Urban Development, 2023). Homelessness arises from the complex interplay of individual, relational, socioeconomic, and environmental factors. Homelessness in the United States is defined as lacking access to a fixed, regular, and adequate nighttime residence (McKinney-Vento Homeless Assistance Act, 1987). Unhoused people may be living in places not meant for human habitation, in shelters or other temporary living arrangement such as hotels subsidized by governmental programs or transitional housing, or are at immanent risk of losing their housing. Unhoused young adults can experience overlapping forms of marginalization and structural disadvantage that increase vulnerabilities to substance misuse. Compared to adolescents and older adults experiencing homelessness, younger adults report shorter duration of homelessness but have reported higher rates of recent life stressors related to social relationships, mental and physical health, education/job training, and employment (Sample & Ferguson, 2020; Tompsett et al., 2009). These stressors are often coupled with the challenges of age-appropriate milestones, such as identity development, navigating friendships and relationships, and development of autonomy (Munson et al., 2017). Initiation of substance use and risk of substance use disorder (SUD) increases in early adulthood, although young adults display heterogeneity in their substance use trajectories (Chen & Jacobson, 2012; Maggs et al., 2023).
Substance misuse, which has biological, personal, relational, and environmental antecedents, has been cited by some unhoused people as a contributing factor for either entering into or finding it difficult to exit homelessness (Barile et al., 2020; O’Toole et al., 2004; Sample & Ferguson, 2020). Substance misuse can contribute to homelessness by disrupting engagement in occupational tasks such as work or school/training programs that can impact persons’ financial resources and create risks for housing instability (Stablein et al., 2021). In addition, substance misuse can disrupt relationships with peers, family, or community members and decrease social, emotional, and instrumental supports that might protect against housing instability (Stablein et al., 2021).
Young adults in addiction treatment have been found to have more severe risk profiles than older adults that could contribute to poorer addiction treatment outcomes in the absence of adequate psychosocial and clinical supports (Andersson et al., 2021). Although unhoused youth aged 13–26 who used injected drugs were found to have higher odds of utilizing mental health services, prior research with unhoused youth found an inverse association between substance use and help-seeking behaviors (Crosby et al., 2018; Solorio et al., 2006). Addressing substance misuse among unhoused young adults might facilitate their engagement in supportive services. In addition, supporting substance use recovery may promote retention in housing services, as both substance use and drug-related activity can contribute to eviction from housing services (Cole, 2024). Early adulthood is a crucial time to change substance use patterns to promote more positive trajectories and mitigate risks of violence, accidental injury, criminal justice system involvement, poor physical and mental health outcomes, and premature mortality. Understanding the predictors of treatment completion among unhoused young adults may aid in the development of strategies to facilitate successful engagement into substance use treatment services.
Insight into behavioral health, service-related, and demographic features predictive of substance use treatment completion in unhoused young adults may inform strategies to promote positive health and social outcomes for this vulnerable group. Improvement in substance use behaviors and treatment completion have been shown to be positively associated, although clinical improvement and treatment completion are theorized to be two distinct constructs (Sahker et al., 2022). Prior research indicates that treatment participation and completion is associated with more positive quality of life outcomes (Gonzales et al., 2009). Outcomes for those who do not complete treatment are heterogeneous and treatment dropout is not always associated with poor social or clinical outcomes (Szafranski et al., 2019). Nevertheless, treatment services, particularly those delivered by caring, recovery-oriented service providers and peers, have been cited by people who use drugs as playing an important role in their own recoveries (Dell et al., 2022a; Schoenberger et al., 2022).
Publicly available administrative data in the United States, such as the Treatment Episode Data Set-Discharges (TEDS-D), has often been utilized to understand factors associated with treatment completion or discontinuation for diverse subpopulations receiving addiction treatment. Such studies have identified disparities in treatment completion by people of minoritized racial/ethnic backgrounds, psychiatric comorbidities, insurance status and other socioeconomic factors (Friesen & Kurdyak, 2020; Goldman et al., 2020; Krawczyk et al., 2017; Suntai, 2021). Increasingly, researchers have sought to develop and evaluate predictive models using publicly available, routinely conducted administrative data to predict treatment completion or non-completion among different subpopulations of people receiving substance use treatment services (Acion et al., 2017; Baird et al., 2022; Kong et al., 2022; Nasir et al., 2021; Shikalgar et al., 2024; Stafford et al., 2022). Less research has focused on predictive models and their relevance for people experiencing homelessness. Using machine learning (ML) approaches to predict treatment outcomes from routinely collected baseline data may provide clinicians and program managers with insights into both population-level and individual-level needs to improve service outcomes.
The use of ML tools in psychiatry and other behavioral health disciplines has been characterized as a “paradigm shift” from classical hypothesis testing to evaluating the predictive performance of models among unseen or untested data to inform clinical treatment (Chekroud et al., 2021). An advantage of ML approaches is their ability to handle high dimensional data and non-linear relationships. Although ML approaches are becoming more popular in addiction medicine, additional research is needed to explore the generalizability of ML models to diverse populations and treatment settings (Mak et al., 2019). The development and evaluation of ML models to understand patterns in treatment completion among large real-world datasets may inform theory-driven research and specific clinical practice to improve treatment completion for persons at risk of non-completion.
Aims
The present study is novel in its use of ML methods to study treatment completion among unhoused young adults. Our first aim was to compare the performance of two predictive models using data collected upon enrollment into services to classify treatment completion for young adults experiencing homelessness enrolled in substance use treatment services. Our second aim was to identify features most important for accurate prediction in the best performing predictive models. These aims assess the extent to which routinely collected administrative data can be leveraged to make accurate predictions about treatment completion. This study provides insight into the features that are most important for accurately assessing the probability of treatment completion for young adults. If models are successful in using administrative data to accurately predict treatment completion, such models may be scalable across different service settings and could provide actionable insight for equitably allocating resources and services to unhoused young adults who experience lower risk of completing treatment.
Materials and Methods
Data and Sample
We used publicly available data from the 2020 TEDS-D (Substance Abuse and Mental Health Services Administration [SAMHSA], 2020). Data are collected for all 50 states, Washington, D.C., and Puerto Rico, although Idaho, Maryland, New Mexico, Oregon, Utah, and West Virginia were missing data in 2020. The TEDS-D is a repository of routinely collected, publicly funded treatment data collected primarily for administrative purposes. TEDS-D documents the demographic and clinical characteristics of discharges from publicly funded substance use treatment services in the United States of America. TEDS-D data primarily documents individual-level characteristics. Potentially meaningful data about treatment dose and service quality, as well as many social factors that influence health status, are not collected. The TEDS-D does include both baseline and discharge-level data for certain characteristics, such as living arrangement, employment status, and substance use behaviors. Although both baseline and discharge characteristics are contained in TEDS-D, we selected only characteristics collected at baseline, as we wanted to assess the extent to which data collected at intake could be predictive of treatment completion. The sample included treatment discharges for unhoused adults aged 18–24 years whose living arrangements were coded as ‘homeless’ (N = 12,273).
Target
The target variable, treatment completion, showed mild class imbalance (completed: 32%, n = 3988; non-completed: 68%, n = 8285). In the TEDS-D, treatment completion is met if the service setting documents that “all parts of the treatment plan or program were completed,” whereas non-completion, for this analysis, included those who were coded as having dropped out of treatment, having been terminated by the facility, having transferred to a facility, were incarcerated, died, or experienced some other reason for treatment non-completion (SAMHSA, 2020).
Features
The majority of features are dichotomous, indicating the presence or absence of a condition and include: whether the client had any prior treatment episodes, employment, substance use as a minor, co-occurring psychiatric illness, past 30-day arrests, male sex, insurance (non, public, private/commercial); race (non-Hispanic White, non-Hispanic Black, non-Hispanic other, or Hispanic), referral source (self, criminal justice, community), education level (less than high school, high school or equivalency, greater than high school), and service setting (detoxification, ambulatory, residential).
The following substance use variables indicated whether the treatment episode documented any primary, secondary, or tertiary use at admission: alcohol, cocaine, cannabis, heroin, other synthetic opioids, hallucinogens, methamphetamine, other amphetamines, stimulants, benzodiazepines, or other substances. The ‘other substances’ variable combined several drug categories labeled in the TEDS-D that had low cell counts in the sample and included barbiturates, phencyclidine (PCP), sedatives, over the counter medications, over the counter drugs, including non-prescription methadone, inhalants, and other substances. As the TEDS-D documents only the first three substances that led to the treatment episode, all substances used by the client may not be enumerated by the TEDS-D. The absence of a particular substance does not necessarily imply that the substance was not used by the client.
Data Preparation and Analysis
Random forest (RF) and penalized logistic regression using elastic net (LR) were conducted in R (R Core Team, 2022). First, we assessed missing data in the dataset. Missing data were rare for most features, with only three of 36 features having missing data greater than 5.0%, including whether the patient had a prior treatment episode (5.8%, n = 712), whether a co-occurring psychiatric condition was documented (13.5%, n = 1654), and the patient’s insurance status (50.3%, n = 6168). Although the proportion of missing data were large for insurance status, prior research suggests that the proportion of missing data does not adequately inform whether to use multiple imputation methods or complete case analysis (Madley-Dowd et al., 2019). Therefore, we selected the missForest package in R to impute missing data, which is an imputation method based on RF (Stekhoven & Bühlmann, 2012).
After imputing missing data, we one-hot encoded features. One-hot encoding is a method of transforming categorical features with multiple levels to separate binary features. For example, the “referral source” feature had three levels: self-referred (1), criminal justice referral (2), or community-based referral (3). Each of these levels becomes separate dichotomous features when transformed. To reduce redundancy, k-1 features were created. For example, the feature “methamphetamine.No” is redundant with the feature “methamphetamine.Yes,” so “methamphetamine.No” was dropped from the dataset. Next, data were partitioned into training (80%; completers = 3250/9902) and test (20%; completers = 738/2371) datasets. Models were initially estimated on the slightly imbalanced training data. However, we wanted to understand whether models on a balanced training set performed similarly to models estimated on the imbalanced training set. Therefore, the ROSE (Random Over-Sampling Examples) package in R was used to balance the training dataset. The ROSE function is a bootstrap-based technique that generates synthetic samples to create a balanced dataset (Lunardon et al., 2014). We then re-ran each analysis on the balanced data.
RF was conducted using the randomForest package (Liaw & Wiener, 2002). RF is a tree-based method that can be used for either regression or classification tasks. RF is an ensemble method that involves aggregating a collection of uncorrelated classification trees (Genuer & Poggi, 2020; Tibshirani et al., 2017). The “random” of RF operates on two levels. First, a prespecified number of classification trees are grown from random bootstrapped training samples. Next, at each node, only a random subset of predictors is chosen to partition the feature space. RF can be tuned by specifying the number of trees grown (ntree) and by specifying the number of features that are included as split candidate at each node (mtry). For classification tasks, the default mtry value is sqrt(x), where x represents the number of features in the model. We searched for optimal mtry by different prespecified numbers of trees. For the RF model trained on imbalanced data, mtry was set to 5 and ntree set to 300, whereas the RF model trained on balanced data, mtry was set to 5 and ntree set to 500.
Penalized (or regularized) LR models were conducted using the caret package with the “glmnet” (or generalized linear models via penalized maximum likelihood) method (Kuhn, 2008). This method is similar to the elastic net and balances the regularization of ridge and LASSO (least absolute shrinkage and selection operator) penalization methods, both of which aim to reduce the complexity of the model. Ridge regression shrinks parameter estimates towards 0 based on a penalty term but retains all predictors in the model, whereas LASSO regression sets some parameters at 0 based on a penalty term and retains only the most significant contributors to the model. In glmnet, an alpha value of 1 corresponds to a pure LASSO penalty. An alpha value of 0 corresponds to a ridge penalty. The function selects the amount of penalization (lambda), although analysts can specify their own regularization parameter. For the model trained on imbalanced data, the alpha was set to 1 and lambda equaled 0.0021. For the model trained on balanced data, alpha was set to 0.11 and lambda equaled 0.0002.
Model performance was assessed by inspecting the following evaluative metrics: Area Under the Receiver Operating Characteristic Curve (AUC) and 95% confidence interval (95%CI), which is a measure of how well the model discriminates between classes. Accuracy refers to the proportion of total cases accurately predicted relative to the total sample. Accuracy can be high in models with class imbalance but result in models with low precision. Specificity refers to model’s ability to accurately distinguish true negatives, whereas recall (sensitivity) refers to the model’s ability to accurately predict positives cases. Recall allows us to know how many relevant cases are identified by the model. Precision, or positive predictive value, is the proportion of true positives relative to the sum of true positives and predicted false positives. Precision identifies how many cases identified by the model are relevant. F1 is the harmonic mean of precision and recall, whereas balanced accuracy is the average of recall and specificity.
Feature importance was inspected for the optimal RF and LR models. For RF, mean decrease accuracy (MDA) was calculated for each feature and inspected to identify who much the accuracy of the model would be impacted if each feature were excluded. Higher MDA scores indicate that a particular feature is relatively more important for accurately classifying cases compared to other features in the model. For the optimal LR model, the varImp function in caret computes variable importance for linear models using the absolute value of the t-test statistic. In addition, we computed partial dependence profiles using the DALEX package to view the marginal effect of each predictor on treatment completion for the five predictors with the highest importance scores for each model (Biecek, 2018).
Results
Sample Characteristics
Supplementary Table 1 presents the demographic and clinical characteristics of unhoused young adults by treatment completion status. Males comprised 62.69% of treatment completers and 56.74% of non-completers. The sample was mostly non-Hispanic white (completers: 59.40%; non-completers: 57.62%), followed by Hispanic (completers: 18.63%; non-completers: 18.87%), non-Hispanic black (completers: 11.51%; non-completers: 13.41%), and non-Hispanic other (completers: 8.43%; non-completers: 13.41%). Most patients had at least a high school education. Insurance status was missing for 44.78% of completed episodes of treatment and 52.89% of episodes that were not successfully completed. Similar rates of no and public insurance were found by completion status, whereas completers had higher rates of private insurance (9.68%) relative to non-completers (3.90%). Only 12.76% of completers and 10.31% of non-completers were employed at baseline. For each completion status, approximately 10% of episodes documented arrest in the 30 days prior to beginning treatment. Mental health problems were documented in 37.89% of completed treatment episodes and 45.07% of episodes that were not completed. For completed episodes, the most frequently reported substances documented included alcohol (39.49%), cannabis (36.94%), methamphetamine (36.28%), heroin (32.82%) and cocaine/crack (18.48%). For episodes ending in treatment non-completion, the most frequently reported substances documented included methamphetamine (46.80%), cannabis (46.08%), heroin (31.29%), alcohol (30.97%), and cocaine/crack (14.44%).
Most treatment episodes were initiated through self-referral (completers: 43.05%; non-completers: 44.61%), a community-based referral (completers: 33.70%; non-completers: 33.51%), or a criminal justice referral (completers: 22.62%; non-completers: 20.57%). Over half of treatment completers had at least one prior treatment episode (52.53%), whereas 45.81% of non-completers had prior treatment. Nearly 60% of completers and non-completers reported initiation of substance use as a minor. One-third of treatment completers received detox services, 39.94% residential services, and 26.86% ambulatory services, whereas 14.79% of non-completers received detox services, 33.94% residential, and 51.27% ambulatory. Over half of treatment episodes lasted one month or less (completers: 58.43%; non-completers: 61.16%), followed by 1–3 months (completers: 20.89%; non-completers: 20.93%), and 4 or more months (completers: 20.69%; non-completers: 17.91%)
Aim 1: Model Performance
Performance Metrics for Models Predicting Treatment Completion for Unhoused Young Adults in Substance Use Treatment Services Using Random Forest Imputed Dataset (N = 12,273).
AUC = Area under the receiver operating characteristic curve; F1 = harmonic mean of precision and recall; 95%CI = 95% confidence interval; ntree = number of trees in each random forest; mtry = number of features tried at each split; α = elastic net mixing parameter; λ = regularization parameter; bolded metrics indicate highest value relative to other models.
Aim 2: Feature Importance
Variable Importance Plots
Figure 1 shows variable importance scores for both RF and LR trained on balanced data. The top two most important features for each model were ranked in the same order: receiving ambulatory treatment services and having private insurance. The next most important features for the RF model included first use of substances as a minor, methamphetamine use, and having any prior treatment. The next most important features for accurate classification in the LR model included other stimulant use, being referred from a criminal justice setting, and receiving residential services. Several top features were ranked similarly across models, such as receiving ambulatory services, having private insurance, and methamphetamine use, criminal justice referral, and prior treatment. However, some features were inconsistently ranked across models; for example, first use as a minor was ranked as the third most important feature in the RF model but was ranked second to last in the LR model. Variable Importance Plots for Models Trained on Balanced Data.
Partial Dependence Profiles
To understand the marginal effect of each of the top predictors on treatment completion, partial dependence profiles are presented in Figure 2. As indicated in Figure 1, the top two features were the same for the optimal RF and LR model. The PDPs visualized in Figure 2 show that each of these top three features had a similar predicted marginal effect on treatment completion. For the RF model, receiving ambulatory services was associated with 0.32 probability of treatment completion, relative to 0.59 probability of completing treatment if not receiving ambulatory services. For the LR model, receiving ambulatory services was associated with 0.29 probability of completing treatment, relative to 0.65 probability of completing treatment if not receiving ambulatory services. Similarly, probability of treatment completion was comparable across RF (0.62) and LR (0.68) models. Partial Dependence Profiles: Top 5 Features for Each Model.
Discussion
Our study contributes to the growing use of supervised ML approaches to advance behavioral health science that promotes positive mental health and substance use outcomes among young adults (Dell et al., 2022b; Han & Seo, 2022; Kundu et al., 2022; Rakovski et al., 2023). We found that models trained on routinely collected administrative data can yield moderately accurate predictions of treatment completion in unhoused young adults. We envision that the models developed in this study could be further refined and potentially integrated into substance use treatment services that already collect TEDS-D data elements, particularly those programs serving unhoused young adults. Models may be useful for facilitating individualized treatment planning between clinicians and clients or for administrators in developing programming to meet the needs of unhoused young adults. However, further evaluation is needed before findings can be implemented and scaled to individual settings, as it is unknown whether the models, trained on national data, might yield meaningful predictions for individual sites with smaller numbers of clients. In the present study, findings reveal multiple features important in the accurate classification of treatment completion, which may be useful for the development of individualized interventions to support clients’ engagement and retention in treatment. These two main findings are discussed in more depth below.
Recent research has found that RF has outperformed LR in classifying different clinical conditions (Dell et al., 2022b, 2022c). However, not one method will outperform under all conditions, as in the present study the penalized LR using elastic net and RF yielded similar results when both were trained on balanced data. In addition, the evaluation of each approach requires interpretation of multiple metrics, as no single metric can fully capture the performance of a model. Although models trained on balanced data had slightly lower overall accuracy, AUC, precision and specificity, higher recall may be prioritized over specificity when the goal is to identify a rare outcome (Byrne et al., 2019). Overall, our models showed similar AUC as other modeling approaches utilizing similar features from TEDS data on distinct populations in SUD treatment. Although a few other studies (discussed below) have used a wider array of modeling techniques, we evaluated our models with a comprehensive array of performance metrics that facilitate assessment of the tradeoffs between models (Acion et al., 2017; Baird et al., 2022;Stafford et al., 2022).
An additional consideration when selecting between models for use in any setting is how explainable each model is for different stakeholders who directly use or are impacted by the predictions generated by the model. Inspecting feature importances can provide insight into how each model operates and can influence stakeholders’ levels of trust in the model’s predictions and ultimately how or whether the model is used to guide practice. As was evident in the present study, several features were ranked of similar importance across the RF and LR models considered and these consistently important features may be considered as reflecting “genuine aspects of the data considered” (Saarela & Jauhiainen, 2021, p. 9). Similar levels of accuracy were found between LR and RF models despite having some inconsistencies in feature importance scores across models. Selection of an optimal model in practice settings requires consideration of both the technical performance of the algorithm to classify the outcome of interest and the degree to which stakeholders trust the model to advance the service setting’s goals, cohere with the setting’s people, culture, and processes, and the degree to which the use of a predictive algorithm can be supported by the agency’s technology and infrastructure.
As detailed below, our findings complement and build upon other research using predictive modeling approaches to classify treatment completion or non-completion for different populations receiving SUD treatment. Acion et al. (2017) sought to predict treatment completion among Hispanic adults in outpatient treatment who had no prior treatment episodes using the 2006-2011 TEDS-D. The range of AUC (0.793–0.820) found by Acion et al. (2017) was slightly higher than what we found (0.7234–0.7753) and might be attributable to differences in substance use and demographic patterns between each sample and that our study included different treatment settings rather than focusing solely on outpatient settings. We did not rely solely on AUC and included metrics such as precision, recall, and F1 score to assess how well each predictive model was at correctly identifying relevant cases rather than looking at the model’s overall ability to discriminate between classes. In addition, Stafford et al. (2022) generated several predictive models to identify reasons for OUD treatment discontinuation using more recent 2015-2019 TEDS-D data. The highest performing RF model generated by Safford et al. showed comparable accuracy (0.69) and AUC (0.73) relative to our models. In Stafford et al. (2022), the feature most important for accurate classification of treatment discontinuation among opioid use disorder (OUD) treatment episodes (service setting) was ranked highly in both the RF and penalized LR models considered in this study. Finally, prior analyses using decision trees to identify inequities in SUD treatment completion have found similar patterns, namely the absence of a mental health condition, having income, and being non-Hispanic white were associated with higher probabilities of treatment completion (Baird et al., 2022).
The lower probability of treatment completion in ambulatory settings found in our study is consistent with other studies using TEDS data to study other populations, such as those who use opioids and pregnant women receiving treatment for cannabis use (Kitsantas et al., 2023; Stahler & Mennis, 2018). Several implications could follow from this finding. On the one hand, we may question whether ambulatory settings are the most appropriate settings for unhoused young adults to engage in substance misuse treatment, as lacking secure and stable housing can create immense barriers to accessing treatment regularly. However, a limitation of the current study is that what constitutes treatment completion likely varies across detoxification, ambulatory, and residential services. Furthermore, TEDS-D does not contain data on the complete array of services offered during each ambulatory episode and whether harm reduction or abstinence-only approaches are being implemented. Person-centered care emphasizes client choice and requires flexibility in the types of services offered and how they are delivered. As most services are offered in ambulatory settings, strategies to retain unhoused young adults into are needed, and could include an increase in the availability of housing first services, for example, and intensive outreach approaches tailored to those clients who use methamphetamine and experience treatment barriers due to lack of commercial insurance.
Having no private insurance and receiving outpatient services was predictive of lower probability of treatment completion. This could be interpreted in light of the difficulties the unhoused population can face when engaging in ambulatory treatment services. Prior literature has found that unhoused clients’ previous negative experiences of care, including provider stigma or difficulty navigating complex administrative processes or keeping appointments, self-stigma, or meeting food or shelter needs taking priority over attending clinic services create substantial barriers to accessing and engaging in treatment services (Miler et al., 2021). Research on residential and outpatient settings have found mixed results for the relationship between prior treatment episodes and treatment completion (Baker et al., 2020; Cacciola et al., 2009; Mutter et al., 2015). Some evidence suggests that persons with multiple treatment episodes may have greater substance use severity and psychiatric comorbidities (Rash et al., 2008; Villagrana & Lee, 2020).
Use of methamphetamine (for each model), use of heroin (in the RF model), and use of other stimulants (in the LR model) were important substance use-related features for accurately predicting treatment completion. Co-occurring methamphetamine and opioid, particularly fentanyl, use is associated with greater risk of overdose and death (Han et al., 2021; Mattson et al., 2021). Although the unhoused population has higher risk of opioid and methamphetamine use disorder, they are less likely to engage and be retained in care, although low-threshold programs are associated with higher retention in SUD treatment services (Gaeta Gazzola et al., 2022; Wakeman et al., 2022). In addition, as psychiatric disorders are common among people experiencing homelessness relative to the general population, low-threshold programs that integrate mental health and substance use treatment may be advantageous to the health of unhoused people, especially as comorbid psychiatric conditions are risk factors for treatment non-adherence and non-completion (Krawczyk et al., 2017).
Although these features are associated with lower probability of completing treatment, several evidence-based and evidence-informed interventions can be delivered in outpatient settings that specifically address methamphetamine use, opioid use, and psychiatric comorbidity among young adults. Housing First models, which provide immediate access to housing and rehabilitative services, are considered crucial to address and treat SUD and co-occurring and mental health issues among unhoused young adults (Slesnick et al., 2023). Motivational interviewing (MI) is an evidence-based, person-centered practice that strengthens motivation for behavior change (Miller & Rollnick, 2023). MI strategies with unhoused young adults may promote both positive substance use behaviors as well as sexual health behaviors (Tucker et al., 2020, 2021). Next, contingency management interventions, which are informed by operant conditioning to reward and reinforce healthy substance use behaviors, have shown effectiveness for unhoused men who have sex with men with methamphetamine use disorders and for persons receiving medication for opioid use disorders (MOUD) (Bolivar et al., 2021; Brown & DeFulio, 2020; Reback et al., 2010). For psychiatric comorbidity, Assertive Community Treatment leverages the skills of a multidisciplinary team to assertively engage and support persons with comorbid serious mental illness and substance use disorders who have high rates of hospitalization, criminal justice system-involvement, or housing instability (Morse et al., 2006; Penzenstadler et al., 2019). In addition, although employment was not a highly important feature for classifying treatment completion, qualitative evidence from unhoused young adults suggests the importance of employment and educational opportunities as a potential pathway from homelessness (Sample & Ferguson, 2020). To this end, integrating supported employment and supported education into housing and behavioral health interventions may be supportive of recovery (Ferguson et al., 2012; Hofstra et al., 2023).
Limitations
First, the unit of analysis was the treatment episode, rather than the unique patient. Identification of unique patients is not possible in the publicly available TEDS-D. TEDS-D also relies on a simplistic operationalization of homelessness that does not detail the type (sheltered vs. unsheltered) or duration of homelessness and lacks nuanced categories to document gender, a full enumeration of problematic substances, and important details about the type, quality, and intensity of services received across treatment settings. Models generated using linked records on individual patients over time may yield different results from models estimated on cross-sectional treatment episode data. Furthermore, the data reflect only publicly funded treatment admissions might not be representative of treatment completion among unhoused young adults with SUD receiving services in other settings. The true population of unhoused young adults is unknown, so the sample may not be representative as those receiving treatment services might differ in important ways from those not connected to publicly funded treatment, or who are receiving non-publicly funded treatment. Given the importance of service setting in the penalized LR model, specific predictive models for unique service settings, particularly ambulatory services, are warranted. In addition, a fuller, more theoretically informed selection of features could yield more accurate predictive models. However, the purpose of this study was to evaluate the predictive performance of models using routinely collected administrative data and to identify overarching patterns.
Conclusions
The present study suggests that treatment completion for unhoused young adults can be adequately predicted using a routinely collected, although limited range of individual-level and treatment setting-related factors. Considering that each model performed adequately across different evaluative metrics, further research is warranted to evaluate the feasibility of translating models using routinely collected baseline administrative data into effective tools to support treatment completion among unhoused young adults with SUD. This aim could be further evaluated using transfer learning approaches to maximize the generalizability of models to settings serving unique or smaller patient populations (Bailey & DeFulio, 2022). In sum, this study complements and extends prior research on the feasibility of developing predictive algorithms by investigating treatment completion among unhoused young adults. Data science approaches can aid researchers, program administrators, and clinicians in making use of existing administrative data to inform service delivery. Models may be useful for tailoring clinical programming to support retention in services for specific patients and patient subgroups but additional research is needed to translate these findings and evaluate their impact in clinical practice.
Supplemental Material
Supplemental Material - Substance Use Treatment Completion Among Unhoused Young Adults: A Predictive Modeling Approach
Supplemental Material for Substance Use Treatment Completion Among Unhoused Young Adults: A Predictive Modeling Approach by Nathaniel A. Dell, Charvonne Long, Christopher P. Salas-Wright, Michael G. Vaughn, Hannah S. Szlyk, Patricia Cavazos-Rehg in Journal of Drug Issues
Footnotes
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This work was supported by theNational Institute on Drug Abuse (K01DA058750) and Center for Substance Abuse Treatment (H179TI081697).
Supplemental Material
Supplemental material for this article is available online.
Author Biographies
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
