Abstract
Can mental health apps be a tool in tackling the declining mental health in the United States and around the globe? Our review of the literature shows that mental health apps can improve symptoms of some disorders, such as depression and anxiety, when compared to no treatment. Such evidence of efficacy, however, is lacking in populations that are underserved by existing mental health services, including youth, seniors, men, and racial minorities. Additionally, a large number of apps claim to be evidence-based without any publicly available peer-reviewed evidence. This unregulated marketplace risks misleading consumers and exacerbating health disparities. We propose multi-level policy solutions, including oversight by the Food and Drug Administration (FDA), voluntary adoption of scientific disclosure standards by app stores, and the creation of accessible, AI-enhanced databases to help consumers identify validated tools. Mental health apps promise to expand access to care, but without stronger evidence and regulation, they may deepen—rather than bridge—the mental health treatment divide.
Social Media
Mental health apps promise accessible care—but most lack scientific validation. Our review finds that while some apps help with depression & anxiety, the majority make unverified claims. Regulation & transparency are urgently needed to protect consumers. #MentalHealth #Science
Key Points
Mental health apps offer scalable, low-cost tools that can reduce symptoms of depression and anxiety—especially when paired with professional guidance.
Most apps have not been tested using rigorous scientific methods, and few show real-world effectiveness compared to standard care.
Research rarely includes youth, seniors, men, or racial minorities—the very groups that apps claim to help most.
Of 1,244 mental health apps analyzed, about one-third made evidence-based claims, but only 27% of those making claims had any peer-reviewed, publicly available support.
Stronger FDA oversight, app-store accountability, and consumer tools are needed to help users identify validated, safe, and effective apps.
Mental health apps can democratize care—but without rigorous testing and regulation, they risk widening, not narrowing, mental health inequalities.
Sophia, a 27-year-old marketing professional in New York City, enjoys her fast-paced job, weekend dinners with friends, and quiet nights at home with her cat, Muffin. Lately, though, she has felt a constant edge of anxiety she can’t shake. A quick online screening suggests an anxiety disorder, and her doctor recommends therapy—but the next available appointment is weeks away. Scrolling through her phone one night, Sophia finds an app that promises to ease anxiety with short daily exercises. It appears under the “medical” category, which feels reassuring—but what does that label really mean? The App Store is full of similar apps, all claiming to reduce stress or lift mood. Curious, Sophia searches for scientific reviews of mental health apps. Some studies report benefits for certain apps, yet she notices that the ones shown to work aren’t the ones she can actually download. The discovery leaves her uneasy: if most apps aren’t backed by evidence, how is anyone supposed to know which to trust?
Though fictional, the story of Sophia illustrates the dilemma of millions of people who suffer from mental illness while facing long waitlists and sizable bills for care. Globally, the prevalence of mental health disorders had risen by more than 30% since 1990—even before the COVID-19 pandemic (James et al., 2018). In the United States, mental health also continues to decline amongst adults (Wright et al., 2025), while at the same time almost two-thirds of American college students report experiencing “overwhelming anxiety” each year (American College Health Association, 2019). Existing infrastructure is failing to meet this increased demand for mental health services. In the United States, less than one in three people with mental health disorders receive treatment (Olfson et al., 2016), while the inequities in the prevalence of poor mental health continue to grow (Wright et al., 2025). Can mobile apps for mental health be part of a comprehensive solution to the mental health crisis in the United States and around the globe?
In the present article, we will show that despite some promising evidence, most mental health apps marketed to consumers are unvalidated to treat mental illness. As a result, an unregulated market of mental health apps might have little effect or even exacerbate the current mental health crisis. Accordingly, we propose regulatory and other solutions to protect consumers by ensuring that they can easily discern between validated and unvalidated apps. Our focus is on smartphone apps, but we also consider web-based apps where relevant. We examine mental health apps designed to deliver treatment without input from healthcare professionals (self-guided apps) and apps that deliver app-based treatment with the guidance of a professional (provider-guided apps). Our review focuses largely on mental health app research predating the AI-chatbot revolution in 2022, though we integrate recent findings where available. These recent technological developments underscore why policy action is now more pressing than ever. While AI chatbots show promise for democratizing mental health care, they also hold the potential of exposing millions of users to untested interventions that could delay appropriate care or cause harm.
We organize the present work into three sections. First, we review the existing evidence of the efficacy and effectiveness of mobile health (mHealth) apps for mental health. Efficacy trials test whether an intervention works under ideal and controlled conditions. In contrast, effectiveness trials assess how well an intervention performs in real-world settings, where users vary in motivation, context, and adherence. In other words, efficacy tells us if an intervention can work, while effectiveness tells us if it does work when deployed outside the lab. We highlight critical gaps in knowledge, including a lack of efficacy evidence compared to established treatments like talk or drug therapy (i.e., standard-of-care treatments), a lack of evidence in underserved populations, and a lack of evidence for real-world effectiveness. Notably, most of the existing evidence comes from studies in the United States, but we do note evidence from other countries where relevant. Second, we conduct an original analysis that demonstrates that the vast majority of apps available to consumers are unvalidated. Finally, we explore the implications of our findings for public policy, proposing specific ideas to protect and inform consumers.
We propose a multi-pronged approach to ensure the efficacy and safety of mental health apps. Our solutions are primarily focused on the U.S. regulatory landscape, but the overall principles may be applicable elsewhere. First, the U.S. Food and Drug Administration (FDA) should regulate these apps as medical devices or akin to dietary supplements, necessitating stringent efficacy and safety data or explicit disclaimers. To strengthen these efforts, we suggest refining and expanding the FDA's guidelines, such as extending regulatory oversight to consumer-facing mental health apps, requiring standardized disclaimers for untested claims, and/or requiring platforms such as Apple's App Store and Google Play to adopt consistent labeling and reporting standards. Apps should be listed under the medical category of Apple's App Store, for example, only if they meet the standards for medical devices. Every app claiming to treat psychiatric conditions should be accompanied by a disclaimer that clearly indicates whether peer-reviewed research supports its efficacy claims, ensuring transparent communication to users. Second, we propose the development of accessible databases and tools that aid consumers in selecting apps based on evidence. Custom-designed AI agents that help consumers use such databases interactively to reach better evidence-based decisions could be a way to use AI in a productive way in this space. These resources promise to simplify the evaluation of scientific support for such apps and clarify their relevance to specific demographic groups.
Part 1: Review
In our review, we use evidence from meta-analyses and existing reviews of randomized controlled trials (RCTs), where participants are randomly assigned to an app treatment or a control condition and where their mental health outcomes are compared across conditions. When meta-analyses are not available, we also consider evidence from individual RCTs. Where relevant, we evaluate the strength of the evidence by specifying both the sample size and the effect size. We use Cohen's conventions of effect size, where an effect of d/g = .20 is considered small, d/g = .50 is considered medium, and d/g = .80 is considered large (Cohen, 1988). Further details on our methodology are available in our Supplementary Materials (https://doi.org/10.17605/OSF.IO/5DKV).
Are Mental Health Apps Efficacious?
First, we consider existing evidence on the efficacy of mental health apps, that is, whether mental health apps can reduce symptoms of mental illness under ideal conditions. Several meta-analyses have concluded that mental health apps can be efficacious for a range of outcomes, such as symptoms of depression and anxiety (Doğan et al., 2024; Firth, Torous, Nicholas, Carney, Pratap, et al., 2017; Firth, Torous, Nicholas, Carney, Rosenbaum, et al., 2017; Linardon et al., 2019; Luo et al., 2025). One meta-analysis concluded that smartphone-based apps have small-to-medium improvements in symptoms of depression (number of studies = 54) and anxiety (number of studies = 39) from baseline to post-test, immediately after treatment (Linardon et al., 2019). Another meta-analysis of 16 studies concluded that self-guided digital interventions—the majority of which have been web-based rather than mobile-based—lead to a small but significant reduction in suicidal ideation at post-test (Torok et al., 2020). In contrast, meta-analyses have shown little evidence for the efficacy of mental health apps in reducing symptoms of panic disorder and posttraumatic stress disorder (Linardon et al., 2019) or reducing self-harm behaviors (Witt et al., 2017)—though this lack of evidence may be due to the relatively fewer studies that have examined the effects of apps on these disorders.
Efficacious Compared to What?
Mental health apps can be efficacious in reducing symptoms of anxiety and depression, but efficacious compared to what? Active control conditions involve comparing an intervention to another credible treatment or activity—such as psychoeducation or relaxation training—rather than to doing nothing. Active controls may also include treatment-as-usual (TAU) care, which refers to the care participants would typically be receiving had they not been part of a study, such as primary care, talk therapy, or drug therapy. The majority of active controls in existing literature, however, have been placebo/attention controls, such as listening to music or playing video games, rather than TAU (Linardon et al., 2019, 2025). In contrast, inactive controls, like waitlist or no-treatment groups, provide no alternative intervention. Comparing an app's efficacy against an inactive control condition, therefore, may reflect placebo or expectancy effects rather than specific benefits of the app itself. Comparing efficacy to active controls, in contrast, offers a more rigorous test of whether the app provides unique therapeutic value beyond generic support or being in a study.
The efficacy of mental health apps is larger when compared to inactive controls (e.g., waitlist control or no treatment control) than to active control conditions (Linardon et al., 2019). Whereas mental health apps show medium effects in reducing symptoms of both anxiety and depression when compared to inactive controls, they show only small effects when compared to active controls (Firth, Torous, Nicholas, Carney, Pratap, et al., 2017; Firth, Torous, Nicholas, Carney, Rosenbaum, et al., 2017). Still, the fact that mental health apps are more efficacious compared to active controls suggests that the mental health apps have specific therapeutic benefits.
A specific subtype of TAU is known as standard-of-care treatment, which refers to established and recommended courses of treatment. For a mental health disorder, standard-of-care may be talk therapy or drug therapy. A comparison of a new treatment with a standard-of-care treatment, therefore, can tell us how the new treatment compares to the currently best treatment options. We were not able to identify any studies that directly compared the efficacy of mobile-based mental health apps to standard-of-care, including talk or drug therapy. Thus, there is no direct evidence to indicate that mental health apps can serve as a replacement for established treatments. Rather, the evidence for their efficacy can be interpreted to mean that mental health apps may be better than doing nothing while waiting to receive validated in-person treatments.
Finally, when considering the comparative efficacy of mental health apps, we must consider whether mental health apps can improve the efficacy of existing treatments when these treatments are supplemented by a mental health app. The evidence here is mixed. One meta-analysis found no significant benefit of adding mental health apps to TAU compared to TAU alone on symptoms of depression, though this evidence comes from only four studies (Linardon et al., 2019). Supplementing an existing treatment with mental health apps may, however, be beneficial when patients are resistant to the treatment alone. A study of 164 patients recruited in Japan with antidepressant-resistant major depression compared patient outcomes for drug therapy alone versus drug therapy plus a self-guided app based on cognitive behavioral therapy (CBT). After 9 weeks of treatment, the drug-resistant patients in the app condition showed improvements in depression compared to drug therapy alone in the range of a small-to-medium effect size (Mantani et al., 2017). Apps may also be beneficial when used to maintain improvements after initial treatment. In one study, 349 patients with alcohol dependance were randomly assigned to either receive treatment as usual—as part of residential ‘rehab’ programs—or to receive the same treatment plus continued care after discharge via a smartphone app. The app provided monitoring, information, and support, including contact with therapists. Compared to TAU alone, patients who received app-facilitated continued care engaged in risky drinking less often (Gustafson et al., 2014).
In sum, there is evidence that mental health apps can be efficacious. However, it remains unclear how efficacious mobile apps are compared to standard-of-care, including psychiatric drugs and in-person therapy. The conclusions from our meta-review correspond to the conclusions of three published systematic meta-reviews of mental health apps (Astafeva et al., 2022; Goldberg et al., 2022; Lecomte et al., 2020). These meta-reviews agree that initial evidence is promising, but due to methodological concerns, small sample sizes, and other issues, these apps cannot yet be seen as robust alternatives to established modes of treatment.
Which Apps Are Efficacious?
Most of the research on the efficacy of mental health apps comes from studies on self-guided apps, where participants are using the app without direct supervision by a trained provider or practitioner. When comparing the efficacy of these self-guided apps to provider-guided apps, meta-analyses have concluded that self-guided apps appear to be less efficacious than provider-guided apps. One meta-analysis found that provider-guided apps produced a medium effect in reducing symptoms of anxiety and depression, whereas self-guided apps only produced small effects (Linardon et al., 2019). Another meta-analysis focusing on the efficacy of self-guided apps among users with more severe symptoms found a small-to-medium effect of self-guided apps on depression, but no effect on anxiety (Weisel et al., 2019). Overall, one meta-review (i.e., a summary of multiple meta-analyses) concluded that self-guided apps can be a promising stand-alone treatment for the management of symptoms of anxiety and depression, but that apps are more efficacious when combined with professional guidance or when used as supplements to existing treatments (Lecomte et al., 2020).
Recent research suggests that AI-enhanced apps—especially those featuring AI-based conversation agents, or chatbots—may improve efficacy compared to self-guided apps without AI. One meta-analysis, for example, found that AI-based conversation agents showed medium efficacy in reducing symptoms of depression. The review further concluded that these effects are stronger when the AI conversation agent is integrated into a mobile app (Li et al., 2023). Evidence further suggests that chatbots may improve the efficacy of mental health apps at least in part by facilitating the formation of human-like therapeutic bonds. One analysis of 36,000 users of a therapeutic chatbot showed strong working alliances between users and the chatbot, on par with those found between patients and traditional therapists in outpatient settings (Darcy et al., 2021). Overall, evidence suggests that AI-enhanced mental health apps and chatbot-based apps may improve efficacy by increasing engagement and strengthening the therapeutic alliance. However, studies are needed to directly compare chatbot-based mental health apps with standard-of-care treatments, such as talk and drug therapies. Until such evidence exists, it remains unknown if these apps might be assumed effective and lead users to delay or forgo more effective forms of care.
Who Are Apps Efficacious For?
Mental health apps have shown some promise as a tool to address the need for on-demand mental health treatments. But their true promise can only be realized if they prove to be an effective tool to improve the mental health of high-risk and traditionally underserved populations. But researchers investigating the impact of these apps have primarily relied on convenience samples consisting of participants who are easy to recruit—samples that frequently exclude the very same populations that are underserved by existing mental health services. In this section, we will take a critical look at the demographic composition of the samples in existing research.
Age
The efficacy of apps has primarily been assessed in middle-aged adults (ages 18–65), but less so in adolescents (under 18 years old) and seniors (over 65 years old). In a meta-analysis of 18 studies examining the efficacy of smartphone apps for depression, for example, the age of participants across all studies ranged from 18 to 58 (Firth, Torous, Nicholas, Carney, Pratap, et al., 2017). In a meta-analysis of 9 studies on anxiety, the age range was even more restricted: from 18 to 43 (Firth, Torous, Nicholas, Carney, Rosenbaum, et al., 2017). A more recent meta-analysis focused specifically on the efficacy of mobile applications on depression in adolescents and young adults (up to 35 years old). Across 12 trials, they found trends but no significant effects in this younger group (Lee et al., 2025). This calls into question the efficacy of mental health apps in youth, whose mental health has been of particular concern in recent years. The picture may change, however, as more apps integrate AI chatbots. A recent meta-analysis concluded that AI-driven conversation agents can be efficacious in reducing depression symptoms, specifically in youth and young adults ages 12 to 25 (Feng et al., 2025). There is little evidence for the efficacy, or lack thereof, of mental health apps amongst seniors who are 65 years or older.
Gender
Men are less likely than women to seek help for their mental health problems in part because of the greater perceived stigma of mental illness in men (Chatmon, 2020). By providing more private access to mental health therapies, apps have the potential to help men seek treatment. The problem is that men are also underrepresented in the literature on mental health apps. In a meta-analysis of app-based interventions for anxiety, only about a third (34.8%) of participants were male (Firth, Torous, Nicholas, Carney, Rosenbaum, et al., 2017). Similarly, a recent meta-analysis of mental health apps for depression, 72.8% of participants were female (Luo et al., 2025). That said, the authors concluded that the effects of the interventions did not vary by the gender composition of the study. Still, more evidence is needed to directly examine the efficacy and acceptability of mental health apps amongst men.
Race, Ethnicity, and Social Class
Racial minorities are also underserved in terms of mental health services, but in addition to stigma, the sources of the problem here are also structural. Structural factors include barriers that come from the way systems and institutions are organized—not from individual choices or attitudes. Structural barriers include greater shortages of providers in certain neighborhoods, longer waitlists, or a greater lack of insurance. Black people under the age of 60 in the United States, for example, are more likely than other demographic groups to report structural barriers as a reason for not receiving mental health treatment (Green et al., 2020). Similarly, due in part to barriers in help-seeking, suicide in Australian Indigenous communities is two times higher than amongst non-indigenous populations (Tighe et al., 2017)
By democratizing mental health treatment, apps could help remove some of those structural barriers. But is there existing evidence for their efficacy in racial and ethnic minorities? Users’ race and ethnicity have been underreported in existing studies, making it difficult to assess whether and to what extent the overall efficacy of mental health apps observed in existing meta-analyses applies to racial and ethnic minorities. One recent review of app-based digital interventions found that more than half of existing clinical trials do not report the racial composition of their samples (Kirvin-Quamme et al., 2024). Similarly, a 2024 systematic review of 62 phone-based and internet-based interventions using CBT for depression noted that only 27% of these RCTs reported participants’ race (De Jesús-Romero et al., 2024). Among the 3,623 patients whose race was documented, the vast majority were White (∼75%), with far smaller proportions identified as Black (∼7.6%) or Asian (∼5.8%).
Though racial minorities are generally underrepresented in existing research and meta-analyses, a smaller subset of studies have examined the efficacy of culturally-adapted digital interventions with promising results. Culturally-adapted interventions are specifically designed to provide care in specific groups or communities, often with input from members of the community during the design process. In a meta-analysis of 12 randomized trials (total sample of 653 participants), culturally-adapted digital interventions (both app and web-based) showed a large positive effect on mental health outcomes compared to control conditions. Even amongst those interventions, however, there was a paucity of research with Black and Indigenous communities (Ellis et al., 2022).
In this context of underrepresentation of underserved racial groups in existing research, several studies are notable for focusing primarily on populations that are typically underserved. A small study in the United States with 52 participants, who were primarily Black (55.8%), found that both a self-guided app based on Behavioral Activation and a self-guided app based on CBT were efficacious in reducing symptoms of depression compared to primary care with small to medium effects (Dahne et al., 2019). In addition, an Australian study conducted in remote communities in the Kimberley region of Northwestern Australia enrolled 61 participants, 94% of whom identified as Aboriginal or Torres Strait Islander. The sample faced socioeconomic challenges: 38% were unemployed and looking for work, and only 5% had completed a college degree. The treatment arm of the study used a self-guided app for six weeks. After the intervention period, the app group showed medium-to-large improvements in psychological distress and symptoms of depression compared to the waitlist control (Tighe et al., 2017).
In further support of the efficacy of mental health apps in racially and socioeconomically diverse populations, an 8-week coach-guided app treatment reduced symptoms of depression and anxiety compared to treatment as usual in a sample of 146 primary care patients suffering from anxiety and depression (Graham et al., 2020). Black patients (32%) were overrepresented compared to their composition of the U.S. population. Another study notable for its sociodemographic diversity included 315 participants from 45 states, with most of the participants being unemployed and 10% describing their living situation as homeless or residing in an assisted living space. The treatment group used an app called CORE, designed to teach users to increase their mental flexibility. Following 30 days of CORE, the treatment group improved significantly compared to a waitlist control with a medium effect on symptoms of depression and a small effect on symptoms of anxiety (Ben-Zeev et al., 2021).
Summary
In sum, a handful of small studies have shown that mental health apps can improve symptoms in certain racial and ethnic minoritized populations as well as low-income populations; these findings remain preliminary because of their small numbers and sample size. Larger, more rigorous studies are needed to establish efficacy. Similarly, more research is needed to establish the efficacy of mental health apps in men, youth, and seniors. Overall, more robust evidence is needed to determine whether mental health apps are a viable tool to improve care in underserved populations.
Are Apps Effective in the Real World?
Though there is some evidence for the efficacy of mental health apps, there is little evidence for their effectiveness in real-world settings. Efficacy trials attempt to establish whether an intervention can work in a controlled research setting, whereas effectiveness trials (or pragmatic trials) are designed to see if an intervention is effective in the real world, after factoring in things like (non)compliance and attrition (Gartlehner et al., 2006). Most existing studies on mental health apps focus on establishing efficacy rather than effectiveness. For example, a search of the literature for anxiety and depression apps identified 70 RCTs but only three studies with real-world evidence (Leong et al., 2022).
Even in well-controlled efficacy trials, about a quarter of participants drop out. In a meta-analysis of smartphone apps for depression, for example, researchers estimated a pooled dropout rate of 26.2% across eighteen studies (Torous et al., 2020). When people use a mental health app without the structure and the incentives of a research study, the attrition rate is even higher. Indeed, an analysis of real-world use of 93 popular mental health apps for anxiety, depression, and emotional well-being found that approximately 96.7% of users had stopped using the app 30 days after downloading it (Baumel et al., 2019). Despite these high dropout rates, the sheer scale of mental health app adoption means they still reach substantial numbers of people (Baumel et al., 2019). Even if only 3.3% of people continue to use a mental health app after 30 days, this could result in a practically meaningful impact across the population.
Of course, the possibility that millions have been helped by mental health apps rests on the assumption that the majority of apps available to consumers are effective in the real world. But such an assumption would be premature. Though some mental health apps are efficacious, the vast majority of apps available to consumers are not empirically validated for either their efficacy under ideal conditions or for their real-world effectiveness. Next, we examine the massive gap between the scientific evidence and the apps available to consumers.
Part 2: Can Consumers Trust Mental Health Apps?
Thousands of wellness and mental health apps are available to consumers on the iPhone and Android app stores (Buss et al., 2024; Larsen et al., 2019), but only a small number—in the dozens— have been tested using randomized control trials, or RCTs (Linardon et al., 2019, 2025), and most of those RCT-tested apps are not directly available to consumers (Buss et al., 2024). Yet, one analysis found that while 64% of mental health apps available to consumers made claims about their effectiveness, and 44% used scientific language to make those claims, only one app provided a citation for their claims (Larsen et al., 2019).
Given the rapidly evolving landscape of mobile mental health apps, we conducted our own analysis to understand the extent to which mental health apps available to consumers are supported by evidence. Following the approach of Larsen et al. (2019), we identified consumer mental health apps by searching the Apple App Store for terms corresponding to the nine most prevalent mental disorders worldwide based on the Global Burden of Disease Study 2019 (GBD 2019 Mental Disorders Collaborators, 2022): anxiety disorders, depressive disorders, intellectual disabilities, ADHD, conduct disorders, bipolar disorders, autism spectrum disorders, schizophrenia, and eating disorders. Guided by subcategories of the DSM-5 (Diagnostic and Statistical Manual, 5th edition, American Psychiatric Association, 2013), we included additional disorder-specific terms (e.g., agoraphobia, panic disorder). Searches were conducted on September 25, 2025, using Apple's iTunes Search API (Application Programming Interface) and limited to English-language mobile apps. The initial search yielded 2,313 unique apps. We then filtered apps to include only those explicitly referencing treatment (e.g., “therapy” or “treatment”), excluding 1,018 apps that did not meet criteria. We additionally removed 51 apps containing the word “test” in their titles (e.g., “Depression Test”) to eliminate self-assessment tools. The final analytic sample comprised 1,244 distinct mental health apps. We did not search for Android apps since the Google Play Store does not provide official code (API) for searching available Android apps.
Next, we found that about a third of the apps we identified (424 out of 1244 apps) are listed under the medical category in the Apple App Store. Though consumers may perceive such apps to provide evidence-based medical treatments, Apple classifies apps as ‘medical’ based on the claims of the marketer rather than on the presence of validated information attesting to their efficacy as medical treatments. Third, we found that 31% (391/1244) of app descriptions claim that the app is backed by scientific evidence. Specifically, an app was classified as making a quality claim if its description contained at least one of the following terms: evidence, research, or science. Yet, a search of the peer-reviewed literature revealed that only 27% (81) of the 391 apps that made empirical claims were backed by any publicly available scientific evidence. For our literature search, we used a large language model (LLM), GPT-4.1-mini (April 2025 release). See Supplementary Materials for details of our method (https://doi.org/10.17605/OSF.IO/5DKV).
Importantly, our review of the literature was quite broad, including correlational studies, experimental studies, and clinical trials (RCTs). This means that the number of apps with strong empirical support based on RCTs is likely even smaller. Indeed, an earlier review of mobile apps for depression and anxiety identified 179 apps available to consumers, but only three apps—less than 2%—had evidence from both RCTs and real-world effectiveness trials (Leong et al., 2022). Not only are commercially available apps largely unvalidated, but many apps are also not even based on established treatments. For example, CBT is a validated treatment for anxiety. One analysis of 361 free iPhone apps for anxiety and worry available to consumers showed that about three-quarters (269) of the apps were not based on CBT—the most validated treatment for anxiety (Kertz et al., 2017). If only a fraction of people download apps with validated efficacy, and most apps do not even use validated treatments, the real-world effectiveness of the mental health apps available to consumers may be both statistically and practically insignificant.
Part 3: Policy Recommendations and Other Solutions
To the extent that some people turn to unvalidated mental health apps for treatment instead of seeking validated treatments, mental health apps might not only be ineffective but also harmful: the time and effort spent on these apps represents a missed opportunity to seek treatments that could actually help. What can be done to protect consumers? Our solutions focus primarily on the approaches relevant to the regulatory landscape in the United States, but similar approaches could be applicable in other countries, as well.
One solution is for regulatory agencies, such as the U.S. Food and Drug Administration (FDA), to regulate mental health apps as they regulate other medical devices. Medical devices are defined by the FDA as devices “used to diagnose, prevent, or treat a medical disease or condition without having any chemical action on any part of the body” (Jin, 2014). This description sounds similar to what many mental health apps are claiming to be doing. If mental health apps claim to be medical devices, then they should be regulated as such. Importantly, the FDA does consider software functions provided by some mobile apps to be medical devices, but currently, the FDA focuses on mobile apps that use phone sensors, complement existing medical devices, or are used for active patient monitoring in the category of medical devices (U.S. Food & Drug Administration, 2022). Thus, the FDA's current regulatory focus excludes most mental health apps. Even classifying mental health apps as a Class I medical device—the lowest risk category, which includes devices such as bandages, hospital beds, or electronic toothbrushes—would provide some regulatory control, such as requiring app developers to keep records of complaints and report adverse effects from consumers. It could be argued, of course, that apps claiming to treat potentially life-threatening psychiatric disorders like depression are in a higher risk category than bandages and toothbrushes, which would make them Class II medical devices. In that case, regulation would include requirements for premarket authorization based on efficacy and safety data, as well as post-market surveillance studies with actual consumers (U.S. Food & Drug Administration, 2018).
Another regulatory solution is for the FDA to regulate mental health apps similarly to dietary supplements. In that case, the FDA will not need to approve each app based on clinical trials, but will require all mental health apps to include a disclaimer similar to the one found on food supplements, such as: “These claims have not been evaluated by the Food & Drug Administration.” Interestingly, the FDA already provides guidance for disclaimers that mental health apps should include, but these recommendations are currently nonbinding. In addition, the recommendations apply exclusively to prescription-only “computerized behavioral therapy devices for psychiatric disorders.” This FDA category refers to software-delivered therapy programs—not general wellness apps—that function as a medical treatment and require medical supervision, similar to how some medications require a prescription. This definition, however, excludes the majority of mental health apps available to consumers—even those making claims that they can provide treatment or those categorized as ‘medical’ in the App Store. These FDA recommendations encourage app developers to include (1) a statement about when and how to contact a healthcare provider, (2) a prominent label that the app “does not represent a substitution for a patient's medication,” (3) an explicit statement that the app has not been clinically tested or the results and methods of any clinical trials, and (4) clear identification of any functions that are not cleared by the FDA (U.S. Food & Drug Administration, 2020). A logical next step would be for the FDA to begin requiring all mental health apps marketed for the treatment or diagnosis of psychiatric disorders to include those statements.
It is important to note that the FDA does not currently have any regulatory guidance or requirements for commercial platforms that host mental health apps, such as the Apple App Store and the Google Play Store. As we saw in our analysis of apps available to consumers, this has allowed Apple to classify apps as ‘medical’ without following the same FDA guidelines for the classification and approval of medical devices. Even without explicit FDA regulation, Google and Apple could voluntarily adopt the FDA guidelines by only classifying apps under the medical category if they have met the burden of proof for medical devices. In addition, Apple and Google could require each app claiming to treat psychiatric conditions to, at the very least, be accompanied by a disclaimer, such as: “These claims have not been validated by the scientific community.” Unlike strictly regulatory solutions, such company-endorsed and company-initiated solutions have the notable advantage of influencing mental health outcomes globally, not just in the United States.
Looking beyond solutions that depend on underfunded regulatory agencies and the goodwill of for-profit corporations, consumers also need access to user-friendly central databases where they can check the evidence for the efficacy of each app. Several such tools, such as MindTools (Neary & Schueller, 2018) and MindApps (Beth Israel Deaconess Medical Center, Division of Digital Psychiatry, n.d.), are already available to consumers. MindApps, for example, is a free tool available to all, with a list of filters to tailor content. These tools provide a rating for each app based, in part, on scientific evidence. To ensure consumers are aware of these resources, funding is needed to increase consumers’ awareness and use of these databases.
Furthermore, future versions of such tools should make it even easier for consumers and providers to evaluate the strength of the existing evidence. An example of such a tool is the website examine.com, which evaluates the evidence for the efficacy of herbal and nutritional supplements. For each supplement, the website provides an evidence table that easily shows consumers how strong the evidence is for the efficacy of each supplement on each outcome, provides information on the magnitude of the effect, indicates how consistent the findings are, and supplies a list of all relevant citations. In addition, future tools should also include information on the efficacy of these apps in specific populations. Given that most studies identified in meta-analyses did not include people over 65, for example, consumers need to know that efficacy has not been established in older populations. Finally, future consumer-facing tools to evaluate the efficacy and effectiveness of apps can be enhanced by custom AI agents powered by large language models. This can help to improve the engagement of consumers with otherwise complex databases while also making it easier for providers to maintain and update these tools over the long term.
Conclusion
Overall, evidence suggests that mental health apps, including some self-guided apps, can be efficacious in reducing symptoms of some mental health conditions, primarily depression and anxiety. And recent evidence shows the promise of AI-driven chatbots in making mental health apps more engaging for consumers and efficacious in their impacts. Consumers should know, however, that the efficacy of the vast majority of apps available to them has not been established. In addition, there is very little evidence demonstrating the effectiveness of mental health apps in real-world settings. Evidence is also scant to test the efficacy of these apps in particular underserved populations, such as men, the elderly, and Black people, undermining their promise to provide better care to the underserved. Furthermore, the efficacy of digital treatments has not been tested compared to standard-of-care, such as talk or drug therapy. Thus, for now, apps and digital treatments cannot be recommended as a replacement for established in-person therapies, but they can be used to supplement treatment and offer treatment sooner when in-person therapy is not immediately available (Lecomte et al., 2020). Without further evidence and clear guidelines, however, we caution that this approach reinforces a two-tier system of mental health treatment, whereby people who can afford therapy receive standard-of-care treatment while underprivileged populations are left to rely on lesser, digital treatments.
Mental health apps could become a powerful tool in combating the global mental health crisis. To realize their true potential in democratizing mental health care, we need more evidence to determine what works and for whom, and more regulation to curtail the spread of unvalidated apps that may be doing more harm than good. There are many paths forward: regulatory agencies like the FDA can extend their oversight to mental health apps that claim to treat psychiatric conditions, requiring, at a minimum, that these apps include clear disclaimers about the evidence supporting their claims. App store platforms like Apple and Google can voluntarily adopt FDA guidelines and only classify apps as ‘medical’ if they meet appropriate standards of proof. Finally, consumer-facing databases that evaluate app efficacy can be supported with increased funding and visibility. With the right combination of rigorous research, thoughtful regulation, and accessible information, mental health apps can move from being a largely unvalidated marketplace to a legitimate component of evidence-based mental health care.
Supplemental Material
sj-docx-1-bbs-10.1177_23727322251405255 - Supplemental material for The Promise and Peril of Mental Health Apps
Supplemental material, sj-docx-1-bbs-10.1177_23727322251405255 for The Promise and Peril of Mental Health Apps by Kostadin Kushlev, Kibum Moon, Maureen Harris and Grace Falgoust in Policy Insights from the Behavioral and Brain Sciences
Footnotes
Acknowledgments
This project received no funding. We thank Robert L. Longyear for his help in conceptualizing this project.
Author Contributions
KK conceptualized the project, supervised the data collection and literature review, interpreted the analyses, wrote the original draft, and edited the manuscript. MH conducted the literature review and helped with the writing of the original draft. KM collected the data from the App Store, analyzed the data, created the visualizations, and helped with the writing of the original draft. GF helped with the literature review.
Declaration of Conflicting Interests
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The authors received no financial support for the research, authorship, and/or publication of this article.
Supplemental Material
Supplemental material for this article is available online.
References
Supplementary Material
Please find the following supplemental material available below.
For Open Access articles published under a Creative Commons License, all supplemental material carries the same license as the article it is associated with.
For non-Open Access articles published, all supplemental material carries a non-exclusive license, and permission requests for re-use of supplemental material or any part of supplemental material shall be sent directly to the copyright owner as specified in the copyright notice associated with the article.
