There are now more mental health apps available than there are trained psychiatrists in most countries. The WHO estimates thirteen mental health workers per hundred thousand people globally, and the pressure to fill that gap with technology is enormous. So it is worth pausing, carefully, to look at what the evidence actually shows.

What the Research on Mental Health Apps Actually Shows

A large meta-analysis published in The Lancet Digital Health in late 2025 found that apps do produce meaningful reductions in symptom severity compared to inactive controls, particularly for depression and anxiety. A separate meta-analysis in the British Journal of Psychiatry in 2025, covering 17 RCTs with over 2,800 participants, found a pooled standardised mean difference of -0.46 for depressive symptoms. That moderate effect held across both adolescent and adult populations.

But the caveats matter. A review in World Psychiatry in 2025 noted that as few as 2% of publicly available wellbeing apps have any scientific evidence supporting their feasibility. User engagement remains a stubborn problem: up to 60% of participants do not complete all prescribed modules, and 30-day retention rates for popular mental health apps can sit as low as 3%.

The implication is not that apps do not work. It is that the gap between clinical trial performance and real-world impact remains large.

LLMs in Mental Health: Evidence and Gaps

Large language models arrived in clinical and consumer mental health settings with very little advance warning. A systematic review published in January 2026 in Electronics examined 205 studies on the use of LLMs in psychiatry, psychology, psychotherapy, and clinical workflows. The review found promising short-term performance across domains, but concluded that most evaluations relied on small, non-longitudinal datasets.

A separate systematic review in JMIR Mental Health in 2025 found that LLMs show useful performance in detecting mental health conditions through text analysis and in providing accessible, destigmatised support. The same review concluded that the current risks associated with clinical use may surpass the benefits, citing hallucinations and the absence of a benchmarked ethical framework.

The Safety Question Is Not Yet Settled

A 2025 study published in JMIR compared the performance of three leading LLMs against expert suicidologists using a standardised suicide intervention inventory. The LLMs showed competence in some dimensions but significant gaps in others. There are also documented cases of AI chatbots reinforcing delusional thinking in users already predisposed to such conditions.

These are not arguments against LLM-based tools. They are arguments for the kind of evidence generation the field has not yet done consistently.

What Builders and Clinicians Should Take From This

The research picture in 2025 and 2026 is not a green light. But it is not a red light either. For founders and clinicians building here, the most honest interpretation of the current evidence suggests three things.

First, the regulatory environment is tightening and will likely continue to do so. Building rigorous evaluation into the product from the beginning is no longer optional. Second, engagement is an unsolved design problem, not a product marketing problem. Third, LLMs are already embedded in the mental health ecosystem at scale, whether the clinical community is ready or not.

The gap is real. The tools are imperfect. The work is worth doing anyway.