Randomised controlled evidence in AI mental health is sparse, which is why a trial published in JAMA Network Open on 14 April 2026 matters. It is the largest and most carefully designed study of a conversational AI therapy tool to date. That is also the reason to read it carefully rather than celebrate it loosely. The signal is real. The confidence interval around the signal matters just as much.

The trial randomised 995 Israeli university students reporting psychological distress into three arms: a conversational AI platform called Kai, in-person group therapy with psychologists, and a waiting-list control. It ran from April to October 2025, with follow-up at 12 weeks. Both the results and their limits deserve equal attention.

What the trial found

The anxiety findings are striking. At 12 weeks, the AI group showed greater reductions in anxiety than both group therapy and control, with a mean difference of -2.17 against group therapy (95% CI -2.67 to -1.67). Among participants who began the trial with clinical-level anxiety, 57.9 per cent in the AI group remitted to nonclinical levels, against 14.4 per cent in group therapy and 9.8 per cent in the control group. That is a large gap, and it is the headline most coverage has led with.

The depression findings are more qualified. Symptoms improved relative to the waiting-list control, but after adjustment they did not differ significantly from group therapy. On post-traumatic stress symptoms, the AI group showed no significant advantage at all. The benefit, in other words, was real but condition-specific, and the researchers were careful not to claim otherwise.

Engagement held up better than these tools usually manage. Around 61 per cent of participants were still using the platform at 12 weeks, at a mean of roughly three days per week, in a category where most products lose users within days.

The finding that should interest builders most

The most interesting result in the paper is not the symptom reduction. It is the therapeutic alliance finding. Participants who reported a stronger sense of being understood by the AI engaged more, and that engagement predicted greater symptom improvement. The perceived bond was not incidental to the outcome. It appeared to be part of the mechanism.

That is a design insight before it is a clinical one. It suggests that the felt quality of the interaction, the sense of being heard, is doing meaningful work, and that building for alliance is not soft or secondary. For anyone designing in this space, it reframes the question from how to deliver an intervention to how to make the interaction feel responsive enough that people return to it.

The caveats, named directly

These results need bounding, and the limits are not footnotes.

The sample was university students, not a clinical population. The findings may not generalise to people with moderate or severe presentations. The comparator was group therapy, which is a weaker benchmark than individual therapy; the study does not tell us how Kai performs against one-to-one human care. The trial was conducted in Israel with a digitally literate student cohort, during a period of national security tension, so generalisability to other populations and contexts is unclear.

Then there is the conflict of interest, which needs saying plainly rather than burying. The lead author, Professor Anat Shoshani, is chief psychologist at Kai.ai, the company whose platform was tested, and reported personal fees and stock options in the company. This does not invalidate the findings. The trial was randomised and the methodology is sound. But it is a significant interest, and it belongs in any honest reading of the result.

Finally, no serious adverse events were reported, which is reassuring but limited. The monitoring period was short, and safety data from uncontrolled, real-world deployment, without the structure of a trial around it, would look different.

Where this leaves the field

None of these caveats erase the result. They locate it. This is what it looks like when digital mental health evidence starts to mature: a genuine randomised trial, with genuine results, and limitations that need naming rather than smoothing over.

For builders and clinicians, the trial offers two durable lessons. The first is that conversational AI can produce measurable benefit for some conditions, in some populations, under some conditions of care, and that precision about which is the difference between a credible claim and an overclaim. The second is that the alliance result points to where the real design work sits. The open questions, how this performs against individual therapy, in clinical populations, and over longer horizons, are exactly the ones the field should be funding next.

🪺