Most digital mental health services sit on a quiet mountain of user feedback. People fill in the box at the end of a conversation and say what helped and what did not. Then that text tends to go nowhere in particular. It gets skimmed, sampled, or filed away. The numbers we track instead are the easy ones. Response times. Session length. Completion rates. All useful, and all describing how a service runs. They say very little about how it felt to the person on the other end.

A recent report from Mental Health Innovations (MHI) offers a way through this. Listening to Shout user feedback at scale, written by Dr Emilia Piwek from MHI's Data Insights team, looks at close to 40,000 pieces of anonymous free-text feedback left by users of Shout, the UK's 24/7 crisis text service. The report is about one service. Its lesson reaches much further.

Why user feedback matters

Operational metrics answer operational questions. They tell you whether a service is fast and whether people finish their conversations. They do not tell you whether someone felt heard.

In mental health support, that gap matters more than in most fields. Trust, empathy, and the sense of being understood shape whether a person comes back, and whether they reach out again when things get hard. Torous and colleagues make this point across their reviews of digital psychiatry: evaluating these tools well means paying attention to experience and engagement, not only clinical endpoints. And experience tends to live in people's own words.

The trouble is that words are hard to count. Reading 40,000 comments by hand is slow, and it gets slower as a service grows. So most of that feedback goes unread, or gets reviewed once and then forgotten.

Inside the Shout report

Rather than sampling a handful of comments, MHI analysed the feedback from across 2024 and 2025 together.

The headline is warm. 84% of texters who left written feedback also rated their conversation as helpful. When researchers looked at what those people actually said, three things came up again and again: the kindness and compassion of the volunteer, having the space to work through their thoughts, and getting a prompt, useful reply. One in five went out of their way to praise their volunteer as a person. Feeling listened to, without judgement, sat at the centre of nearly every positive comment.

The report does not look away from the harder feedback. Among those who found their conversation unhelpful, the most common complaints were specific and fixable. Conversations that ended too abruptly. Replies that felt robotic or scripted, to the point that some texters wondered whether they were talking to a bot. Long waits. Advice that missed the mark. Around one in ten described the exchange as impersonal.

None of this was left as a chart. MHI traced these complaints back to the conversations that produced them, then used what they learned to update volunteer guidance and to build monitoring that flags conversations for extra support. Feedback went in one end and better training came out the other. That loop is the whole point.

How the analysis actually works

It helps to be precise about the method, because "AI read the feedback" both undersells it and oversells it.

MHI used language models to turn each comment into a numerical representation of its meaning, then grouped similar comments together automatically. The result is a map of recurring themes across tens of thousands of notes, surfaced without anyone having to read every line. Less common themes that a hand-picked sample would miss still show up.

This sits on peer-reviewed foundations. In a 2021 study in Frontiers in Digital Health, researchers from Imperial College London and MHI applied natural language processing to Shout conversations and classified conversation stages, volunteer behaviours, and texter demographics with accuracy in the high 80s and 90s. The feedback report is the applied, service-facing cousin of that academic work.

AI that supports, rather than replaces

The Shout project is a good example of where AI actually earns its place in care. It did not judge conversations or make decisions on its own. It helped a small team notice patterns across a volume of feedback no human could read in full, so that people could decide what to do about them.

Topol has argued for years that AI does its best work when it gives clinicians back their time and attention. Standing in for them was never the point. The Shout report is that argument in miniature. The machine finds the patterns. Human judgement, clinical knowledge, and organisational context decide what those patterns mean and what should change.

What this means for the rest of the field

Most services already collect qualitative feedback. Far fewer do anything systematic with it. As demand for digital mental health keeps climbing, reading everything by hand stops being realistic, and the insight that could improve a service quietly goes to waste.

There is a real opportunity here. Analyse feedback routinely and you can catch emerging problems earlier, see whether a change actually helped, and keep service development anchored to what users experience rather than what the dashboard reports. It is a more honest way to build.

One caution the report is careful to name: survey feedback is shaped by who chooses to leave it. People who had a good experience may be more likely to write in. No single method removes that skew. Using several lenses at once, from large-scale text analysis to smaller in-depth user groups, keeps the picture honest.

The part technology cannot do

For all the machine learning, the thing texters valued most was deeply human. Being met with kindness. Being given room to speak. Being believed. AI can help an organisation hear those experiences at scale and learn from them. It cannot supply the warmth itself. That still comes from a trained volunteer, awake at 3am, choosing the right words for a stranger who needs them.

That is the balance the Shout report gets right. Use the technology to listen better. Leave the caring to people.

🪺