AI/ ai · mental-health · chatbots · healthcare

AI Mental Health Chatbot Pilot Finds Uneven Results

A year-long pilot of a purpose-built mental health AI showed strong average symptom drops, but most of the 299 participants didn't actually improve.

A foundation model built specifically for mental health just wrapped a year-long real-world test, and the topline numbers look good as long as you don't look too closely.

Researchers tracked 299 US adults with at least moderate depression or anxiety symptoms for up to 12 months as they used a generative AI chatbot designed for mental health. By 10 weeks, average depression and anxiety scores had dropped substantially, with effect sizes (Cohen's d of 0.93 and 0.79) in the range typically seen in in-person therapy trials. Participants also reported less loneliness, more social interaction, and more behavioral activation. The chatbot used ten distinct intervention types, shifted toward more clinical content as symptoms worsened, and its safety escalations were confirmed appropriate by clinicians.

Averages hide the real story. Only 37.1% of participants were classified as improving and just 5.7% as rapidly improving. The remaining 57.2%, a clear majority, were non-responders. That split matters because health systems are increasingly pointing people toward chatbots to cover therapist shortages, and "most users saw no benefit" is a very different pitch than the headline effect sizes suggest.

It's also a single-arm study with no control group or placebo chatbot to compare against, so the averages are a reason for cautious interest, not a replacement for a randomized trial.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →