Researchers tested whether making an AI depression-screening method more auditable also makes it more accurate. It does not.
The study compared two ways of using large language models to rate depression severity from Reddit posts: asking a model to output a score directly, or walking it through the reasoning first (chain-of-thought), versus having it flag which clinical symptoms a post displays and letting simple code add up the count. The second approach, called criteria extraction, is easier to check since a clinician can review each flagged symptom. The team ran three models, from a 9-billion-parameter system up to frontier-scale ones, against two Reddit datasets and two standard depression questionnaires, PHQ-9 and BDI-II. Criteria extraction beat chain-of-thought on only one dataset, and only after its scoring thresholds were calibrated on labeled data; with thresholds set in advance, it showed no edge at all, and neither model's gain was statistically significant.
That matters because the whole pitch for criteria extraction is trustworthiness through transparency. But across nearly every comparison, extraction missed more severe cases than chain-of-thought, which missed more than direct prompting, even as its overall agreement score looked comparable or better. On the main dataset, a bare-bones model counting words in the post did no worse than a frontier model carefully flagging clinical criteria.
A method that is easier to audit isn't automatically one worth trusting, especially when the cases it quietly misses are the most severe ones.