AI/ ai · mental-health · research · bias

Depression Detection AI Often Just Reads Interview Scripts

A new study finds depression detection models often key off interviewer scripts, not patient speech, raising doubts about their real accuracy.

A study of AI depression detectors finds many of them are grading the interviewer's script, not the patient's words.

Researchers examined three widely used interview datasets, ANDROIDS, DAIC-WOZ, and E-DAIC, built from semi-structured clinical interviews meant to screen for depression. They found that when models are trained on the interviewer's turns, the models can separate depressed from non-depressed participants with high accuracy just by learning the fixed prompts and where they fall in the conversation, without ever using the patient's answers. Restricting the same models to only the participant's utterances spread the decision-making evidence across a wider range of language cues, which the researchers describe as more genuine signal. The pattern held across all three datasets and multiple model architectures, meaning it is not a quirk of one corpus or one design.

Semi-structured interviews use the same script with every patient specifically so results are consistent and comparable. But that consistency is exactly what lets a model cheat: it learns the shape of the script instead of the content of the patient's speech. A depression-detection tool that is quietly grading the interviewer instead of the patient is not measuring what anyone thinks it is measuring, and that gap does not show up in a headline accuracy number.

High benchmark scores on these datasets, in other words, may say more about interview scripts than about depression.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →