AI/ ai · healthcare · clinical-ai · llm-benchmarks

LLMs recommend more unnecessary care than doctors, study finds

A new benchmark of over 6,000 clinical scenarios finds AI triage advice shifts more than physicians' under irrelevant changes like gender and tone.

AI models recommending medical care are more trigger-happy than the doctors they might one day assist.

Researchers built a benchmark of more than 6,000 clinical triage scenarios, collected 7,000 physician annotations, and gathered 225,000 model responses to compare large language models against practicing physicians. They then tweaked the case descriptions in ways that should not change the clinical picture, such as the patient's gender or the tone of the message, to see if recommendations held steady. At baseline, the models recommended unnecessary care more often than physicians did. Under those harmless text changes, that gap widened, and the models' advice shifted more than the physicians' did, particularly in response to gender and tone.

Triage is exactly where hospitals and insurers want AI to save time, deciding who needs care right away versus who can wait. If a model's recommendation changes because a message sounds terse or mentions a different gender, that is a reliability problem, not a minor quirk, and it could translate into real inequities in who gets sent for further care. The study hands physicians and hospital systems an actual benchmark to check vendor claims against, instead of taking marketing language about validated AI at face value.

Call it the tell on overcautious software: when in doubt, the machine orders the test, and it is more easily rattled by phrasing than the doctors it is meant to assist.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →