AI/ ai · healthcare-ai · benchmarking · research-methodology

ChatGPT's Health Triage Failures Were Mostly a Format Problem

A replication of the study behind that scary 51.6% under-triage number finds the real culprit is how questions are asked, not the AI's medical judgment.

The scary stat that ChatGPT Health misses medical emergencies 51.6% of the time looks a lot shakier once you change how the test is given.

A new study re-examined the Nature Medicine research behind that number, which forced five models to answer in a rigid A/B/C/D format, blocked follow-up questions, and suppressed background medical knowledge. Researchers first ran a 17-scenario test bank through naturalistic, patient-style messages instead of that exam format, and accuracy rose by 6.4 points; on one vignette, models jumped from 0-24% accuracy to 100% once the forced-choice format disappeared. In a second, stricter pass, they replicated the original study's 60 published vignettes across six frontier models under four different formats, with clinicians checking the rewritten prompts and auditing the AI graders. This time the direction flipped: free-text answers actually scored slightly below the original rigid prompt (78.7% versus 81.8%), while naturalistic questions paired with a forced categorical answer beat both.

The real finding isn't that chatbots are dangerously bad at spotting emergencies - it's that a one-shot, multiple-choice benchmark can't capture how people actually talk to a health chatbot. Under the original four-point scale, every free-text miss was actually a model telling the patient to seek care that same day, just graded as the wrong letter, and 60% of those answers made escalation depend on information the patient still had to check.

Headlines about AI health risk travel fast; the methodology section in the fine print rarely catches up.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →