AI/ ai · healthcare ai · benchmarks · model evaluation

Benchmark Finds AI Nails Diagnoses but Fumbles Patient Visits

A clinician-built benchmark finds top AI models diagnose well but fail most full clinical encounters, succeeding under 30% of the time.

A new benchmark says AI models can nail a diagnosis on paper but fall apart once they have to actually see the patient.

Researchers built KlinikeBench, a set of 333 clinician-authored cases that place a language model in a sandboxed exam room with a virtual patient, a menu of orderable tests, and a fixed number of turns to work with. More than 35 clinicians wrote the cases and graded the results, and in a quality check they rated the simulated patient dialogues higher than conversations adapted from real visits. Rather than handing the model a finished chart, the test makes it ask its own history questions, decide which exams to order, and follow constraints before it commits to a diagnosis. Each step is scored on its own, and then again as part of the full encounter.

Across 31 models from seven families, the best performers, including GPT-6-astra and Claude Opus 5, hit 90.7% diagnostic accuracy when simply handed the facts of a case. Run those same models through the full encounter, and fewer than 30% of them complete the task successfully, since gathering the right information turns out to be a separate skill from recognizing a disease once it is described. The split is uneven too: some models get better by talking to the patient, while others ace a complete chart and then stumble as soon as they have to ask for it themselves.

That is a gap of 60-plus points between reading comprehension and bedside manner, and it is a familiar pattern: AI coding tools that breeze through canned test suites often struggle with messy, real-world debugging too. A model that is good at a multiple-choice exam is not automatically good at being a doctor.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →