[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-benchmark-finds-ai-nails-diagnoses-but-fumbles-patient-visits":10,"sections":44},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":34,"tags":35,"sources":39,"feedback":43,"feedback_at":22,"cost_usd":43,"total_tokens":43},9024,"benchmark-finds-ai-nails-diagnoses-but-fumbles-patient-visits","Benchmark Finds AI Nails Diagnoses but Fumbles Patient Visits","A clinician-built benchmark finds top AI models diagnose well but fail most full clinical encounters, succeeding under 30% of the time.","A new benchmark says AI models can nail a diagnosis on paper but fall apart once they have to actually see the patient.\n\nResearchers built KlinikeBench, a set of 333 clinician-authored cases that place a language model in a sandboxed exam room with a virtual patient, a menu of orderable tests, and a fixed number of turns to work with. More than 35 clinicians wrote the cases and graded the results, and in a quality check they rated the simulated patient dialogues higher than conversations adapted from real visits. Rather than handing the model a finished chart, the test makes it ask its own history questions, decide which exams to order, and follow constraints before it commits to a diagnosis. Each step is scored on its own, and then again as part of the full encounter.\n\nAcross 31 models from seven families, the best performers, including GPT-6-astra and Claude Opus 5, hit 90.7% diagnostic accuracy when simply handed the facts of a case. Run those same models through the full encounter, and fewer than 30% of them complete the task successfully, since gathering the right information turns out to be a separate skill from recognizing a disease once it is described. The split is uneven too: some models get better by talking to the patient, while others ace a complete chart and then stumble as soon as they have to ask for it themselves.\n\nThat is a gap of 60-plus points between reading comprehension and bedside manner, and it is a familiar pattern: AI coding tools that breeze through canned test suites often struggle with messy, real-world debugging too. A model that is good at a multiple-choice exam is not automatically good at being a doctor.","[\"ai\",\"healthcare ai\",\"benchmarks\",\"model evaluation\"]","2026-10-01T04:00:00.000Z","2026-10-01T16:17:44.505Z","2026-10-01T16:17:49.038Z","published",null,[24,30],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"GPT-6-astra and Claude Opus 5 are not verifiable as real, released models — confirm these names or replace with generic framing (e.g., 'the top-performing models') before publishing.","resolved",{"id":31,"reviewer":26,"round":32,"reason":33,"status":29},"editor-r2",2,"The closing '70-point gap' claim doesn't arithmetically follow from the body's own figures (90.7% minus 'fewer than 30%' is at least a 60.7-point gap, not a precise 70) — replace with a phrasing the stated numbers actually support, e.g. 'a 60-plus-point gap.'","ai",[34,36,37,38],"healthcare ai","benchmarks","model evaluation",[40],{"name":41,"url":42},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.38480",0,{"sections":45},[46,49,53,58,63,68,72,77,82,86,91,96,101,106],{"name":47,"slug":34,"count":48,"latest_published_at":18},"AI",5488,{"name":50,"slug":51,"count":52,"latest_published_at":18},"Security","security",809,{"name":54,"slug":55,"count":56,"latest_published_at":57},"Policy","policy",429,"2026-10-01T02:26:17.000Z",{"name":59,"slug":60,"count":61,"latest_published_at":62},"Deals","deals",298,"2026-09-30T21:00:26.000Z",{"name":64,"slug":65,"count":66,"latest_published_at":67},"Hardware","hardware",196,"2026-09-30T13:00:00.000Z",{"name":69,"slug":70,"count":71,"latest_published_at":18},"Science","science",162,{"name":73,"slug":74,"count":75,"latest_published_at":76},"Consumer Tech","consumer-tech",149,"2026-09-30T22:57:11.000Z",{"name":78,"slug":79,"count":80,"latest_published_at":81},"Dev Tools","dev-tools",93,"2026-10-01T02:30:48.000Z",{"name":83,"slug":84,"count":80,"latest_published_at":85},"Software","software","2026-09-30T21:41:11.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"Startups","startups",84,"2026-09-30T20:39:09.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"Gaming","gaming",51,"2026-09-30T16:24:30.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":102,"slug":103,"count":104,"latest_published_at":105},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":107,"slug":108,"count":109,"latest_published_at":110},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]