[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-ai-clinical-agents-falter-on-messy-hospital-records":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},8938,"ai-clinical-agents-falter-on-messy-hospital-records","AI Clinical Agents Falter on Messy Hospital Records","A new benchmark of 365,000 hospital patients finds AI clinical agents' success rate drops from 62.2% to 37.9% when records turn noisy.","AI agents reading hospital records get a lot less reliable the moment those records turn messy, according to a new benchmark.\n\nResearchers built EHR-RobustGym, an interactive testing environment grounded in MIMIC-IV, a hospital database covering 365,000 patients, 31 tables, and more than 500 million records. The benchmark includes 5,486 matched pairs of clean and noisy clinical questions, corrupting data at the record, value, or query level. Across six clinical intents and both patient-level and population-level queries, the team evaluated multiple proprietary and open-weight large language models. Average task success dropped from 62.2 percent on clean questions to 37.9 percent once noise was introduced, and most models answered consistently less than half the time when asked the same question four times.\n\nThat gap matters because real hospital records are rarely clean. A clinical agent that returns a confident, plausible-sounding answer instead of flagging missing or contradictory evidence is more dangerous than one that is merely slow, and this benchmark isolates exactly that failure mode by pairing clean and noisy versions of the same question.\n\nThe researchers found that fine-tuning and reinforcement learning on EHR-RobustGym improve robustness, with gains carrying over to five external EHR benchmarks - useful, but a reminder that this problem gets solved through training, not better prompting.","[\"ai\",\"healthcare\",\"llm-benchmarks\",\"clinical-ai\"]","2026-10-01T04:00:00.000Z","2026-10-01T11:47:20.290Z","2026-10-01T11:47:25.395Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Fix the dek: it says the benchmark is 'built on 365,000 real hospital records,' but the body clarifies MIMIC-IV has 365,000 patients and over 500 million records — correct the dek to say patients, not records, so the two numbers don't contradict each other.","resolved","ai",[30,32,33,34],"healthcare","llm-benchmarks","clinical-ai",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.39371",0,{"sections":41},[42,45,50,55,60,65,70,75,80,84,89,94,99,104],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",5351,{"name":46,"slug":47,"count":48,"latest_published_at":49},"Security","security",801,"2026-09-30T22:18:23.000Z",{"name":51,"slug":52,"count":53,"latest_published_at":54},"Policy","policy",429,"2026-10-01T02:26:17.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Deals","deals",298,"2026-09-30T21:00:26.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Hardware","hardware",196,"2026-09-30T13:00:00.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":69},"Science","science",157,"2026-09-30T15:00:56.000Z",{"name":71,"slug":72,"count":73,"latest_published_at":74},"Consumer Tech","consumer-tech",149,"2026-09-30T22:57:11.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Dev Tools","dev-tools",93,"2026-10-01T02:30:48.000Z",{"name":81,"slug":82,"count":78,"latest_published_at":83},"Software","software","2026-09-30T21:41:11.000Z",{"name":85,"slug":86,"count":87,"latest_published_at":88},"Startups","startups",84,"2026-09-30T20:39:09.000Z",{"name":90,"slug":91,"count":92,"latest_published_at":93},"Gaming","gaming",51,"2026-09-30T16:24:30.000Z",{"name":95,"slug":96,"count":97,"latest_published_at":98},"General","general",50,"2026-09-30T21:37:54.000Z",{"name":100,"slug":101,"count":102,"latest_published_at":103},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":105,"slug":106,"count":107,"latest_published_at":108},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]