[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-researchers-train-a-30b-medical-model-that-beats-gpt-5":10,"sections":44},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":34,"tags":35,"sources":39,"feedback":43,"feedback_at":22,"cost_usd":43,"total_tokens":43},7296,"researchers-train-a-30b-medical-model-that-beats-gpt-5","Researchers Train a 30B Medical Model That Beats GPT-5","A 30B parameter medical model trained with rubric-guided reinforcement learning beats GPT-5 on a brutal health reasoning benchmark.","A 30B parameter model trained on synthetic medical scenarios now outperforms GPT-5 on one of the toughest health reasoning tests.\n\nResearchers built a two-stage training pipeline called Fathom-Vaidya. The first stage sharpens diagnostic reasoning, the step-by-step work of turning symptoms and lab results into a diagnosis, using MedBullets exam questions and reinforcement learning guided by rubrics rather than a single right answer. The second stage tackles clinical reasoning, the messier skill of managing a multi-turn conversation with a patient where there isn't always one correct move. For that stage, the team generated 5,300 synthetic multi-turn scenarios, each graded against multi-dimensional rubrics instead of a simple pass-fail check.\n\nThe payoff: the resulting 30B model scores 50.1% on HealthBench-Hard, ahead of GPT-5 in thinking mode, and shows more than a 10% jump on MedXpertQA. That matters because both benchmarks were built specifically to expose where medical LLMs fail, not to flatter them. Beating a much larger proprietary model on a benchmark designed to be brutal is a stronger signal than another leaderboard win on an easier test.\n\nRubric-based reward training has become a common way to teach models judgment instead of just facts, and this result is a reminder that a well-designed reward function can matter more than raw parameter count.","[\"ai\",\"healthcare ai\",\"medical ai\",\"benchmarks\"]","2026-09-23T04:00:00.000Z","2026-09-23T05:51:59.564Z","2026-09-23T05:52:05.452Z","published",null,[24,30],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Fix the confusing opening sentence ('diagnosing patients and talking to them like one') — the pronoun referent is unclear and reads as garbled; also clarify in the lede that the GPT-5 win is specific to the HealthBench-Hard benchmark, not an across-the-board result.","resolved",{"id":31,"reviewer":26,"round":32,"reason":33,"status":29},"editor-r2",2,"The headline claims this is an 'Open Model' but the source material never states the model is open-source\u002Fopen-weight — either verify and cite that release detail or drop 'Open' from the headline.","ai",[34,36,37,38],"healthcare ai","medical ai","benchmarks",[40],{"name":41,"url":42},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.24480",0,{"sections":45},[46,49,53,58,63,68,72,77,82,87,92,97,102,107],{"name":47,"slug":34,"count":48,"latest_published_at":18},"AI",4265,{"name":50,"slug":51,"count":52,"latest_published_at":18},"Security","security",707,{"name":54,"slug":55,"count":56,"latest_published_at":57},"Policy","policy",369,"2026-09-23T02:13:52.000Z",{"name":59,"slug":60,"count":61,"latest_published_at":62},"Deals","deals",202,"2026-09-22T23:00:04.000Z",{"name":64,"slug":65,"count":66,"latest_published_at":67},"Hardware","hardware",168,"2026-09-22T23:56:03.000Z",{"name":69,"slug":70,"count":71,"latest_published_at":18},"Science","science",133,{"name":73,"slug":74,"count":75,"latest_published_at":76},"Consumer Tech","consumer-tech",110,"2026-09-22T20:00:00.000Z",{"name":78,"slug":79,"count":80,"latest_published_at":81},"Software","software",80,"2026-09-22T23:32:52.000Z",{"name":83,"slug":84,"count":85,"latest_published_at":86},"Dev Tools","dev-tools",79,"2026-09-22T22:21:13.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"Startups","startups",65,"2026-09-22T22:06:48.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"Gaming","gaming",45,"2026-09-22T15:35:06.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"General","general",43,"2026-09-21T23:48:56.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"Reviews","reviews",27,"2026-09-22T13:00:00.000Z",{"name":108,"slug":109,"count":110,"latest_published_at":111},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]