AI/ ai · healthcare ai · medical ai · benchmarks

Researchers Train a 30B Medical Model That Beats GPT-5

A 30B parameter medical model trained with rubric-guided reinforcement learning beats GPT-5 on a brutal health reasoning benchmark.

A 30B parameter model trained on synthetic medical scenarios now outperforms GPT-5 on one of the toughest health reasoning tests.

Researchers built a two-stage training pipeline called Fathom-Vaidya. The first stage sharpens diagnostic reasoning, the step-by-step work of turning symptoms and lab results into a diagnosis, using MedBullets exam questions and reinforcement learning guided by rubrics rather than a single right answer. The second stage tackles clinical reasoning, the messier skill of managing a multi-turn conversation with a patient where there isn't always one correct move. For that stage, the team generated 5,300 synthetic multi-turn scenarios, each graded against multi-dimensional rubrics instead of a simple pass-fail check.

The payoff: the resulting 30B model scores 50.1% on HealthBench-Hard, ahead of GPT-5 in thinking mode, and shows more than a 10% jump on MedXpertQA. That matters because both benchmarks were built specifically to expose where medical LLMs fail, not to flatter them. Beating a much larger proprietary model on a benchmark designed to be brutal is a stronger signal than another leaderboard win on an easier test.

Rubric-based reward training has become a common way to teach models judgment instead of just facts, and this result is a reminder that a well-designed reward function can matter more than raw parameter count.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →