A new multi-agent system called MedBenchAgent doesn't just generate test questions for medical AI models. It builds the test itself.
Researchers describe MedBenchAgent as a framework that treats benchmark creation as "constrained compilation": deriving what to test, which annotations justify each task, and how to turn that evidence into graded questions. It splits the work into a planning stage that locks a specification and an instantiation stage that writes and audits items against it. In testing, the system identified the right evaluation tasks with a 90.9% F1 score, beating simpler automated approaches that scored between 79% and 85%. Of 1,000 sample items drawn from correctly identified tasks, human reviewers passed 994, and the team used the resulting benchmark to evaluate twelve vision-language models, including a run in a specialized medical domain.
Most automated benchmark tools skip the hardest step: deciding what counts as a fair, clinically grounded question in the first place, usually leaving that curation to overworked clinicians. MedBenchAgent's results also showed something aggregate leaderboard scores tend to hide: model performance swung widely by task and setting, meaning one headline number can mask a system that nails imaging description but fails at diagnostic reasoning.
Medical AI has a trust problem that better scores alone will not fix, but a benchmark that can show its work, task by task, is at least a start.