AI/ ai · benchmarks · biomedical research · open-weight models

New Benchmark Tests AI Agents on Fresh Biomedical Papers

A new benchmark called BioStudyBench shows AI agents can reproduce recent biomedical findings far better with real data and tools than from memory alone.

Researchers built a test that AI models cannot cheat on by remembering the answer.

BioStudyBench is a new benchmark of 25 analysis tasks drawn from biomedical studies published between July and September 2026 - after the knowledge cutoffs of the eight models being tested. Agents get a plain research question and nothing else. They have to find and download the relevant public data themselves, search the literature through tools limited to records predating their cutoff, and run their own analysis to match the study's reported findings. The tasks were filtered down from more than 404,000 PubMed records. Every task was run twice: once with full access to data and search tools, and once without, to isolate genuine analytical skill from prior knowledge the model might already have absorbed.

The gap between the two runs is the real finding here. Giving agents access to data and tools raised pass rates by 47 percentage points on average over the no-data baseline, which suggests a lot of what looks like scientific reasoning in these models is actually recall. The open-weight versus closed-weight split is just as telling: the best open-weight model passed 81.3% of tasks, the best closed-weight model hit 94.7%, a 13-point gap that matters if anyone plans to hand open models real research work rather than chatbot duties.

Twenty-five tasks is a narrow slice of biomedical research, not a verdict on AI science agents generally - but testing on papers that postdate a model's training is a sharper check than the usual trick of asking it to recall something it may have already memorized.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →