A new benchmark says today's best AI models still fumble basic evidence-citation in biophysics research.
Researchers released BioPhys-Bridge, a benchmark of 500 cases and 1,517 tasks spanning six biological domains and nine physical-model families. It is built to test whether language models can ground biological claims in the right evidence, equations, and units rather than produce plausible-sounding text. Each case pairs quantitative data, source equations, and stated assumptions with a specific biological mechanism, and scoring checks whether a model's answer traces back to the correct evidence ID. The team ran quality checks covering schema, evidence integrity, and unit normalization, and had domain experts hand-review 81 cases.
In preliminary tests described in the paper, the authors report an evidence-ID F1 score of 0.360 for the top-scoring system, 0.316 for the runner-up, and 0.294 for GPT-4o-mini. A score below 0.4 for even the best performer is the real headline here: it shows models pitched as science assistants still struggle to show their work, which is exactly what matters when the output is meant to inform real experiments rather than just sound plausible.
One caveat worth flagging: the paper credits those top two scores to systems named DeepSeek-V4-Flash and Qwen3.7-Max, version labels that do not match any publicly documented release from either company, so read those specific rankings as the paper's own claim rather than a verified result.