AI/ ai agents · benchmarking · ai for science · research

New Benchmark Tests Whether AI Science Agents Actually Learn

ScienceClaw benchmarks 23 disciplines to see if AI agents' one-off fixes turn into lasting skills, not just lucky one-time wins.

A new benchmark asks a blunt question: when an AI agent fixes a broken science workflow, does it actually get better, or does it just get lucky once?

Researchers behind ScienceClaw built a framework and a companion benchmark, ScienceClaw-Eval, spanning 23 disciplines across the natural and social sciences. The system treats improvement as fixed-parameter program self-evolution, meaning the agent cannot simply retrain its weights; it has to repair and update its own executable workflows through multi-turn interaction. A fix only counts as a real update if replaying the original failing task reproduces the repair, and the change also helps on separate, independent tasks. ScienceClaw-Eval then tracks correctness, evolutionary gain, retention, cross-dataset transfer, and evolution cost across sequential task streams.

Most claims that AI science agents 'get better over time' rest on anecdote, a single success story rather than a measured trajectory. By forcing every update to survive replay and independent testing, ScienceClaw sets a higher bar than most coding-agent benchmarks, which usually just check whether a task passed once. That distinction matters because science agents are being pitched as lab collaborators that accumulate expertise, not disposable one-shot tools.

The code is open-source on GitHub, so other labs can test the retention claims themselves instead of taking the paper's word for it.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →