A new benchmark asks a blunt question: when an AI agent fixes a broken science workflow, does it actually get better, or does it just get lucky once?
Researchers behind ScienceClaw built a framework and a companion benchmark, ScienceClaw-Eval, spanning 23 disciplines across the natural and social sciences. The system treats improvement as fixed-parameter program self-evolution, meaning the agent cannot simply retrain its weights; it has to repair and update its own executable workflows through multi-turn interaction. A fix only counts as a real update if replaying the original failing task reproduces the repair, and the change also helps on separate, independent tasks. ScienceClaw-Eval then tracks correctness, evolutionary gain, retention, cross-dataset transfer, and evolution cost across sequential task streams.
Most claims that AI science agents 'get better over time' rest on anecdote, a single success story rather than a measured trajectory. By forcing every update to survive replay and independent testing, ScienceClaw sets a higher bar than most coding-agent benchmarks, which usually just check whether a task passed once. That distinction matters because science agents are being pitched as lab collaborators that accumulate expertise, not disposable one-shot tools.
The code is open-source on GitHub, so other labs can test the retention claims themselves instead of taking the paper's word for it.