A new benchmark asks a simple question: do AI coding agents patch broken scientific code better when handed a runnable check instead of a paragraph of rules.
The project, called Rules to Tools, pairs matched groups of agents repairing code from the SciCode benchmark, which packages equations, boundary conditions, and output requirements from real scientific software. Each group starts from identical code, the same model, and the same budget. One group gets the requirements as written text; the other gets a callable tool that actually runs the check. Across the two task-ID cohorts tested, the tool group completed repairs on 29 of 30 problems versus 26 of 30 for the text group, though the per-task picture was messier: three task IDs favored tools, one favored text, and eleven ended in a tie.
Zoom into the eight-task subset where the researchers ran their one statistical test, and the gap narrows to 15 of 16 versus 13 of 16, with a bootstrapped 95% confidence interval of -12.5 to 43.75 percentage points for that subset, a range wide enough to include "no difference at all." On a larger shared-definition set of SciCode tasks, both groups tied at 13 of 24.
The more telling split shows up when the starting code doesn't match what a model likely memorized from training: on five such tasks, the tool group hit 7 of 10 versus 3 of 10 for text, suggesting runnable checks matter most when agents can't just pattern-match an answer. In a separate PDE comparison, checks also cut reported model output by 31.2% while matching text's accuracy, a real cost saving even where the correctness gap disappears.
Translate the small samples and overlapping intervals honestly: this reads as a plausible argument for executable specs over prose ones, not proof that handing an agent a tool beats handing it instructions.