AI/ ai · benchmarks · coding-agents · swe-bench

SWE-bench's Top Coding Agents Are Statistically Tied

An audit of 254 SWE-bench submissions finds top coding agents share nearly all their wins and losses, meaning small leaderboard gaps are mostly noise.

The SWE-bench leaderboard's top coding agents are basically tied, a new statistical audit finds.

A team examined 254 SWE-bench submissions across four test splits without rerunning a single model. On the Verified split, the top two entries each resolved exactly 396 of 500 instances, and the top ten agents shared 285 successes and 51 failures in common, leaving just 164 instances where their results actually diverged. Paired statistical tests found no significant difference between any of the 29 adjacent pairs in the Verified top thirty. The gap separating the best and worst of that top thirty, 8.8 percentage points, was dwarfed by a 29.8-point swing observed when the same underlying model was paired with different scaffolding.

That scaffolding number is the real story. If swapping the tool harness around a model moves its score more than three times as much as separates the entire top thirty, the leaderboard is measuring scaffold engineering nearly as much as model capability. That gives vendors an incentive to tune wrappers for the benchmark rather than for messy real-world coding, and it undercuts headlines built on a one-point leaderboard lead. The pattern held across splits, though the larger Test split did separate more pairs, 14 of 23.

So the next state-of-the-art coding agent headline probably deserves an asterisk, not applause.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →