AI/ ai · ai-agents · benchmarks · research

Branching Search Lets AI Agents Improve Themselves Faster

A new arXiv preprint splits AI harness optimization into competing branches, beating Meta-Harness by up to 34.8% on math and smaller gains on coding tasks.

A new arXiv preprint says letting an AI coding agent split its own self-improvement process into competing branches, instead of following one search path, makes it notably better at solving math and coding benchmarks.

Posted September 30, 2026 as "Mixture of Self-Improving Branches For Agent Harness Optimization" (arXiv:2609.37834), the paper builds on a prior system called Meta-Harness, which has an AI agent rewrite its own tooling and test the results, feeding each round's outcome into the next. Meta-Harness relies on one fixed set of test cases and one strategy for proposing changes, which the authors argue can trap the process in a local optimum. Their alternative runs several branches at once, each keeping the test cases its top harnesses solve better than rivals', dropping cases every branch already handles, and evolving its own proposal strategy from its own search history. A router then assigns each new task to whichever branch's harness looks best, using only training data rather than the actual test outcomes.

This matters because agentic coding tools are increasingly built to modify their own scaffolding, and a process that reliably self-improves without babysitting is worth more than one that plateaus. The reported gains - 34.8% on Olympiad-level math, 11.6% on Terminal-Bench 2.0, and 3.8% on SWE-bench Lite, all relative to Meta-Harness - suggest branching helps most on tightly scoped, puzzle-like problems and less on messier, real-world coding tasks.

The comparison is against one predecessor system, not the field at large, and the shrinking margin from math to SWE-bench Lite hints that this is an incremental evolutionary-search tweak rather than a breakthrough in how agents learn. Still, the core idea - stop forcing every improvement attempt down a single path - is a sensible fix, and it's the kind of unglamorous infrastructure work that ends up quietly baked into the next generation of coding agents.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →