AI/ ai agents · reinforcement learning · swe-bench · arxiv research

New Training Method Boosts AI Coding Agents Without Human Labels

Counterfactual Rollout Replay trains coding agents on forked what-if branches, lifting benchmark pass rates without added human supervision.

A new training trick lets AI coding agents learn from paths they never actually took.

The paper, "Counterfactual Rollout Replay: Forkable Environments as Free Process Rewards for Software Engineering Agents" (arXiv:2609.33875, posted September 30, 2026), introduces Counterfactual Rollout Replay, or CRR. Instead of only rewarding an agent for finishing a coding task successfully, CRR picks a handful of decision points mid-task, rewinds the environment to that exact state, tries a different action, and plays the alternate path forward. It then compares the outcome of that alternate path to what actually happened and uses the difference to correct the training signal at that step. Tested with a 14B parameter policy on SWE-bench Verified, SWE-bench Live, and SWE-rebench, the method improved pass@1 scores across all three, and on SWE-bench Verified it hit 41.7% versus 36.7% for an extended outcome-only GRPO baseline under an equal-wall-clock comparison on the same hardware, a 5-point gain even after accounting for the extra compute the forking requires.

That matters because most reinforcement learning for coding agents today only tells the model whether the whole task succeeded or failed, leaving it to guess which of dozens of intermediate steps actually helped. Getting that step-level signal has usually meant paying humans to label decisions or training a separate reward model, both of which are slow, expensive, and prone to being gamed. CRR gets a similar signal for the cost of extra simulation, no labels or reward model required, which is a meaningfully cheaper path to the same kind of credit assignment.

The paper is upfront that this only works where the environment can be cheaply and reliably rewound, which leaves out plenty of real coding work involving live APIs, databases, or irreversible file changes. Even with that caveat, a free 5-point bump on a benchmark as heavily gamed as SWE-bench Verified is worth watching, especially as labs increasingly compete on process-reward tricks rather than just bigger models.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →