A new training trick lets AI agents learn even when every attempt at a task scores exactly the same.
Researchers have built on reinforcement learning with verifiable rewards (RLVR), the technique that trains AI agents by scoring batches of attempts at a task and nudging the model toward whatever scored best. The catch: when every attempt in a batch gets the same score - all successes, or more often all failures - there is no contrast to learn from, and the signal disappears. The new method, called Self-Retrospection Distillation (SRD), fixes that by having the model study its own completed attempts after the fact, then training a version of itself to predict useful moves before it acts, using what hindsight revealed. Across 10 tool-use and long-horizon agent tasks, adding SRD on top of standard RLVR produced gains of up to 24.2 percentage points, with the biggest payoff when reward-uniform batches were common - as much as 37-98% of them, depending on model size.
The sharpest example: on a 2-billion-parameter model, 98% of training batches were complete failures with nothing to differentiate between attempts. Standard RLVR training stalled at 0.0% success. Adding SRD pushed the same model to 60.6% success on the same budget of attempts. That is not a marginal tuning gain. It is the difference between a training run that goes nowhere and one that works.
It is a reminder that a lot of RL-for-agents research keeps hitting the same wall: reward sparsity. As labs push RL onto harder, longer-horizon tasks - the kind where an agent might fail dozens of times before succeeding once - the all-failure batch becomes the norm, not the edge case. SRD looks like one more patch on that problem, alongside self-distillation and curriculum tricks others have tried, rather than a wholesale fix. Worth noting: the foresight predictions SRD trains are never actually used when the agent runs for real, only during training, and how the approach holds up on bigger, frontier-scale models is still an open question.