AI/ reinforcement-learning · llm-agents · ai-research · arxiv

New RL Method Salvages Failed AI Agent Training Runs

A new training algorithm called MVPO extracts usable signal from failed attempts, boosting two benchmark scores without slowing training much.

A new reinforcement learning method lets AI agents extract training signal from their failed attempts, not just their successes.

Researchers describe Milestone Viability Potential Policy Optimization (MVPO), a training method built to fix what they call "zero-credit failure" - when popular methods like GRPO and GiGPO compare an AI agent's different attempts at a task but find no variation to learn from, because early in training almost every attempt fails outright. That leaves potentially useful partial progress, buried inside those failed attempts, completely unused. MVPO fixes this by mapping which partial action sequences were still on a viable path toward success, using a technique called Union-Find, then assigning those sequences a score so the model gets credit for getting partway there. Tested on the Qwen2.5-1.5B-Instruct model, MVPO beat eight rival training methods, improving success rates over GiGPO by 4.4 points on the ALFWorld benchmark and 5.3 points on WebShop, while adding under 0.2% extra computational overhead.

That efficiency is the real story. Long-horizon agent tasks - the kind that require ten or twenty correct steps in a row - are notoriously hard to train because reward only shows up at the very end, if at all. A method that squeezes usable signal out of near-misses without adding meaningful compute cost addresses a genuine bottleneck in how these agents get built, not just a benchmark curiosity.

The catch is scale: this is one 1.5-billion-parameter model tested on two benchmark suites, not a production agent stack, so the real test is whether these gains hold as labs run MVPO on larger models and messier real-world tasks. If they do, it's a plausible building block for agents that actually finish what they start - reason enough to keep watching this line of work.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →