AI/ reinforcement-learning · llm-agents · ai-research · training-stability

New Method Stabilizes Reinforcement Learning for AI Agents

A new algorithm called SORL tackles the gradient instability that causes reinforcement learning to collapse when training AI agents on multi-turn tasks.

Training AI agents to handle multi-step tasks keeps breaking in the same predictable way, and a new paper claims to have found why.

Reinforcement learning methods like PPO and GRPO are the standard tools for teaching large language models to act as agents across multiple turns, such as answering a question through several reasoning steps. Researchers found that in off-policy training, where the model learns from data generated by an earlier version of itself, these methods become unstable and can collapse outright. They trace the problem to two causes: a mismatch between optimizing at the level of individual tokens and the actual turn-by-turn structure of agent interactions, plus noisy gradient updates from imprecise estimates of which actions were actually good. Their proposed framework, SORL, applies importance sampling at the turn level instead of the token level and adds a normalization step that only kicks in when updates get clipped, producing two variants, SO-PPO and SO-GRPO, tested on open-domain QA, multi-hop QA, medical multiple-choice QA, and math reasoning via DAPO-Math-17k validated on AIME-2024.

This matters because agent training has a babysitting problem. Runs degrade unpredictably, and researchers often resort to early stopping or manual tuning just to get a usable checkpoint, which makes scaling agent training expensive and finicky. A fix aimed at the structural mismatch, rather than one more heuristic patch, is the kind of unglamorous work that actually unblocks further progress.

The real test isn't the benchmark numbers in the paper, it's whether other labs training agents stop needing to babysit their own runs.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →