A new training trick teaches AI agents to learn from their best moments, not just their final score.
Researchers describe ProVer, a reinforcement learning framework detailed in an arXiv preprint (arXiv:2609.36178, posted September 30, 2026). The paper targets a weakness in Group Relative Policy Optimization (GRPO), a popular method for training large language model agents: GRPO credits every token in a trajectory equally, whether or not that token mattered. ProVer instead uses an 'agentic judge' to compare successful and failed runs and flag a segment of the trajectory that likely caused the split outcome, then checks that guess by resampling the policy before and after the segment and comparing success rates. Only segments that actually move the needle get extra credit during training.
The approach matters because credit assignment is one of the quiet bottlenecks in agentic reinforcement learning: models that cannot tell which mid-task decision led to success waste compute reinforcing noise instead of skill. On the ALFWorld, WebShop, and SearchQA benchmarks, the preprint reports ProVer beating plain GRPO by 9.91% on a 2B-parameter Qwen3.5 model and 7.12% on a 4B version, with only modest extra generation overhead from the verification step.
It is a narrow, technical fix rather than a new architecture, and the gains come from the paper's own authors, not an independent replication. But if it holds up, it is the kind of unglamorous efficiency tweak that tends to show up in production training pipelines long before it shows up in flashy agent demos.