A new training tweak called PEPO scores how much each individual word in an AI's answer actually mattered, instead of crediting every word in a response equally.
The method targets reinforcement learning with verifiable rewards (RLVR), the technique used to train reasoning models by rewarding correct final answers. The standard approach, GRPO, hands out that reward credit evenly across every token in a response, even though some words clearly matter more than others. Some newer methods try to fix this by measuring each token's "entropy" - essentially how uncertain the model was when it picked that word - but they calculate it across an entire batch of prompts, which muddles a token's real importance with how hard the overall question was. PEPO instead measures entropy relative to a token's immediate neighbors, which the researchers say removes that confound. Across three open models - Qwen3-1.7B, Qwen3-4B, and Llama-3.2-3B-Instruct - it beat both GRPO and prior entropy-based methods on math reasoning.
Credit assignment, or figuring out which part of a long answer actually earned the reward, is one of the fiddlier unsolved problems in training reasoning models, and most setups still use the bluntest possible version. A fix this narrow and well-isolated is the kind of plumbing improvement that could raise reasoning scores across many existing RLVR pipelines without needing bigger models or more compute. The paper's claim that it also rescues single-stream RL training, where the batch-based version fails outright, suggests this is a structural correction, not a tuning trick that just happens to win one benchmark.
Still, this is one arXiv paper tested on three small-to-mid-size open models. A clever engineering fix is not the same as proof it holds up at frontier scale.