New research says a popular way to train AI models to reason without a built-in answer-checker has a blind spot: the longer the model thinks, the less useful its own reward signal gets.
The method is called verifier-free reinforcement learning, used when there's no outside checker to confirm whether an answer is correct. Instead, the system scores the model based on how likely the correct answer looks given the reasoning trace it produced. Researchers found that as traces get longer, that likelihood score collapses into a narrow, nearly identical range no matter the output, a pattern they call the Posterior Concentration Phenomenon. In common training setups like GRPO, that flattening leaves almost no usable signal, making optimization unstable and wasteful. Their proposed fix, RLCPR, filters out rollouts likely to hit this trap before generation and penalizes traces that run needlessly long once the reward signal flattens.
This matters because verifier-free RL is one of the few scalable ways to train reasoning on tasks with no ground truth to check against, open-ended writing, judgment calls, general problem solving. If the reward signal quietly breaks down on exactly the long, hard problems labs care most about, that undercuts a technique many are counting on as reasoning chains get longer, not shorter.
The reported gains, up to 4.0 percent over the strongest existing verifier-free baseline on six of seven benchmarks, are incremental rather than a breakthrough. Naming the failure mode clearly may end up mattering more than the specific fix.