AI/ rlvr · reinforcement learning · llm training · ai research

AI Training Difficulty Labels Are Less Reliable Than Thought

A new analysis finds the hard prompts AI models supposedly cannot learn from actually do improve, but the labels used to call them hard are unreliable.

A new paper says a weird training glitch in AI reasoning models is real, but the way researchers identify it is shakier than assumed.

The method is called RLVR, reinforcement learning with verifiable rewards, and it is a common way to train reasoning models after initial training. Researchers had noticed that some especially difficult prompts barely improved during RLVR training even when the model occasionally got them right, a pattern dubbed unlearnability. This paper retests that finding and shows those hard prompts do improve, just at about a third of the rate of easier ones. It also finds that the difficulty labels used to sort prompts into easy and hard buckets are far less reproducible than expected, since they come from a small number of sampled responses, and averaging across different sampling runs can shift which prompts count as hard rather than just reducing noise. A separate explanation for unlearnability, based on how similar gradients look across prompts, partly turns out to be a side effect of hard prompts simply yielding fewer correct answers to calculate gradients from.

That is a meaningful correction for anyone building curricula or benchmarks around RLVR difficulty tiers. A measurement that moves depending on which random seed you used is not a measurement you want informing training decisions or papers' causal claims.

The slow-learning effect holds up. The confidence researchers had in measuring and explaining it does not.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →