A new training trick lets language models get better at reasoning using far less trial and error.
Training a model to reason with reinforcement learning usually means asking it tons of questions and only rewarding it when it gets the right answer. That reward signal is sparse, so progress is slow. A new method called Goldilocks adds a second neural network, a Selector, that predicts how much a model's answers to a given question will vary. It then prioritizes questions that are neither too easy, where the model already succeeds, nor too hard, where it always fails, feeding those into the standard GRPO training loop. Tested on the OpenMathReasoning and Polaris datasets, the approach matched plain GRPO's results using up to 78% fewer optimization steps.
Curriculum learning, the idea of ordering training examples by difficulty, has existed for years but never scaled cleanly to today's large language models, partly because nobody agrees on what "difficulty" means for a given model at a given moment. Goldilocks sidesteps that by letting the model's own performance define difficulty in real time, adapting as it improves rather than following a fixed schedule.
Still, a 78% efficiency gain on two math-and-reasoning benchmarks is not the same as a 78% cheaper way to build the next flagship model. Call it a promising lab result, not a training strategy anyone has shipped at scale.