A new training technique lets groups of language models get smarter by picking their own homework.
Researchers built Stackelberg Alignment, a framework where an algorithm called EXP3 acts as a 'leader' that decides which instructions a pool of models should practice on, based on how hard each instruction is and how much the models' answers still differ. The models themselves act as 'followers': they answer the selected prompts, judge each other's responses using an Elo-style reputation system, and learn from those judgments through standard preference-tuning methods, DPO or GRPO. The team tested the setup across three different model pools and 12 benchmarks covering reasoning, code, science, instruction-following, and general knowledge. Stackelberg Alignment beat the strongest existing training-time method by up to 7.4%, and beat static inference-time setups, where models are not actively re-trained, by 12-25%.
The real contribution here is not the models but the curriculum. Multi-model self-improvement schemes often burn compute re-running instructions long after the models agree on the answer. By letting an adaptive leader redirect practice toward instructions that still produce disagreement, the system squeezes more signal out of the same training budget.
The 7.4% gain over other training methods is the number that actually matters, and it is single digits. The flashier 12-25% figure is a comparison against inference-time tricks that were never designed to improve through training in the first place, not a measure of how far this method has pushed the field.