AI/ ai · llm-training · research

Teaching AI With Instructions Instead of Answers Works Better

A new study finds short, general instructions train AI models better than exact answers, especially outside their training distribution.

A new study finds that a short set of instructions can train an AI model better than handing it the exact correct answer.

Researchers tested a method called on-policy context distillation, where a teacher model sees privileged information and a student model learns to match its outputs. The default privilege has usually been the gold, instance-specific answer. The team swapped that out for short, general instructions, just a few sentences, that flag the mistakes students commonly make, then tested the approach on logic-based benchmarks ProverQA, ProofWriter, and ProntoQA using Qwen3-Thinking and Olmo3-Thinking models. In 7 of 8 experiments, the instruction-based approach beat gold answers on out-of-distribution accuracy by 4 to 17 points, while matching gold on in-distribution tasks. In a separate autoformalization task, a single formatting instruction beat gold by an even wider margin.

That gap matters because gold answers are the default privilege in distillation pipelines, and they are expensive to produce at scale since each one is tailored to a specific example. A short instruction, by contrast, costs little to write and applies uniformly across an entire training set, source domain and target domain alike. If it also generalizes better, that is a cheaper and more robust signal than the industry's go-to method for teaching smaller models to imitate bigger ones.

The catch: this is tested on three logic-puzzle datasets and one formalization task, not on messier real-world reasoning. It is also still an unpublished arXiv preprint, not peer-reviewed. Worth watching, not yet worth rewriting your fine-tuning pipeline around.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →