AI/ ai · model-training · distillation · research

New Distillation Method Cuts AI Teacher Costs by 95 Percent

A new training technique teaches AI models using a fraction of the computing power a teacher model normally needs, without hurting accuracy on math benchmarks.

Researchers have found a way to make a popular AI training technique dramatically cheaper, without giving up accuracy.

The technique is called on-policy distillation: a smaller "student" model tries a task, and a bigger "teacher" model corrects its mistakes token by token. That correction is thorough but expensive, since it means running the teacher model constantly. A new method called Success-Referenced On-Policy Distillation (SR-OPD) cuts that cost by being selective. When a student gets some attempts right and others wrong on the same prompt, SR-OPD uses the successful attempt as a reference point to find exactly where the failed attempt went off track, then sends only that portion to the teacher. Tested across three teacher-student model pairs and six math reasoning benchmarks, it matched the performance of full teacher supervision while using just 3.46 to 5.02 percent of the teacher's token budget.

This matters because teacher compute, not student compute, is usually the bottleneck in distillation-heavy training pipelines. A nearly 20-fold reduction in teacher-token usage means labs with smaller budgets can run the kind of dense supervision that was previously the province of well-funded players. It is an efficiency story, not a capability story, but efficiency is what determines who can actually afford to train competitive models.

The catch: every result here comes from math reasoning benchmarks, a domain where "successful" and "failed" rollouts are unusually easy to tell apart. Whether this selection trick holds up on messier tasks, like open-ended writing or coding, is still an open question.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →