AI/ ai · machine-learning · model-training · llm-research

Training AI on Fewer Tokens Now Beats Using More of Them

DIAL-OPD trains student AI models on just 40 percent of supervision tokens and still beats full-token training on math reasoning benchmarks.

A new token-filtering trick lets AI models learn more by training on less.

Researchers introduced DIAL-OPD, a method for on-policy distillation, the process of training a smaller student model by having it generate its own outputs and then scoring those outputs token by token against a larger teacher model's predictions. The team found that giving the student fewer, better-chosen tokens beats giving it every token, a reversal of the usual assumption that more supervision is better. The problem was that standard methods for picking which tokens matter rely on disagreement between teacher and student probabilities, but ignore scale: tokens both models consider unlikely, called low-low tokens, generated outsized reward signals that confused training. DIAL-OPD fixes this by weighting each token's reward by the combined log-probability of teacher and student, controlled by a tunable parameter called beta, then keeping only the top-scoring tokens.

Tested across four teacher-student model pairs and seven math reasoning benchmarks, DIAL-OPD retained only 40 percent of tokens yet lifted mean accuracy by up to 5.25 percentage points over standard on-policy distillation and doubled AIME25 pass rates from roughly 13 percent to 27 percent. A 4-billion-parameter teacher using DIAL-OPD even beat an 8-billion-parameter teacher using full-token training, suggesting smarter token selection can substitute for a bigger, costlier teacher model.

The results so far are confined to math benchmarks and a handful of model pairs, so it is still an open question whether this kind of token triage holds up on messier, harder-to-verify tasks like writing or coding.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →