Researchers have found a cheaper way to teach language models to draft multiple tokens per step, without baking that skill in from day one.
Multi-token prediction (MTP) lets a model draft several next tokens per forward pass, then a quick verification check confirms they match what the full model would have generated anyway. That speeds up generation without changing the model's outputs. Until now, every open MTP system, including MiMo-7B, DeepSeek-V3 and Qwen3, trained its drafting heads jointly with the base model across tens of trillions of pretraining tokens. The new paper instead starts with a frozen Qwen3-8B model, bolts on three chained MTP heads, and post-trains them on roughly 2.5 billion tokens of chain-of-thought data using plain cross-entropy, then shows that recipe matches or exceeds the speedups MiMo-7B achieved through full joint training on math, coding and knowledge benchmarks, using 1,000 to 10,000 times fewer MTP-specific training tokens.
The result suggests the expensive part of MTP speedups was never the sheer volume of training data, just the absence of a targeted recipe. The paper also adds a relaxed verification rule that tolerates a bounded drift from the base model's output distribution, lifting speedups another 12 to 16 percent per benchmark, plus an adaptive controller that adjusts how many drafting heads run at inference time, recovering up to 11 to 14 percent of lost speed when draft length is capped.
None of this makes the underlying model smarter. It just makes serving it faster and cheaper to retrofit, which matters more to anyone paying inference bills than to anyone chasing benchmark headlines.