AI/ ai · llm-inference · mixture-of-agents · model-routing

Study Measures How Much Compute Each Token Actually Needs

A panel of fifteen language models finds most tokens need far less compute than LLMs spend, with the costliest 10 percent eating most of the budget.

Not every token an AI model generates deserves the same amount of math.

A new paper tests that assumption with a mixture-of-agents setup: fifteen language models of increasing size, drawn from the Qwen, OLMo, and R1-distilled families, each try to reproduce a reference sequence one token at a time, using the correct preceding tokens as context. The smallest model that gets a given token right defines that token's sufficient compute, an upper bound on what the token actually needed. Across three core benchmarks, a tiny 0.5B model alone reproduced 92-95% of tokens correctly. But the hardest 10% of tokens ate up 64-80% of the estimated compute across all three model families.

That gap is the real finding. Techniques like speculative decoding and model routing have always assumed most tokens are cheap and a few are expensive, but nobody had measured it directly. Applying this compute map to routing on the MATH-500 benchmark cut projected latency from 7.59 to 5.12 seconds while slightly improving accuracy versus the best existing routing baseline, and applying it to drafting trimmed draft-token counts by 32.6% for a similar latency gain at matching accuracy.

Worth noting: this is an upper-bound estimate on curated benchmarks, not a deployed system, and sufficient compute measured against a 15-model panel is not the same as the true minimum. Still, it's the first real evidence for what routing and drafting systems have been betting on for years.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →