Not every token an AI model generates deserves the same amount of math.
A new paper tests that assumption with a mixture-of-agents setup: fifteen language models of increasing size, drawn from the Qwen, OLMo, and R1-distilled families, each try to reproduce a reference sequence one token at a time, using the correct preceding tokens as context. The smallest model that gets a given token right defines that token's sufficient compute, an upper bound on what the token actually needed. Across three core benchmarks, a tiny 0.5B model alone reproduced 92-95% of tokens correctly. But the hardest 10% of tokens ate up 64-80% of the estimated compute across all three model families.
That gap is the real finding. Techniques like speculative decoding and model routing have always assumed most tokens are cheap and a few are expensive, but nobody had measured it directly. Applying this compute map to routing on the MATH-500 benchmark cut projected latency from 7.59 to 5.12 seconds while slightly improving accuracy versus the best existing routing baseline, and applying it to drafting trimmed draft-token counts by 32.6% for a similar latency gain at matching accuracy.
Worth noting: this is an upper-bound estimate on curated benchmarks, not a deployed system, and sufficient compute measured against a 15-model panel is not the same as the true minimum. Still, it's the first real evidence for what routing and drafting systems have been betting on for years.