A new decoding trick lets small language models run faster without needing extra memory for a second model.
Researchers describe BitNest, a speculative decoding framework that speeds up text generation on 7B-8B parameter models by predicting several tokens at once and verifying them in parallel. Most speculative decoding setups need a separate, lower-precision draft model sitting alongside the main one, which eats up memory on phones and other resource-constrained devices. BitNest avoids that by building a low-precision draft directly inside the same weights as the full-precision target model, then layering residual refinement on top to recover full quality. The same progressive-precision trick is applied to the KV cache, extending the savings to longer context windows.
On edge hardware, memory is often the real bottleneck, not raw compute, so folding the draft model into the target's own weights is a cleaner fix than the usual workaround of shrinking or distilling a second model. The results back it up: a 95.2% average acceptance rate for the draft's guesses, and up to 1.61x faster decoding than standard FP16 generation, without much quality loss.
It's an incremental fix rather than a new paradigm, but for anyone trying to run capable models on a laptop or phone, incremental often beats clever-but-bulky.