AI/ ai · diffusion-models · kv-cache · inference-efficiency

Researchers Cut Memory Use in Diffusion AI Models by 20x

A new training-free technique trims the memory diffusion language models need by up to 20 times, with only a modest accuracy cost.

A new technique lets diffusion-based language models hang onto far less memory while they generate text.

Researchers built a method called MaskAhead that decides, without any extra training, which parts of a model's short-term memory, the key-value cache, are worth keeping and which can be discarded. Block diffusion models generate text by filling in whole chunks, or blocks, at once rather than one word at a time, but they still have to track every token they've already produced, which eats up memory and slows generation. MaskAhead uses the model's own masking signals to rank which cached entries matter most for the current block, and which ones can be evicted before the next one starts. Tested on three models, Fast-dLLM-v2, DreamReasoner, and LLaDA2.0-mini, across long-document question answering and needle-in-a-haystack retrieval, it cut memory use by 9.5 times on average for long-prompt QA, with only a 1.2-point drop in F1 accuracy. A more aggressive, low-bit version, Q-MaskAhead, pushed the reduction to 20.1 times for a 2.3-point accuracy hit.

Memory is the bottleneck that decides how long a prompt a model can handle and how many users a server can serve at once, so a 9x to 20x cut is a real capacity gain, not a marginal tweak. It also suggests block diffusion models, still a newer alternative to standard autoregressive transformers, can borrow the same cache-pruning tricks that have already sped up conventional LLMs.

The paper reports a modest 1.23x end-to-end speedup in systems tests, a reminder that memory savings on paper do not always translate one-for-one into faster responses.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →