Diffusion language models just got faster by cutting out work they didn't need to redo.
A research team describes SpecFold, a technique for speeding up multi-branch speculative decoding in diffusion large language models (DLLMs). DLLMs generate text by iteratively denoising blocks, and speculative decoding speeds that up by checking a main branch plus several draft branches in a single forward pass. The team found that those draft branches were quietly re-running nearly identical computation, since each one inherits most tokens from its parent and only unmasks a handful of new positions. SpecFold selectively reuses that shared computation, using token-level residual gating with what the researchers call folded attention and FFN, then relies on a custom Triton kernel to turn the savings into real throughput. Across two model families, five models, and five benchmarks, it hit up to 1.64x the throughput of a prior method called Spiffy and up to 1.99x over standard decoding, with no meaningful drop in task performance.
Most DLLM speedups to date have targeted redundancy between denoising steps over time. SpecFold goes after a different kind of waste: redundancy between branches within the same step, a gap that existing caching tricks don't touch. Because the authors say it's compatible with those existing methods, it looks less like a one-off trick and more like a layer other optimizations can stack on.
Diffusion models have been pitched as a faster alternative to standard autoregressive transformers; unglamorous fixes like this one, which just stop the model from redoing work it already did, are what actually decide whether that pitch survives outside a benchmark table.