A new paper argues that diffusion language models can drop one of their core design assumptions: that they need to know exactly how much noise a sequence has at each step in order to predict what comes next.
Researchers behind a paper titled "Does Uniform Discrete Diffusion Need Time?" studied uniform discrete diffusion models (UDMs), a class of language models that generate text by gradually removing noise from garbled sequences instead of predicting one word at a time. These models are normally told explicitly how far along that denoising process they are, a signal known as time conditioning. The authors show that while the theoretically optimal predictor does depend on that signal, in practice, with the finite datasets used to train language models, a corrupted training sequence usually stays much closer to its original clean version than to any other example in the dataset. That makes the time signal largely redundant across most of the diffusion trajectory. When the researchers tested time-agnostic models against time-conditioned ones across multiple datasets and training objectives, the time-agnostic versions matched or beat their time-aware counterparts.
This chips away at a design choice long treated as load-bearing in diffusion-based language modeling, an architecture increasingly pitched as a parallel-generation alternative to autoregressive, GPT-style transformers. If time conditioning is optional, builders can simplify these models without sacrificing performance, a small but concrete gain in a subfield still working out its basic recipe.
The paper notes its own guarantee weakens near the high-noise end of the process, so nobody should strip out time conditioning wholesale. Still, it is one fewer assumption diffusion language models need to defend.