A research paper lays out a diffusion-based language model that thinks in smooth curves instead of discrete word swaps, and it keeps pace with standard text generators on math and coding tests.
The model, called Sigma, comes in 3 billion and 8 billion parameter versions. Unlike existing diffusion language models that denoise discrete tokens in one noisy jump, Sigma denoises continuous, Gaussian-corrupted token embeddings along smooth ODE/SDE trajectories, and it learns the shape of that embedding space as it trains. The team warm-started training from existing autoregressive model weights to cut training cost, then leaned on classifier-free guidance and score temperature at inference to keep outputs accurate. Tested against both masked diffusion models and autoregressive baselines, Sigma matched them on GSM8K, Minerva, HumanEval, and MBPP, and held up on harder reasoning sets like MATH-500 and AIME after fine-tuning.
That parity matters less than what the continuous design unlocks. Discrete diffusion models decode in parallel, which is fast, but their jagged, high-dimensional space makes it hard to nudge a generation toward better answers. Sigma's smoother embedding space lets researchers steer the trade-off between answer quality and variety, improves pass@k scores, degrades gracefully when the model is asked to use fewer denoising steps, and compresses more easily into smaller, faster models.
Autoregressive models still rule chatbots and coding tools; diffusion language models have mostly been a speed pitch that struggled on reasoning. Sigma is a research prototype, not a product, and competitive is not better - but it is the first real sign that diffusion text generation can be steered like an image model, not just sped up.