Researchers have found a way to shrink diffusion language models without quietly breaking their ability to do math.
Diffusion language models generate text by gradually unmasking noisy tokens, and that creates a mismatch when you try to compress them. The standard compression technique, low-rank approximation, is usually calibrated on clean, fully visible activations, but the model actually runs on partially masked states during inference. A team describes a technique called Traj-MC that calibrates compression using a Monte Carlo estimate of the trajectory second moment, essentially sampling activations across the actual corruption levels and masking patterns a model sees while generating, rather than just its clean end state. Tested under matched compression budgets, this trajectory-aware calibration reconstructed the model's behavior more accurately across the generation process and preserved substantially more performance on math reasoning benchmarks than standard clean calibration. The code is posted on GitHub.
Diffusion-based language models are a newer alternative to the autoregressive transformers behind most chatbots, and making them smaller and cheaper to run matters if they want to compete on cost. This result suggests that compression tools built with autoregressive assumptions in mind can silently gut a diffusion model's reasoning if applied without accounting for how it actually generates text.
The gains are benchmark-specific and dLLM compression is a young enough field that "preserves reasoning" should be read as a comparative claim against clean calibration, not a guarantee it holds up everywhere else.