Researchers have published a new way to train AI image and video generators that produce results in just a handful of steps, without the usual computational baggage.
The method, called DMAD, tackles a problem with an existing technique known as Distribution Matching Distillation, or DMD. DMD trains a fast student model by comparing it to a slow but high-quality teacher model, but it needs a separate auxiliary model running alongside the student just to make that comparison, adding memory and compute overhead. DMAD swaps that auxiliary model for a simpler classifier: two discriminator heads that just try to tell real images, teacher-generated images, and student-generated images apart. The researchers say this swap mathematically recovers the same training signal as DMD, minus the extra model, and they add a reweighting trick that focuses training where the student is weakest.
On paper, the gains are real: a one-step image generator scoring 1.04 on a standard image-quality benchmark, where lower is better, on ImageNet, a four-step version of Stable Diffusion XL beating slower multi-step teachers, and a four-step video model that human evaluators preferred over two competing speed-up methods, DMD2 and rCM, more than three-quarters of the time. For an industry racing to make generative AI fast and cheap enough for real-time products, think instant AI video drafts or on-device image generation, cutting both the number of steps and the extra model overhead is a meaningful win, not just a benchmark flex.
The numbers come from the authors' own tests against their own benchmarks, so treat the wins over DMD2 and rCM as a strong opening argument rather than a verdict - this is a preprint, and claims of being state of the art in diffusion distillation have a short shelf life.