A new theoretical framework explains why identical-cost tweaks to a diffusion model's sampling process can produce wildly different images.
The researchers built a framework that pairs dynamical analysis of the sampling process with information theory to track how small perturbations ripple through to the final output. They measure perturbation strength using KL divergence between a perturbed trajectory and the original, a metric they call path cost, and show it only sets an upper bound on how much the output changes rather than predicting the actual change. Testing on pretrained diffusion models at equal path cost revealed that sensitivity varies sharply by sampling stage and by spatial frequency, meaning two tweaks that cost the same can produce very different results. The team then pointed the framework at cache-based acceleration, a common speed trick, and used it to identify exactly which sampling intervals cause the largest image errors when caching kicks in.
Diffusion models, the engines behind most image generators, are routinely accelerated with shortcuts like caching and step-skipping, usually justified by eyeballing the output afterward. This gives developers a way to predict which shortcuts will quietly degrade quality before shipping them, rather than finding out from a blurry or distorted image later.
It is pure theory for now, with no new model or product attached, but if adopted it could become a standard pre-flight check for anyone selling a faster diffusion pipeline.