A new paper argues that the way AI image generators pass information between layers has been quietly wasting training time, and proposes a fix.
Diffusion Transformers quietly became the default architecture for AI image and video generation, inheriting their residual stream, the mechanism that shuttles information between layers, wholesale from the original Transformer. Researchers diagnosed three problems with that inherited design: signal magnitude balloons as the network gets deeper, gradients used in training decay sharply on the way back, and nearby layers end up doing redundant work. Their fix, called Diffusion-Adaptive Routing (DAR), replaces fixed residual addition with a learnable mechanism that weighs outputs from earlier layers differently depending on the denoising timestep. Tested on ImageNet at 256x256 resolution, DAR improved a SiT-XL/2 model's FID score from 9.67 to 7.56 and matched the original model's final quality using 8.75 times fewer training iterations, with gains doubling when combined with another efficiency method called REPA.
That 8.75x figure matters because training these models is expensive, and most efficiency research chases savings through better data, better objectives, or smarter attention, not the plumbing that moves information between layers. DAR is pitched as compatible with those other tricks rather than competing with them, which is the more interesting claim: it suggests cross-layer routing is a design axis that has gone unexamined simply because everyone inherited the same default.
Still, this is a benchmark result on ImageNet, not a shipped product, and FID gains on a research benchmark do not always survive contact with messier, larger-scale training runs.