AI/ ai · diffusion-transformers · image-generation · caching

AutoTarget Chooses What to Cache in Speedy Image AI

A new technique tests how much error each cached tensor introduces and keeps the one that distorts images least, speeding up AI image and video generation.

A new calibration trick tells AI image and video generators exactly which shortcut to take without mangling the picture.

Diffusion Transformers power many of today's top image and video generators, but each output requires running the model many times, which is slow and expensive. Developers have sped this up two ways: distillation, which cuts the number of steps, and caching, which skips some of those steps by reusing a tensor computed earlier. Researchers introduce AutoTarget, which figures out which tensor is safest to reuse for a specific model, solver, and reuse schedule. It runs a handful of passes without caching, measures how much error each candidate tensor would introduce if reused, and picks whichever candidate has the lowest measured error.

This matters because distillation and caching fight each other - fewer steps means each skipped step covers more ground, amplifying whatever error caching introduces. Picking the wrong tensor to reuse can degrade quality in ways that only show up after the fact. The researchers found the best tensor to cache changes depending on the model, resolution, and solver, meaning a one-size-fits-all caching rule was probably costing quality for free.

On PixArt-LCM and FLUX.1-schnell, AutoTarget's calibration ranking lined up with results from held-out cached runs - a reassuring sign for a field with no shortage of caching tricks that don't hold up outside their original test case.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →