A new technique lets flow-based image and video generators skip steps without retraining, cutting generation time by more than half.
Researchers describe COFLOW, a method that adaptively picks how many steps a flow-matching model needs for each individual generation, based on the prompt itself. Flow matching is the technique behind many modern visual generators: it builds an image or video by running a model through a series of small updates, and more steps generally means higher quality but slower output. COFLOW is trained online using an unsupervised reward that weighs speed against fidelity, and it plugs into existing models without retraining them. The team reports over 2.5x speedup across both image and video generation while preserving perceptual and semantic quality, backed by a theoretical error bound for the step-reduction approach.
The pitch here is not a new generative model, it is a tax cut on the ones that already exist. Most prior speedup methods force a tradeoff: retrain the model, accept worse output, or use a fixed step count that ignores the fact that a simple prompt needs less computation than a complex one. By adapting step count per input and skipping retraining, COFLOW targets the actual bottleneck in deployed systems, inference cost, rather than training cost, which is what most efficiency research still optimizes for.
If the speedup holds up outside the paper's benchmarks, it is the kind of unglamorous change that actually shows up in your cloud bill.