A new technique tunes fast image and video generators without ever touching their internals.
Researchers built ZeNOVA, a method that optimizes the random starting noise fed into one-step and few-step generative models, nudging that noise toward outputs that score higher on some external reward, all without computing a single gradient. It combines three gradient-free tricks: annealed soft-value guidance, Langevin dynamics constrained to the sphere that Gaussian noise naturally lives on, and occasional Metropolis-Hastings jumps to escape bad local optima. The team tested it on image and video models and reports it beats other zeroth-order baselines at pushing reward scores higher, more stably and more efficiently.
This matters because the fast generators now shipping in products - the one- and few-step distilled models that make image and video generation near-instant - are hard to align with typical reinforcement-learning tricks, since there is no multi-step trajectory to backpropagate through. Many real-world reward signals, like a proprietary aesthetic scorer or a safety classifier sitting behind an API, are black boxes anyway, so a stable gradient-free alignment method could matter more as production systems lean on sealed reward models researchers cannot differentiate through.
The comparisons here are against other zeroth-order methods, not full gradient-based alignment, so the real test is whether this closes that gap in practice rather than just outscoring other noise-whisperers.