AI/ generative-ai · diffusion-models · model-alignment · black-box-optimization

A Gradient-Free Way to Align Fast Image and Video Generators

A new method tunes the noise behind one- and few-step generators toward better rewards without gradients, aimed at black-box alignment setups.

A new technique tunes fast image and video generators without ever touching their internals.

Researchers built ZeNOVA, a method that optimizes the random starting noise fed into one-step and few-step generative models, nudging that noise toward outputs that score higher on some external reward, all without computing a single gradient. It combines three gradient-free tricks: annealed soft-value guidance, Langevin dynamics constrained to the sphere that Gaussian noise naturally lives on, and occasional Metropolis-Hastings jumps to escape bad local optima. The team tested it on image and video models and reports it beats other zeroth-order baselines at pushing reward scores higher, more stably and more efficiently.

This matters because the fast generators now shipping in products - the one- and few-step distilled models that make image and video generation near-instant - are hard to align with typical reinforcement-learning tricks, since there is no multi-step trajectory to backpropagate through. Many real-world reward signals, like a proprietary aesthetic scorer or a safety classifier sitting behind an API, are black boxes anyway, so a stable gradient-free alignment method could matter more as production systems lean on sealed reward models researchers cannot differentiate through.

The comparisons here are against other zeroth-order methods, not full gradient-based alignment, so the real test is whether this closes that gap in practice rather than just outscoring other noise-whisperers.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →