AI/ diffusion models · reinforcement learning · generative ai · ai research

Researchers Find Better Way to Tune AI Image Generators

A new paper argues a popular reward-tuning fix for diffusion models actually hurts the diversity it was supposed to protect, and proposes a fix for the fix.

A new paper claims to fix a tradeoff that has been quietly hurting AI image generators: training them to chase a reward score tends to kill variety and sometimes quality too.

Researchers studied reinforcement learning methods used to fine-tune diffusion models after their initial training, including a technique called Denoising Diffusion Policy Optimization, or DDPO, which nudges a model's step-by-step denoising process toward a reward signal. The paper shows mathematically that updating only the final steps of that denoising process, a trick used in earlier work, actually hurts diversity, the opposite of what that earlier work concluded. Building on the corrected theory, the authors propose a new training approach called incremental Feynman-Kac training. Across three tasks and a set of ablation experiments, the method outperforms existing diffusion policy optimization approaches on both alignment and diversity.

Reward-based fine-tuning is the standard route for getting diffusion models to produce outputs that score well with human raters or automated classifiers. Push that optimization too hard and you get the generative-AI version of regression to the mean: safe, repetitive outputs that technically score well but look alike. A method that holds onto variety while still hitting the reward target matters for anyone shipping these models at scale, not just a theoretical fix.

For now it is one arXiv preprint validated on three tasks, so treat the best yet tradeoffs claim as one to watch, not a verdict.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →