A new technique called Best-of-N Guidance nudges AI image generators toward better outputs while they are still generating, not just after the fact.
Diffusion models are often steered with a reward model after the fact: generate a batch of N images, score them, and keep the best one. That is Best-of-N sampling, and it only acts at the finish line - it does not change how the earlier denoising steps unfold, so it mostly helps single best-pick outputs rather than the overall batch. The new method, called Best-of-N Guidance (BoNG), builds reward selection into the denoising process itself. It runs online best-of-N selection across a population of denoising particles and uses the current frontrunner as a guidance signal for the rest, pushing the whole batch toward higher-reward regions as it generates.
That matters because it fixes a real limitation of Best-of-N: it improves the average output, not just the cherry-picked one, which is the difference between a demo trick and a method useful when you need several good images at once. In 36 head-to-head tests against Best-of-N and a particle-based rival called SMC, BoNG won 29 of them, about 81 percent, and against the leading sample-based guidance method it hit a 1.3x higher ImageReward score while running 1.6x faster.
Reward-guided generation has quietly become the workhorse behind a lot of 'aligned' AI image output; BoNG's pitch is less a new idea than a sharper version of one everyone already leans on, worth watching once someone runs it outside a benchmark table.