AI/ ai · image-generation · reinforcement-learning · benchmarks

New Post-Training Recipe Lifts AI Image Model Rankings

Blending human preference scores with rubric checks lifted two image models' Arena rankings, though top scorer Ideogram-4 is proprietary, not open-source.

A new post-training recipe is reshuffling the image-generator leaderboard, and the top scorer isn't open-source.

Researchers built a reward system that combines two signals: a preference reward trained on large-scale human judgments of what looks good, and rubric-based rewards that check whether an image actually matches its prompt and resist being gamed. They found that simply averaging the two signals produced worse results, so they designed a composition method that balances preference optimization against rubric satisfaction instead. Applied through reinforcement learning to Flux2dev, the technique added 69 Elo points over the base model on the Arena text-to-image leaderboard. The same recipe applied to Ideogram-4 pushed its score to an Elo of 1223.5, enough to beat every open-source model on the board - but Ideogram-4 itself is a closed, commercial model, not an open-source one.

The real story here isn't which single model won - benchmarks churn constantly - it's that reward design, not more data or bigger models, is driving the gains. Preference scores alone can let a model learn to produce generically pretty images while ignoring the prompt; pairing them with rubric checks on prompt-following is what keeps that shortcut from paying off.

Worth noting: the researchers are grading their own technique against a leaderboard snapshot from one day in September, and they trained the model that came out on top - that's useful signal, not an independent verdict.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →