A new post-training recipe is reshuffling the image-generator leaderboard, and the top scorer isn't open-source.
Researchers built a reward system that combines two signals: a preference reward trained on large-scale human judgments of what looks good, and rubric-based rewards that check whether an image actually matches its prompt and resist being gamed. They found that simply averaging the two signals produced worse results, so they designed a composition method that balances preference optimization against rubric satisfaction instead. Applied through reinforcement learning to Flux2dev, the technique added 69 Elo points over the base model on the Arena text-to-image leaderboard. The same recipe applied to Ideogram-4 pushed its score to an Elo of 1223.5, enough to beat every open-source model on the board - but Ideogram-4 itself is a closed, commercial model, not an open-source one.
The real story here isn't which single model won - benchmarks churn constantly - it's that reward design, not more data or bigger models, is driving the gains. Preference scores alone can let a model learn to produce generically pretty images while ignoring the prompt; pairing them with rubric checks on prompt-following is what keeps that shortcut from paying off.
Worth noting: the researchers are grading their own technique against a leaderboard snapshot from one day in September, and they trained the model that came out on top - that's useful signal, not an independent verdict.