Researchers found an AI model used to pick the 'best' image from a batch often just goes with whatever picture happens to be first.
The study tested vision-language models, AI systems increasingly used to judge and select among multiple AI-generated images, on 300 prompts tied to specific cultures. Researchers compared each judge's pick against human ratings and against random chance, then reran every decision with the same images shuffled into a different order. A 4-billion-parameter judge barely outperformed a coin flip and lost to CLIP similarity, a simpler, non-judging comparison method. It picked whichever image came first in 49% of trials, well above the 28% expected by chance, and swapping the image order flipped its answer on 60% of prompts.
That matters because these judges are meant to be quality filters, deciding automatically what users and downstream systems see. An 8-billion-parameter judge in the same test showed almost no position bias and beat CLIP outright, meaning the fix isn't simply using any AI judge: it's using one that was actually audited for this failure mode. The researchers also found their proposed safeguard, checking whether a judge's pick survives reordering, only helps catch bad decisions from the biased small model; run on the better judge, it throws out good picks too.
Before any company lets a model pick the best photo for its users, it's worth asking whether the model is grading the image or just grading its position in line.