New research finds that image-generating AI models collapse ambiguous words into a single meaning far more aggressively than text models do.
Researchers tested 17 text-to-image models and 15 text-generation models using polysemous words like "bank" or "palm" with no surrounding context to fix a single sense, then measured how many distinct meanings each model produced across many samples. Using a normalized entropy score, where higher numbers mean more variety, image models scored 0.10 and text models scored 0.25 - both well below the 0.47 score researchers got when they asked people to imagine the same words. When the researchers instead asked models to predict how often they would generate each possible meaning, the models claimed they would be far more diverse than they actually turned out to be.
The gap matters for anyone stitching text and image models together, since the ambiguity a prompt seems to carry in text can quietly narrow the moment a model renders it as a picture. It is also a bias problem in disguise: a model that defaults to one sense of a word defaults to one visual stereotype, over and over, across millions of generations.
Text models come out looking better only by comparison - a 0.25 score next to a human's 0.47 still means a model is picking one meaning for you most of the time, not preserving the ambiguity it claims to understand.