Most AI image generators struggle to draw what you ask them to leave out.
Researchers built NegT2IBench, a benchmark of 4,800 prompts testing whether text-to-image models can honor negated requests, like "a non-red cup," not just affirmative ones. The prompts span two attribute types and four relation categories, with positive and negated requirements each varying from zero to two per prompt, letting the team isolate negation difficulty from plain prompt complexity. Instead of relying on costly vision-language judges, they built a detector-based scorer that matches the agreement of VLM judges up to 30 times larger while using a fraction of the GPU memory, validated against three-annotator human labels on 600 images. Running it across eleven text-to-image models and 211,200 generated images, they found nine of the eleven models score worse on a single negated statement than on a single positive one.
The sharpest failures showed up in color - models told to avoid a color render it anyway - while spatial relationships like proximity were barely affected. Worse, 41.5% of failed statements did not just miss the mark; they generated exactly what the prompt forbade. That distinction matters because current compositional benchmarks lump "did it draw the right thing" and "did it avoid the wrong thing" into a single score, burying a specific and fixable failure mode.
If an image generator can draw a cat but cannot reliably avoid drawing one, the industry's compositional benchmarks have been grading on an incomplete rubric.