Researchers have built a way to measure, mathematically, whether an AI-generated visual metaphor actually makes sense.
The system, called PRISM, applies a framework from category theory to represent an analogy as an explicit relational map rather than just a picture. A vision-language model builds that map, and a 'pullback score' checks how well the image's relationships mirror the ones in the original concept. On the AnaloBench benchmark, picking the best analogy by pullback score alone was correct 82.5% of the time. PRISM then uses that same score as feedback to iteratively redraw the image, nudging it toward deeper structural accuracy instead of just surface polish.
Most AI image generators chase visual plausibility, not conceptual correctness - a metaphor can look polished and still miss the point entirely. PRISM is one of the first attempts to put a real number on that gap, which matters for anything that leans on AI to explain complex ideas through imagery, from textbooks to data dashboards.
Human reviewers only preferred the refined images in 57.65% of comparisons, and the researchers' own qualitative review found the refinement step sometimes just piles on visual clutter instead of deeper meaning - a reminder that a good score and a good metaphor are not automatically the same thing.