AI/ clip · image-quality-metrics · adversarial-attacks · ai-security

Researchers show how to fool CLIP-based image scores

A new study shows CLIP-based quality scores like CLIPscore can be gamed with subtle pixel tweaks, and offers a grayscale trick to catch the fakes.

A new technique can trick CLIP's widely used image-quality scorer into giving top marks to images that do not actually match their prompts.

Researchers built a framework called FoCLIP that uses stochastic gradient descent to construct these "fooling" images. It combines three pieces: a feature-alignment module that narrows the gap between image and text representations, a score-distribution-balance module, and a pixel-guard regularization step that keeps the output looking like an ordinary photo. Tested on ten famous art prompts and subsets of ImageNet, the resulting images scored significantly higher on CLIPscore while still looking visually normal to a person, even when they were semantically unrecognizable or mismatched to the prompt that supposedly produced them.

CLIPscore gets used across the industry to judge how well generated images match text prompts and to help spot manipulated images, so a method that inflates those scores without any visible tampering is a real problem. It suggests the alignment that makes CLIP-based tools useful is also what makes them easy to exploit, and the issue likely extends to other metrics built the same way.

The researchers also found a tell: converting the fooling images to grayscale strips out whatever hidden signal lets them cheat the score, which they turned into a color-channel-sensitivity detector that catches these fakes with 91% accuracy. That is a useful patch, not a fix, and a reminder that a metric built on "alignment" was never as tamper-proof as it looked.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →