Security/ ai-generated images · deepfake detection · adversarial attacks · image provenance

Study Finds Hard Limits on Detecting AI-Generated Images

New research sets a mathematical ceiling on pixel-based image verification and shows public AI detectors leak enough to be fooled.

There's a mathematical limit to how well pixels alone can prove an image's origin, and the study says today's detectors fall well short of it.

The researchers frame provenance detection as a verification problem: does this image match its claimed source, human or AI, even after someone edits it before you check? They calculate the best possible outcome for any verifier, a gap that depends only on the statistical distance between the real-image and AI-image distributions after adversarial edits, not on the detector's design. They also show that an attacker who can roughly mimic how a public verifier scores images can build a stand-in attack that gets within about twice that error of fooling the real system. In tests, public CLIP-based verifiers broke down under targeted pixel attacks, and a ResNet-18 model let some AI images pass as real.

That second finding is the uncomfortable one: any verifier that reveals its scores, even partially, gives attackers a blueprint for beating it. That matters for provenance systems and moderation tools moving toward score transparency for auditability. Requiring detectors to answer "not sure" instead of forcing a real-or-fake call cut down successful attacks in testing, but the authors are careful to say that isn't proof of robustness.

Passing a batch of test attacks is not the same as being unbeatable, and that gap between measured and provable robustness is exactly where deepfake detection claims tend to get oversold.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →