AI/ ai-watermarking · image-generation · ai-safety · research

New Attack Defeats AI Image Watermarks Without Ruining Quality

A new method called LiBRA strips AI watermarks by nudging images in latent space, then checks removal with a statistical test instead of guessing.

A new attack method can erase invisible watermarks from AI-generated images while keeping the pictures intact.

Researchers built a system called LiBRA, short for Latent In-band Bidirectional Removal Attack, that targets the invisible watermarks many AI image generators embed to prove an image was machine-made. Earlier removal attacks tried to flip a watermark's decoded bits to the opposite pattern, but that inverted signal often stayed detectable, and pushing harder just degraded the image for no benefit. LiBRA instead nudges an image's representation in a public autoencoder's latent space until the watermark decoder's confidence drops to a coin-flip, pulling from whichever direction requires the smallest, least damaging change. The team also verifies success with a two-sided statistical test rather than assuming a confidence score alone proves the watermark is gone.

Watermarking is the main technical argument regulators and platforms point to when they promise users can tell real images from AI-generated ones. A method that quietly defeats that signal while preserving image quality undercuts that argument more than earlier attacks, which left visible damage or a detectable fingerprint behind.

Watermarking has always been a lock, not a vault - this is one more reminder that anyone who can see the key can eventually open the door.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →