AI/ ai watermarking · image generation · ai security · diffusion models

New Attack Strips Watermarks From AI Generated Images

Researchers show a frequency based attack can erase or weaken invisible AI image watermarks while keeping pictures looking normal.

Researchers just showed a fast way to scrub the invisible watermarks meant to prove an image came from an AI model.

The technique, called Latent Frequency Masking, targets the hidden frequency data diffusion models bake into an image's latent representation as a watermark. Instead of brute-forcing pixels, the attack swaps out selected Fourier coefficients in that latent space, either with Gaussian noise for speed or with diffusion-regenerated values for better image quality. The team tested it against six diffusion watermarking methods using images generated from DiffusionDB and MS-COCO prompts. The attack removed or significantly weakened several of those watermarks while keeping the images looking normal and running faster than existing removal methods.

This matters because invisible watermarking is the main technical bet regulators and AI labs are making to track AI-generated content, from deepfake detection to copyright disputes. If a watermark can be stripped cheaply and quietly, the labeling promise behind tools like this falls apart fast.

Watermarking has always been a cat-and-mouse game, and this is another reminder the mouse is doing fine. The researchers frame their own attack as a call for better robustness testing, not a how-to guide, but the gap between watermark marketing and watermark reality just got wider.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →