Security/ ai-agents · watermarking · provenance · ai-security

A Watermark for AI Agents That Survives Paraphrasing

Researchers built an AI agent watermark that survives paraphrasing and resists forgery, unlike prior systems that broke under simple rewording.

A new watermarking scheme flags which AI agent produced a given trajectory, even after someone rewrites the tool names and instructions around it.

Researchers describe Semantic Behavioral Watermarking (SBW), which embeds an owner ID in an LLM agent's action choices by watermarking semantic clusters of actions rather than exact action symbols, conditioned on the agent's history. It adds keyed, collision-resistant binning so an attacker cannot predict which bucket a watermark lands in without the key. The team tested SBW on five agent models from four vendors, ranging from 3B to 14B parameters, using three different encoders, on two benchmarks: ToolBench and ALFWorld. On ToolBench, detection after rewriting held at 0.49-0.66 for SBW versus 0.05-0.17 for exact-symbol matching at a 1% false-positive rate; on ALFWorld the gap widened to 0.92-0.97 versus 0.00-0.01. Keyed binning also cut a forgery attack's success rate from 100% down to the false-positive floor.

That matters because every prior agent watermark assumed the exact wording around an action stays fixed: rename a tool or paraphrase an instruction, and detection collapsed. One earlier scheme, AgentMark, reportedly fell to 16.8% bit recovery under paraphrasing alone. SBW is also apparently the first of these systems to test whether an attacker can forge someone else's watermark, not just erase one, a question that matters more as agent trajectories start getting cited as evidence of who built or ran what.

The authors are upfront about where the guarantee stops: an attacker who copies a victim's own steps and chain-replays them still verifies at 0.76-0.98 across all five models, a case they flag as unsolved rather than claim to fix. Watermarking AI agents, in other words, is still an arms race; this result just moves the goalposts, not the finish line.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →