A new paper finds that the leading method for flagging misuse of open-weight AI models doesn't survive contact with a motivated attacker.
Researchers studied trigger-tag mechanisms, designed to leave a detectable signal when an open-weight model is used for something like generating phishing content, since developers lose the ability to police a model once its weights are public. They split these into token-level trigger-tags, which hide watermark-style signals during text generation, and weight-level trigger-tags, which build backdoor-style associations between a trigger condition and a detectable response. The team built a unified attack framework called Untag and tested representative examples of both types using phishing generation as the case study. Every mechanism they tried failed: simple output transformations or weight edits erased the detectable signal entirely.
Trigger-tags have been pitched as a practical fix for a real problem - once weights are out, nobody can stop someone from fine-tuning a model for scams. This paper argues that fix is mostly theater: the same openness that makes a model useful also hands attackers everything they need to strip out any built-in snitch.
It's the open-weight world rediscovering a lesson from DRM: a security mechanism that runs entirely on hardware or weights the attacker controls is not really a security mechanism.