Security/ prompt injection · ai security · ai agents · tokenization

One Reserved Token Explains Why Chat Injection Attacks Work

A new preprint finds a single learned token, not clever wording, powers chat-template prompt injection, and the standard fix misses most popular models.

A forged chat-template marker doesn't need clever wording to hijack an AI agent - it just needs to look official to the tokenizer.

Researchers behind a new preprint, posted to arXiv this month as Same Bytes, Different Authority (arXiv:2609.35932), tested how much of a prompt injection's power comes from the model's own control tokens rather than the words around them. A forged marker, like a fake assistant turn, can reach a model either as one reserved control token or as a string of ordinary subwords that decode to identical text, and it's the server, not the attacker, that decides which one gets sent. Swapping the reserved token for its subword equivalent, with the visible text held constant, cut attack success on the InjecAgent benchmark by 39 to 66 percentage points across three of four open-weight model families, and the effect carried over to multi-turn tasks in AgentDojo. Qwen3-8B was the outlier: it caught the forged turn through its own reasoning even without the special token, until the researchers suppressed that reasoning step, which widened the gap to 50 points.

That points to a narrow, fixable failure: the injected instruction's authority lives almost entirely in one learned vector tied to the marker's position, not the surrounding text, and instruction tuning makes models trust that vector more, not less. But the paper also finds the standard fix - forcing tokenizers to treat special tokens as plain subwords - only covers tokens a given configuration bothers to declare special. In 33 of 67 tokenizer setups, spanning 255 of the 400 most-downloaded chat models on Hugging Face, the tool-protocol tokens agents use to read tool output are left untouched, so the attack keeps working through that channel.

It's a reminder that most prompt-injection defenses have been policing the words attackers use, when the model may be listening to something else entirely: the punctuation of its own training.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →