AI guardrails that gate tool calls can be fooled with nothing more than a misleading label.
Researchers tested seven open-weight typed decision models, classifiers that read text and return a probability across a short list of allowed options, the kind of component many agent systems use to approve or block a tool call. On jailbreak and prompt-injection screening, allow-or-block accuracy ranged from 36% to 72%, close to a coin flip's 50% baseline. Adding six lines of irrelevant server log text to a tool call request raised one gate's fail-open rate, meaning it wrongly allowed a blocked action, from 0% to 63%. Simply renaming the permissive option, while leaving its definition and the judged text untouched, pushed that rate to between 93% and 100% on the four models that expose option labels in their own input.
That matters because these models are already being deployed as the decision-maker in agent pipelines, not just a filter. Every defense the researchers tried failed, including escalating low-confidence calls to a human reviewer: a decision an attack flipped was no less confident than the one it replaced. The one approach that held up, parsing policy fields into typed values, also showed a plain deterministic rule over those fields reaches 100% accuracy, which makes the model unnecessary.
A gate that opens because someone renamed a button isn't a security layer, it's a formality.