Security/ ai-security · llm-agents · benchmarks · evaluation

Security Benchmark Scores Shift When You Rename a Tool

Renaming a tool in an AI agent benchmark can swing attack success rates by double digits, raising doubts about what these security scores actually measure.

A new study finds that renaming a tool - with nothing else about the task changed - can swing a security benchmark's attack success rate by double digits.

Researchers tested this by holding everything about an attack scenario fixed, including the task, the harmful action, the security policy, and the environment, while changing only how a tool's name was presented to the agent. On Agent Security Bench, swapping threat-related tool names for neutral ones raised the committed attack success rate by 11.67 percentage points on GPT-5-mini and 13.21 points on Claude Haiku 4.5. On MCPTox, giving a neutral tool an explicit threat-related name instead lowered attack success by 11.00 points on GPT-5-mini and 4.11 points on Claude Haiku 4.5 - and a follow-up test found that a neutral name merely matched for length and token count reproduced 8.54 of those 11.00 points, meaning word choice alone, not threat content, drove most of the shift. On AgentDojo, adding threat-related wording to the attack-relevant tool barely moved attack success (0.50 points on GPT-4o-mini) but cut benign task performance by 5.36 points on tasks that needed that tool.

This matters because agent security leaderboards increasingly shape vendor roadmaps, and a lab chasing a lower attack success rate could rename a function as easily as patch a real vulnerability. The MCPTox control experiment is the sharpest finding here: it shows scores can swing on cosmetic wording rather than an agent's actual judgment, which undercuts any claim that one model is meaningfully safer than another based on a single benchmark number. It also echoes a long-running problem in language model evaluation, where surface-level prompt phrasing has been shown to move scores without reflecting real capability differences.

A security score that moves when you swap a synonym was never measuring security. It was measuring vocabulary.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →