Security/ ai agents · prompt injection · llm security · ai safety

New Guardrail Vets AI Agent Actions Before They Execute

A new system vets AI agent tool calls against user intent, cutting prompt-injection attack success up to 46 percent over prior defenses.

AI agents are starting to install skills the way browsers install extensions, and that is exactly the problem.

A new research system called ActionGuard checks every tool call an AI agent makes against the user's original request before letting it run. The idea addresses a known weak spot: third-party skills can bundle hidden instructions that quietly redirect a benign task into data exfiltration, file deletion, or arbitrary code execution. ActionGuard keeps the poisoned skill text away from its own decision-making component, called the Reviewer, which instead judges each action using a sanitized skill summary, recent tool-call history, and local script contents, then issues an allow or deny call that defaults to deny when uncertain. Researchers tested it against 139 subtle and 180 obvious injected attacks, using three open-source and two commercial models as Reviewers, and compared results against two existing defenses, Dynamic Guardian and SkillGuard.

The results matter because agent skills are becoming a kind of supply chain, with the same trust problems that have long plagued package managers like npm. ActionGuard cut successful attacks by 35.54 to 46.11 percent compared to those existing guardrails, and by 70.44 percent compared to running with no safeguard at all, while still letting legitimate tasks complete.

That is a real improvement, not a solved problem. A notable share of attacks still got through in testing, and the benchmark relies on simulated injections rather than the messier skills marketplaces users will eventually download from.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →