AI/ ai-security · multi-agent-systems · llm-agents · research

New Defense Cuts Multi-Agent AI Attack Success to 3 Percent

A new paper shows cheap, deterministic rules catch most attacks on AI agent systems, with LLM judges only needed for the ambiguous rest.

A new security framework for AI agent teams drops attack success rates from about 30 percent to roughly 3 percent - and most of that improvement comes from plain rules, not smarter AI.

Researchers sorted existing defenses for multi-agent LLM systems into five principles and combined them into DEFER1, a cascade of 28 deterministic checks that screens requests before anything ambiguous goes to a panel of four LLM judges. Tested across four domains, the combined system cut successful attacks roughly tenfold. The deterministic layer did most of the work itself, blocking 78 percent of stopped attacks on its own; in the security-operations domain, only a quarter of flagged proposals ever reached the judges. The paper also flags a specific failure: a risk-score approval gate meant to pre-screen requests approved most attack proposals while wrongly rejecting legitimate ones.

As companies wire LLM agents into tool use, shared memory, and task delegation, the attack surface grows with every integration added. This research is a useful corrective: the fix isn't automatically 'add a smarter judge model' - it's figuring out which threats are outright policy violations that deterministic rules can catch, and reserving judgment calls for cases that actually require them.

It's also a reminder that the weak point is rarely the fancy part. A single miscalibrated approval gate reportedly waved through most of the attacks it was built to stop.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →