Security/ multi-agent systems · ai security · llm agents

Researchers Find a Cheap Fix for AI Agent Message Tampering

MIRROR, a new integrity check for AI agent networks, cuts a near-100% attack success rate to zero without the runaway costs of semantic filtering.

A new communication protocol makes it dramatically harder for an attacker to quietly rewrite messages passing between AI agents, without the costly overhead of today's defenses.

The system, called MIRROR, targets Agent-in-the-Middle attacks, where a corrupted relay alters messages between AI agents without ever compromising the agents themselves. Prior research found such attacks succeed near 100% of the time on structured tasks, and existing defenses either rely on semantic checks that need extra inference or on encryption that does nothing once a middleman legitimately terminates the TLS connection. MIRROR instead copies a single payload across several logical routes and only accepts a message when a strict majority of those routes report the same hash. Tested on MMLU, HumanEval, and MBPP across two frameworks, four communication topologies, and a live MetaGPT deployment, it drove the attack success rate to zero whenever honest routes stayed in the majority, at roughly the same token cost as sending the message once.

That cost comparison is the real finding here. The leading alternative, having another LLM judge every message for tampering, costs 35 times more in token spend and still blocked up to 44.2% of legitimate messages in testing. As companies wire more autonomous agents together to do real work, the channel between them is becoming as much of an attack surface as the agents themselves, and most teams are not defending it at all.

Majority-vote routing is only as strong as its weakest assumption: that most of the routes stay honest. Enterprises adopting this should stress-test that assumption before they trust it.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →