A researcher just showed that AI agents trust each other more than they should, and that trust is exploitable.
Independent researcher Syed Anas Mohiuddin probed AI agent deployments at six organizations - Google, JP Morgan Chase, Weaviate, Rapid7, France's interministerial digital directorate, and the US federal government - all of which rely on the Model Context Protocol, or MCP, the standard that lets agents inside a network talk to each other. His proof-of-concept attacks show that a single compromised agent, say one handling translation, can pass hidden malicious instructions to a second agent, say one handling data analysis, and the second agent follows them simply because it's configured to trust the first by default. That's not a jailbreak of the underlying model - it's prompt injection aimed at the agent layer, where guardrails are often thin or nonexistent. Over the past five months, five of those six organizations have acknowledged the vulnerability; the source does not specify which one has not.
MCP is fast becoming the default plumbing for agent-to-agent communication across millions of organizations, so this isn't a one-off bug in one company's product - it's a structural flaw in how the protocol handles trust. Patching individual deployments won't fix it; the protocol itself needs a way for agents to verify each other instead of trusting by default.
Network security learned this lesson with lateral movement years ago: let one compromised machine vouch for the next, and the whole network is one bad agent away from a breach. Calling MCP the riskiest protocol you've never heard of sounds like headline bait, but the underlying trust gap is exactly as boring and real as that.