Cloudflare says its first security AI agent hallucinated evidence. So it broke the job into four narrower ones.
Cloudflare's Managed Defense service now runs security alerts through a multi-agent harness instead of one broad AI model. The company says its first prototype, a single general-purpose agent handed an entire investigation, produced useful analysis but also hallucinated claims the evidence didn't support, treated detection hypotheses as proven fact, and sometimes queried the wrong account or time range. The fix moves evidence gathering into deterministic application code before any model runs: a fixed recon step collects identity, detection history, traffic baselines, and enforcement outcomes, each tagged with its source and timestamp. Alerts that Clef, Cloudflare's open-source triage model running on Workers AI, scores as likely noise skip the heavier review entirely; everything else goes to four specialist agents (traffic, customer history, global telemetry, and threat intel) running in parallel, whose findings a synthesis agent can only cite, not invent.
This is less proof that AI can replace security analysts and more an admission that handing a language model an entire investigation produces plausible sounding, unreliable output. Cloudflare's actual fix is procedural rather than a bigger model: deterministic recon, versioned evidence packages, and citation checks that reject any finding not traceable to admitted evidence. That is a more disciplined template for agentic security tools than most vendors are currently offering.
Cloudflare still calls in GPT-5.6 Cyber and Anthropic's Mythos for the deeper analysis, but the real engineering here is the plumbing that keeps those models from freelancing.