AI agents built to triage security alerts miss more than four in ten real attacks - until you make them argue with themselves.
Researchers tested five common designs for LLM-based alert triage agents, from single-pass tool use to iterative retrieval and self-review, against a benchmark called ALERT-BENCH that replays real enterprise telemetry through a live SIEM. Across 1,247 alerts tied to a multi-stage attack scenario, every single approach missed at least 40.4% of the alerts that actually mattered. The failure pattern was consistent: agents dismissed alerts whenever a search turned up no matching records, having the model review its own work produced no real correction, and alerts marked for dismissal got no more scrutiny than ones escalated to a human. The researchers then built AIDA, a multi-agent setup that forces a proposed verdict to survive an independent challenge in a separate reasoning context, logs every step in an append-only Investigation Ledger, and sends the dispute to a separate judge model that can demand another round of evidence before closing anything.
That jump - from a 40.4% false-negative rate down to 3.1%, while escalating 18.4% of alerts to human analysts - is the difference between an AI tool that quietly rubber-stamps breaches and one that's actually useful in a SOC. It's also a pointed data point in the broader debate over agentic AI: a model reasoning in one pass is not the same as a model forced to defend its reasoning against a skeptic.
Escalating nearly a fifth of alerts isn't free - someone still has to work that queue - but it beats the alternative, which is an AI that looks decisive right up until the breach it waved through shows up in an incident report.