Security/ ai agents · cybersecurity · benchmarks · threat detection

Benchmark Finds AI Security Agents Miss Most Attack Evidence

A new benchmark testing eleven AI models on simulated cyberattack investigations finds they properly cite evidence for just a quarter of their findings.

A new benchmark says AI agents chasing down advanced cyberattacks are still shaky once security teams change which logs get fed to them.

Researchers built APTInvestBench, a set of 370 investigation cases spanning seven different telemetry setups meant to mirror real security operations centers. The cases are drawn from 56 reconstructed attacks, backed by 16.4 million log records, and task AI agents with turning a vague lead into a documented, cited incident report. Across eleven large language models, the agents gathered enough evidence to confirm only 44.3% of the attack steps they could theoretically have found, and just 25.0% of findings came with the formal, record-level citations needed to back them up. Trimming telemetry down to endpoint logs barely moved raw coverage, a drop of only 1.6 percentage points, but it quietly wiped out the citation support for 35.5% of findings that remained just as recoverable.

That is the gap between an AI agent saying it found something and being able to prove it, and in an incident report, proof is the job. The instability held across four different agent frameworks, which suggests this is not a one-off prompting quirk but a structural weakness in how these agents track their own evidence.

Security teams piloting AI copilots for their SOC should ask vendors what happens to accuracy when the log feed gets thinner, not just what it looks like on the full dataset.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →