A purpose-built cloud-security AI just beat Claude Code, running as a general coding agent, on the same investigation tasks, and it wasn't close.
Researchers tested Sola Security's "Security Brain," a system that resolves cloud relationships offline and evaluates security logic against that model at query time, against Claude Code working the same live, read-only AWS environment through its CLI. Both ran 28 cloud-security investigation tasks, the kind that ask which identities can read a data store or how many resources violate a control. Answers were scored on coverage, a blinded relative-recall metric measuring how much of the full correct claim pool each answer captured, averaged across three grading passes. The Security Brain scored 0.693 coverage against Claude Code's 0.387, a 79.2 percent relative gain, and won on 25 of the 28 tasks, all while costing 17.7 times less in reasoning tokens per task and 31.6 times less per unit of coverage.
The finding pushes back on the assumption that handing a capable coding agent read-only cloud credentials is enough to replace dedicated security tooling. Many cloud-security questions are population questions, ones that require a complete inventory rather than a spot check, and a general-purpose agent working under a turn budget has no way to guarantee it has seen everything before it answers.
The paper's sharpest example: Claude Code sampled 40 of roughly 5,000 S3 buckets, checked four bucket families, and reported that no bucket policies existed, in an account where 65 buckets actually carried a wildcard-principal read grant. That's not a wrong answer so much as a confident answer to a question nobody actually finished asking.