AI/ ai agents · static analysis · cve benchmarks · security tooling

AI coding agents still struggle to build security checkers

A new benchmark shows AI coding agents solve fewer than half of the tasks required to build working security checkers.

AI coding agents still can't reliably build the static-analysis checkers meant to catch known security bugs, according to a new benchmark.

Researchers built CheckerBench, an executable benchmark of 300 tasks drawn from 297 real-world CVEs across 167 repositories, 85 CWE categories, and five programming-language ecosystems. Each task hands an agent a vulnerable code revision, a fixed revision, a pinned analysis environment, and a starter checker scaffold, then asks it to write analyzer logic that tells the two apart. A companion framework called CheckerLab independently rebuilds every submitted checker and scores it on whether it flags the vulnerable version without flagging the fix, whether it points at the right patch location, how many false positives it generates, and how well it uses its tools. Across 21 model-and-harness combinations, each run three times, the average pass rate was 32.30%, and even the best-performing setup only reached 45.33%.

Static-analysis checkers are the unglamorous backbone of vulnerability scanning: get one wrong and you either miss real bugs or bury developers in false alarms. A sub-50% ceiling on writing them end to end, across dozens of model setups, suggests the hard part of security tooling is still well beyond autopilot.

For an industry currently pitching AI agents as drop-in replacements for application-security engineers, this is a useful gap between the pitch and the plumbing.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →