AI models that can chat convincingly about hacking are not the same as ones that can actually type the right command.
Researchers built KaliBench, a benchmark that tests whether language models can translate plain-English security requests into exact command-line syntax for tools on Kali Linux, the distribution security analysts use for penetration testing. The dataset contains 8,504 query-command pairs covering 1,642 tools, split across 23 capability dimensions and five phases of a security engagement. It was built from tool manuals and checked through a multi-stage pipeline combining automated validation, sandboxed execution, and human review, so answers get graded for exact correctness rather than vibes. Across 24 configurations of general-purpose and security-focused open-weight models, none topped 42 percent exact-command accuracy when working without hints about which tool to use.
That number matters because security tooling has zero tolerance for approximation. A wrong flag or misplaced argument does not just produce a slightly-off answer - it can fail silently, miss the vulnerability, or break the exploit entirely. The researchers also showed that training smaller models on KaliBench's verifiable reward signal closed much of that gap, pushing an 8B model to match the accuracy of a 685B mixture-of-experts model.
That last result is the more interesting one. It suggests the bottleneck is not raw model scale but the lack of precise, checkable training signal for command syntax, the kind of thing coding benchmarks solved years ago with unit tests. Until more of the security-tooling ecosystem gets that treatment, pitches for autonomous pentesting agents deserve the same skepticism as any command typed by a junior analyst at 2 a.m.