Six AI tools built to catch software vulnerabilities barely hold up when tested the same way twice.
A paper titled "Retrieve, Reproduce, Reveal: Dissecting Retrieval-Augmented Software Vulnerability Detection" (arXiv:2609.37669), posted September 30, 2026, reran six open-source systems that use retrieval-augmented generation, or RAG, to help large language models spot security bugs. The researchers rebuilt each system with open-weight models instead of the proprietary ones used in the original papers, then tested all six against a shared dataset, metric suite, and model pool. They also broke each pipeline into three stages: input abstraction, knowledge retrieval, and detection, to see which part actually drives performance. Reproducibility varied widely, and published results did not carry over once the models and evaluation conditions were standardized.
The most telling result: even when researchers fed a system perfect, oracle-level retrieval, detection accuracy topped out at 0.51 pairwise accuracy, barely better than a coin flip. That means bolting a search engine onto an LLM does not fix the underlying problem of judging whether code is actually vulnerable. For security teams evaluating these tools, the lesson is that end-to-end vendor benchmarks reported on closed models are close to meaningless without knowing which stage of the pipeline is doing the work.
It's the same lesson RAG chatbots learned two years ago: retrieval can hand a model the right paragraph, but it still has to know what to do with it.