Security/ ai-security · red-teaming · reinforcement-learning · llm-agents

Cybersecurity AI Agents Trade Wins Depending on Network Size

Researchers pitted RL and LLM-based red-team agents across network sizes and found each wins for a different reason, failing at different stages of attack.

Which AI agent makes the best digital attacker depends on how big the network is, according to a new benchmarking study.

Researchers compared two hierarchical red-team architectures - one using reinforcement learning for both planning and execution (RL+RL), the other using large language models for both (LLM+LLM) - against expert autonomous defenders. They ran 18 configurations across three environments: the CybORG CAGE-4 benchmark and two sizes of the Cyberwheel simulator, a 100-host network and a much larger 1,010-host one. RL+RL dominated the smaller setups, disrupting 78.5% of attacks in CAGE-4 versus 18% for the best LLM+LLM setup, and 81% versus 50.5% on the 100-host network. On the 1,010-host network, though, a pretrained cybersecurity LLM agent won decisively, 55% to 0%, as the RL agent's success rate collapsed entirely.

The interesting part is why. RL agents are good at the mechanical work of finding and breaching hosts but stall at privilege escalation once a network gets large and gated. LLM agents have the opposite problem: they talk their way into privileged access easily but rarely convert it into real operational impact. So the "right" architecture for an automated attacker, or a defender trying to anticipate one, depends on which bottleneck you're worried about, not on some general claim that one kind of AI is smarter.

In other words, there's no universal best AI hacker yet, just a lot of specialists that are great at one part of the job and useless at the rest.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →