Security/ ctf · llms · cybersecurity · ai-safety

Study Maps Where AI Now Beats Humans in CTF Hacking Contests

A new study finds AI already clears easy and mid CTF challenges and proposes tiered divisions and AI-resistant design as fixes organizers could adopt.

AI can now clear most of the easy stuff in hacking competitions, and a new academic study says the rules need to catch up.

The paper, posted to arXiv in July, draws on published benchmarks (including a government evaluation), case studies from live Capture the Flag contests across cryptography, web exploitation, and binary exploitation, observation of community forums, and interviews with players and organizers. It finds that easy and intermediate challenges in all three categories are now reliably solved by large language models, though narrower sub-categories still hold out. The authors argue that the community's fight over whether AI should be allowed sidesteps a more basic question: what a CTF competition is actually supposed to measure. Their proposed fix is a four-part framework: divisions tiered by AI use, challenge design meant to resist LLMs, telemetry for flagging suspicious runs, and a draft code of conduct, tied together by a decision tool that matches safeguards to a competition's stated purpose.

CTFs generate a lot of the training and hiring signal that cybersecurity relies on, so if a leaderboard stops reflecting human skill, that signal breaks. The paper's real contribution isn't the benchmark numbers, which mostly confirm what practitioners already suspected. It's the framing: the same validity problem applies to any credential built on a demonstrated result as evidence of ability, from job interviews to security certifications.

None of this is enforced anywhere yet. It's a proposed framework in one paper, not a rulebook any competition has adopted, so expect plenty more arguing in community Discords before anyone standardizes divisions or bans a model.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →