Security/ ai safety · jailbreaks · llm security · adversarial attacks

New Study Maps When AI Jailbreak Attacks Turn Exponential

New research shows jailbreak success rates shift from polynomial to exponential growth as injected prompts get longer and sampling increases.

Researchers have pinned down exactly when AI jailbreaks stop being a nuisance and start becoming close to inevitable.

A new paper on arXiv examined how attack success rate changes as attackers throw more inference-time attempts at a safety-aligned language model, testing models from 3 billion to 70 billion parameters with established attack methods like GCG and AutoDAN against the AdvBench and HarmBench benchmarks. Without an injected prompt, success crept up slowly and polynomially no matter how many attempts were made. A short injected prompt shifted that curve to a power law, while a long injected prompt made it exponential. To explain the jump, the authors built a theoretical model borrowing from spin-glass physics, treating safe and unsafe outputs as clusters in an energy landscape and the injected prompt as a magnetic field pulling generations toward the unsafe clusters - the longer the prompt, the stronger the pull and the sharper the shift to exponential growth.

That distinction matters because many defenses lean on the idea that breaking a model takes too many attempts to be practical, which assumes success rates always grow slowly. This paper shows that assumption only holds when an attacker's injected prompt stays short; once it's long enough, the exponential regime takes over and the number of attempts needed to find a working jailbreak stops creeping up and starts collapsing.

The paper doesn't say how many attempts that collapse takes in practice - it's a theoretical scaling law, not a field report - but it's still a useful reminder that safety margins measured in 'how many attempts would it take' are only as sturdy as the assumptions behind them, and those assumptions apparently break down once an attacker is willing to write a longer prompt.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →