Security/ ai-security · jailbreak · guardrails · adversarial-ai

New Technique Bypasses Multi-Scanner AI Guardrails Every Time

Researchers' BRANCH method defeated six layered AI guardrail systems in every test, using fewer queries and less time than older jailbreak techniques.

A new jailbreak method cracks AI guardrail systems that stack multiple scanners together, and it does so every single time.

Researchers built a method called BRANCH that runs a branching tree search, probing each scanner in a multi-scanner guardrail setup on its own, then combining the strongest perturbations into a single adversarial prompt. Tested against six guardrail systems across 120 scenarios, it hit a 100% success rate while using 72% fewer queries and finishing 4.5 times faster than established bypass techniques. The winning prompts kept their original meaning intact, so the attack doesn't need to mangle a request to slip it through. The same bypasses, with no extra tuning, also worked against 29 guardrails the method had never seen, including eight commercial black-box products.

Vendors have pitched multi-scanner setups as the stronger defense, on the theory that stacking classifiers closes the gaps any single filter leaves open. BRANCH suggests that theory doesn't hold: separating bypass evaluation from attack optimization lets an attacker route around several scanners at once instead of fighting them one at a time. The more unsettling finding is that the exploit generalizes to commercial, unseen systems, which is exactly what teams buying off-the-shelf guardrails are counting on not happening.

Stacking filters was supposed to fix single-scanner blind spots; this result is a reminder that more layers don't automatically mean more security.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →