A new jailbreak method cracks AI guardrail systems that stack multiple scanners together, and it does so every single time.
Researchers built a method called BRANCH that runs a branching tree search, probing each scanner in a multi-scanner guardrail setup on its own, then combining the strongest perturbations into a single adversarial prompt. Tested against six guardrail systems across 120 scenarios, it hit a 100% success rate while using 72% fewer queries and finishing 4.5 times faster than established bypass techniques. The winning prompts kept their original meaning intact, so the attack doesn't need to mangle a request to slip it through. The same bypasses, with no extra tuning, also worked against 29 guardrails the method had never seen, including eight commercial black-box products.
Vendors have pitched multi-scanner setups as the stronger defense, on the theory that stacking classifiers closes the gaps any single filter leaves open. BRANCH suggests that theory doesn't hold: separating bypass evaluation from attack optimization lets an attacker route around several scanners at once instead of fighting them one at a time. The more unsettling finding is that the exploit generalizes to commercial, unseen systems, which is exactly what teams buying off-the-shelf guardrails are counting on not happening.
Stacking filters was supposed to fix single-scanner blind spots; this result is a reminder that more layers don't automatically mean more security.