Anthropic's newest model is also its most jailbreakable one, according to an independent red-team study.
Researchers used the HackAgent framework to throw four families of automated jailbreak attacks at Opus 4.8, Fable 5, and Fable 5.1, covering 7,826 harmful prompts across a ten-category harm taxonomy. Every apparent break was re-checked by a panel of five judge models before it counted. On the two attack types run against all three models - adaptive tree search and one-shot persuasion - Fable 5.1 failed most often, at 8.19 percent, behind Opus 4.8 at 5.76 percent and Fable 5 at 2.72 percent. Across the full study, the attacks produced 1,315 confirmed harmful completions from Opus 4.8, 620 from Fable 5, and 1,282 from Fable 5.1, all generated automatically without a human attacker, usually within one or two refinement steps.
The raw percentages are less interesting than what breaks each model. Adaptive tree-search attacks broke all three at roughly the same rate, but one-shot persuasion prompts split them by more than 10 times over - Fable 5 barely budged at 0.7 percent while Fable 5.1 folded at 9.5 percent. That means newer releases aren't uniformly hardening defenses; a model can improve elsewhere while getting more susceptible to a specific style of manipulation.
It is a reminder that a newer model number does not mean a safer one, and that jailbreak resistance deserves its own release notes.