Ask a chatbot to explain how to make a weapon and it refuses. Ask it to write a scene where a character explains it, and the same model often complies.
A new benchmark called GUISE tests that gap across languages and registers. On Qwen3-1.7B, researchers got the model to answer harmful requests wrapped in role-play or narrative framing 89.4% of the time in English, 93.0% in modern Chinese, and 95.7% in Classical Chinese, versus far lower rates for the same requests stated plainly. The benchmark also counts 'warn-then-answer' responses, where a model objects and then helps anyway, as failures rather than partial credit. The researchers trace the gap to internal representations: switching language barely moves a harmful request away from the model's refusal direction, but wrapping it in a story moves it much farther.
That's the more interesting finding than the jailbreak itself. Safety training is mostly built and tested in English, and this suggests the fix isn't translating guardrails into more languages - it's addressing how narrative framing itself displaces the signal that triggers a refusal, regardless of language.
The paper's proposed fix, AXIS, nudges harmful-request representations back toward the refusal direction during training and penalizes half-hearted refusals. Tested on Qwen3-1.7B, Qwen3-4B, and GLM-4-9B, it scored best on a combined safety-and-usability metric among the methods compared. Whether that holds up against the next wrapper nobody tested for remains the open question with every jailbreak defense.