AI chatbots are easier to jailbreak in some languages than others, and a new paper traces why.
Researchers studied how large language models decide to refuse harmful requests, moving past single-neuron analysis to trace multi-layer pathways that carry safety signals through the network. They found each language has its own safety pathway, but also identified a small, shared subset of pathways connecting high-resource languages, like English, to languages with less training data. That overlap functions as an internal bridge, carrying safety behavior learned in high-resource languages over to lower-resource ones. The team then built an alignment method that updates only the parameters inside those shared pathways, instead of fine-tuning the whole model.
Most safety patches either retrain a model broadly or bolt on per-language fixes, both expensive and prone to degrading performance elsewhere. Targeting the shared pathways instead improved refusal rates for harmful requests in lower-resource languages while largely preserving the model's general capabilities, according to the paper.
It's a mechanistic explanation, not a fix shipped in any product. The usual approach to this problem has been piling on more non-English refusal examples during training - this suggests that was always treating a symptom rather than the cause.