AI/ ai safety · llms · multilingual nlp · interpretability

Study Finds Shared Neural Pathways Behind AI Safety Gaps

Researchers found internal pathways that transfer AI safety training across languages, and tuning just those closes gaps in lower-resource languages.

AI chatbots are easier to jailbreak in some languages than others, and a new paper traces why.

Researchers studied how large language models decide to refuse harmful requests, moving past single-neuron analysis to trace multi-layer pathways that carry safety signals through the network. They found each language has its own safety pathway, but also identified a small, shared subset of pathways connecting high-resource languages, like English, to languages with less training data. That overlap functions as an internal bridge, carrying safety behavior learned in high-resource languages over to lower-resource ones. The team then built an alignment method that updates only the parameters inside those shared pathways, instead of fine-tuning the whole model.

Most safety patches either retrain a model broadly or bolt on per-language fixes, both expensive and prone to degrading performance elsewhere. Targeting the shared pathways instead improved refusal rates for harmful requests in lower-resource languages while largely preserving the model's general capabilities, according to the paper.

It's a mechanistic explanation, not a fix shipped in any product. The usual approach to this problem has been piling on more non-English refusal examples during training - this suggests that was always treating a symptom rather than the cause.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →