A new training method makes reasoning AI models actually listen to their own safety reasoning, even when a jailbreak attempt isn't written in English.
Researchers built a framework called ACTR, short for aligning cross-lingual thoughts and responses, to fix a specific flaw in reasoning LLMs: a model can correctly flag a request as unsafe in its internal reasoning trace, then still generate an unsafe answer, especially when the prompt comes in a non-high-resource language. The team introduced a think gap score to measure how much a model's reasoning actually shapes its final response across languages, then ran neuron-masking experiments to isolate the specific neurons responsible for acting on that safety reasoning, which they call safety think neurons. They fine-tuned only those neurons using a technique called neuron-selective consistency optimization, which uses a separate judge model to reward agreement between the safety reasoning and the safety of the response, without requiring human-labeled training data. Tested on two reasoning models against the jailbreak benchmarks AdvBench-X and MultiJail, ACTR produced lower average attack success rates than the state-of-the-art methods it was compared against, with the safety gains holding up even in languages the method wasn't trained on.
This targets a specific, underappreciated failure mode: safety training that works in English can quietly fall apart in other languages, and most defenses focus on making models refuse more often rather than making them actually act on reasoning they already produced. Because ACTR touches only a narrow set of neurons and relies on an automated judge instead of human raters, it's a cheaper patch than full retraining, which matters as reasoning models spread into markets where English isn't the default input.
The paper does not publish the actual attack-success-rate percentages for ACTR versus the baselines on AdvBench-X or MultiJail, so there's no way to tell if lower means a rounding error or a double-digit swing, which matters given how loosely state of the art gets used in AI safety papers.