AI/ ai-safety · llm-alignment · jailbreaks · research

New Research Turns AI Safety Prompts Into a Tunable Dial

A new arXiv paper (2609.36434) proposes tweaking a single mathematical value inside language models to reduce both harmful outputs and needless refusals.

Researchers say they found a way to turn AI safety instructions into a dial you can turn up or down.

In a paper posted to arXiv on September 30, 2026 titled "The Safety Operator: Modulating the Expression of Safety Instructions via Spectral Optimization" (arXiv:2609.36434), researchers show that the context tokens carrying a safety instruction can be folded into a language model's weights as a single mathematical operator. They found that adjusting the operator's dominant eigenvalue, a single number describing how strongly it stretches the model's internal representations, directly controls how forcefully the safety instruction shapes what the model generates. To exploit this, they built a training method called Contrastive Safety Loss, which uses a tunable suppression weight to push the operator's influence up on harmful queries and down on harmless ones. Turning that suppression weight traces a curve between two failure modes: letting harmful requests through, and refusing harmless ones.

Most safety tuning today is blunt: retrain the model, add more refusal examples, hope the balance improves. This approach treats safety as a single adjustable parameter instead of a retraining project, and the researchers report that picking the right suppression weight can improve both attack resistance and over-refusal rates at once rather than trading one for the other.

That is a big claim from one paper with no third-party replication yet, and "improves the curve" still means somebody has to decide where on it sits.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →