Researchers say they found a way to turn AI safety instructions into a dial you can turn up or down.
In a paper posted to arXiv on September 30, 2026 titled "The Safety Operator: Modulating the Expression of Safety Instructions via Spectral Optimization" (arXiv:2609.36434), researchers show that the context tokens carrying a safety instruction can be folded into a language model's weights as a single mathematical operator. They found that adjusting the operator's dominant eigenvalue, a single number describing how strongly it stretches the model's internal representations, directly controls how forcefully the safety instruction shapes what the model generates. To exploit this, they built a training method called Contrastive Safety Loss, which uses a tunable suppression weight to push the operator's influence up on harmful queries and down on harmless ones. Turning that suppression weight traces a curve between two failure modes: letting harmful requests through, and refusing harmless ones.
Most safety tuning today is blunt: retrain the model, add more refusal examples, hope the balance improves. This approach treats safety as a single adjustable parameter instead of a retraining project, and the researchers report that picking the right suppression weight can improve both attack resistance and over-refusal rates at once rather than trading one for the other.
That is a big claim from one paper with no third-party replication yet, and "improves the curve" still means somebody has to decide where on it sits.