Researchers say they can make AI agents less likely to gamble on catastrophic power grabs - by giving them a personality.
In a paper posted to arXiv (2609.38093) on September 30, 2026, a team describes "character training" for risk aversion in AI agents. They wrote a model constitution encoding constant absolute risk aversion (CARA) over an agent's resources, then instilled it through on-policy distillation rather than training directly on any specific benchmark. The resulting models had never seen the benchmark's decision format during training, yet still matched baselines trained directly on it, and beat those baselines on out-of-distribution generalization for two of the four models tested. The researchers also tested which parts of the constitution mattered most, and found token budget and choice of underlying model made the biggest difference in how risk averse an agent turned out.
The pitch here is narrow but consequential: a misaligned agent that is also risk averse should prefer negotiating with humans over rolling the dice on rebellion, since rebellion is the riskier play. That reframes AI safety less as "prevent misalignment entirely" and more as "shape the disposition of agents that might already be misaligned" - a hedge, not a cure.
It is an early result on four models and one hand-built constitution, not a deployed safeguard, and a technique that works on a benchmark is a long way from holding up against a genuinely deceptive agent trying to game its own personality profile.