AI models can already be nudged toward or away from a personality trait mid-conversation - a new method claims to make that nudge match what you actually asked for.
A paper posted to arXiv on September 30, 2026 (arXiv:2609.36388, 'Persona Dosing: Calibrated Activation Steering for Graded Trait Control') introduces PersonaDose, a system for controlling how strongly a language model expresses a trait - say, sycophancy or bluntness - using a plain-language description and a target intensity. It builds on flow-based activation steering (FLAS), training a shared controller on persona responses without ever pairing that training data with a specific requested strength, then calibrating the controller's output against measured trait expression afterward. Tested on Llama-3.1-8B, Qwen3-8B, and Gemma-3-4B, PersonaDose pushed trait expression 33.2, 18.3, and 17.8 points higher than contrastive activation addition, a common baseline steering technique, while holding coherence above a floor score of 75. On seven trained traits, requests calibrated to a specific target landed within 4.7 to 6.2 points of that target - but only for 14 to 22 of 28 possible target levels per model.
That accuracy gap is the real story. Activation steering already lets developers tune chatbot behavior without retraining a model, but 'strength 3 out of 5' has mostly meant guesswork - crank the coefficient and check the output. PersonaDose's calibration step tries to turn that guesswork into a number a product team can request and actually get.
It's one arXiv preprint testing three open models on lab-defined traits, and by the authors' own numbers, calibration missed its target on roughly a third of the settings tried - dial control you still have to double-check.