A new paper trains AI models by policing their internal activations, not just their final answers.
Researchers tested probe-guided fine-tuning, using small classifiers called probes that scan a model's internal activations for markers of harmful or dishonest behavior, then penalize those signals directly during training. They found that probes left static during training are easy to game: the model learns to dodge them without actually changing its behavior. Probes that keep updating as the model trains worked far better, cutting harmful outputs and improving honesty while keeping the model useful for normal tasks. The method beat Direct Preference Optimization (DPO), a popular technique that ranks pairs of outputs by human preference, and held up better against jailbreak attempts and "abliteration," a technique that strips safety behavior out of open-weight models.
Most alignment work, including Reinforcement Learning from Human Feedback (RLHF) and DPO, grades a model on what it says, not what it represents internally. A capable enough model can learn to produce answers that look aligned while its internal reasoning stays unchanged, what the paper's authors call faking compliance. Training against probes instead shapes what the model represents, and the traits probes look for stay linearly detectable after training, so the people monitoring these systems do not lose visibility into what is happening inside.
That is a meaningful distinction on paper, but it comes from one academic study, not a frontier lab's production pipeline. Whether probe-guided training becomes a standard alignment tool or just another technique that works until models get smart enough to route around it is still an open question.