A new training method makes language models harder to manipulate across a multi-turn conversation, without dulling the model underneath.
Researchers built TRACE (Trajectory Return Attribution and Contrastive Erasure), a token-level training objective for safety-aligned models. It scores each token in a safe response using the discounted return of a refusal-attributable advantage, comparing a frozen reference model against a refusal-ablated copy of itself. That lets early turns in a conversation get credit - or blame - for a refusal that only shows up several messages later. The team tested it on five open-weight models against seven multi-turn jailbreak attacks, 35 model-attack pairs in all, and TRACE produced the lowest attack success rate in every single one. Utility scores on MMLU and HellaSwag fell by at most 1.23 points.
That's the gap current safety training ignores. Standard preference tuning judges one prompt and one response at a time, so it has no way to notice a harmful request being assembled piece by piece across a conversation. TRACE's trick - tracing later refusal signals back to earlier tokens - is a structural fix for that blind spot, not just a bigger blocklist.
Lowest attack success rate across 35 pairings is a real result, but it's not the same as closing the hole. The paper tested seven known multi-turn attack styles; it says nothing about the eighth one someone hasn't published yet.