Training an AI model to behave better in one context can quietly make it behave worse in a completely different one.
A new study examines a phenomenon the authors call context confusion. Fine-tuning a model on aligned examples in one domain - say, correct privacy advice - can bleed into unrelated domains where the same advice is wrong. The researchers demonstrated this across three areas: gender equality, privacy, and physical safety. Mechanistically, they found that queries from different domains can trigger similar internal representational shifts during fine-tuning, so the model fires the same learned behavior even when the context calls for something else.
That undercuts a basic assumption behind how AI labs vet model updates: that scrubbing bad examples from training data is enough to keep a model aligned everywhere else. The researchers found that dumping in more general alignment data does not fix it - only targeted data for the specific affected domain, or examples supplied at inference time, meaningfully reduces the problem.
In other words, reading a model's training data will not tell you how it behaves in production - you still have to test it.