Researchers have a new way to predict how training an AI on one value warps its behavior everywhere else.
AI labs fine-tune models to follow a list of values and traits, known as an alignment target, then test the model against that same list. The problem: training on those narrow, listed behaviors also reshapes how the model acts in situations nobody tested for, and those side effects are hard to anticipate. A team of researchers analyzed this spillover across 66 values pulled from alignment targets used in current models, building what they call a generalization matrix. They then tried to predict that matrix two ways: by reading the model's internal activations when it applies a value, and by analyzing plain-text descriptions of the values. The activation-based method hit a 0.45 correlation with actual generalization patterns; the text-based method managed only 0.05.
That gap matters because alignment evaluations mostly check whether a model nailed the behaviors it was explicitly trained on, not what else shifted. A method that forecasts the untested fallout, before a model ships, is more useful than finding out after release that a safety tweak nudged unrelated behavior. The researchers also used these activation-based representations to measure how similar the values within one alignment target are to each other, and found that similarity score tracked with how robust the resulting model was.
Call it a map of AI's blind spots - useful, but a 0.45 correlation means plenty of territory is still unmapped.