AI/ ai alignment · llm training · ai research · arxiv

Study Charts When One AI Model Can Serve Conflicting Values

A new study finds AI models can sometimes serve two conflicting human values at once, but only reliably when trained on human, not AI, preference data.

Researchers have found a way to predict, before training even starts, whether tuning an AI model to be better at one thing will make it worse at another.

The paper studies steerable alignment, the practice of building one model that can be dialed toward different human preferences instead of picking a single fixed personality. Using a technique called Multi-Objective Direct Preference Optimization, the researchers tested seven pairs of objectives drawn from the HelpSteer and UltraFeedback datasets. They found two measurements taken before training that reliably predict whether two goals will reinforce or conflict with each other, but only when the training data comes from human annotators. When the data is AI-generated instead, the predictions break down because reward scores get skewed by quirks like response length and repetition.

This matters because pluralistic alignment, the idea that no single model can satisfy every user's values, is quickly becoming an engineering problem, not just a philosophy debate. Right now, covering many trade-offs usually means training a separate model for each combination, which is expensive; the paper shows that merging existing models or picking the nearest trained one helps some, but still falls short of training a dedicated model.

In other words, the field has found a cheap way to guess which value combinations are worth building for and which ones will just fight each other no matter how you turn the dial.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →