AI/ ai · reward-models · research · human-feedback

Study Shows AI Reward Models Should Stop Averaging Taste

A new arXiv paper trained AI models on 575,000 wheel design judgments and found that individual taste beats a single averaged preference model.

A new paper argues that AI models trained on human feedback have been averaging away real disagreement instead of learning from it.

Researchers built a system that estimates "individuated utility" functions - preference models tied to a specific person and their decision context - instead of assuming everyone shares one underlying taste. They tested it on more than 575,000 pairwise judgments from 2,398 people comparing automotive wheel designs, a deliberately subjective task with no objectively right answer. The individuated models beat standard universal preference models, including foundation-model baselines, at predicting what a given person would actually pick. The gap held even though both approaches were trained on the exact same raw comparisons.

Most reward models behind chatbots and recommendation systems still pool annotator votes into one consensus score, treating disagreement as noise to average out. This paper's results suggest that noise is often signal: two people can look at the same wheel, or the same chatbot response, and land on different favorites because their utility functions genuinely differ, not because one of them is wrong. That is a direct challenge to how reinforcement learning from human feedback (RLHF) usually handles subjective calls like tone or style.

It is a study about car rims, but the argument generalizes to any AI system asked to judge taste - and that is most of them.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →