AI/ ai · llm-alignment · reinforcement-learning · research

Researchers Target Bias in Reward Models Trained on Clicks and Skips

ImplicitRM learns AI reward models from clicks and skips instead of costly human labels, correcting for the bias in which responses get noticed.

Researchers have built a reward model that learns from clicks and skips - without inheriting their built-in bias.

The method, called ImplicitRM, trains reward models on implicit user feedback like clicks, copies and skips, rather than the explicit human ratings RLHF normally depends on. That's cheaper to collect at scale, but implicit feedback has no clean negative examples - a skip doesn't confirm a bad answer - and it's skewed by selection bias, since some responses simply get more attention than others regardless of quality. ImplicitRM addresses both problems by sorting training samples into four latent groups using a stratification model, then applying a likelihood-maximization objective the authors say is theoretically unbiased. In tests across diverse LLM backbones and benchmark datasets, the approach produced more accurate reward models and improved results on downstream RLHF tasks.

Human-labeled feedback is the most expensive, slowest part of RLHF, and it doesn't scale with how much text models now produce. A reward model that can learn safely from ordinary usage signals - without just learning to chase clicks - would cut that cost without handing labs a new popularity-bias problem to clean up later.

It's the recommendation-engine playbook applied to model alignment, and the hard part was never using clicks as a signal - it was proving you'd accounted for why some answers get clicked and others don't.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →