A new algorithm claims that learning from preferences alone is no less data-efficient than learning from direct reward scores.
The paper studies preference-based bandits: a system repeatedly picks two options and learns only which one a rater preferred, with preferences modeled using the Bradley-Terry framework common in tournament ranking. That setup covers recommender systems, ranking tournaments, and training AI on human feedback, where people reliably say which answer is better but struggle to score it numerically. Earlier approaches to this problem were mostly stuck with simple linear reward models and were slowed by a troublesome constant in the math linking preferences to rewards, one that can grow large and inflate the data needed to learn. The researchers introduce a new complexity measure, the locally sensitive eluder dimension, and an algorithm called GINOP (Generic INformative OPtimism) that builds confidence sets from log-loss and picks arm pairs to balance curiosity about uncertain options against confidence in good ones.
The practical hook is reinforcement learning from human feedback, the method behind tuning chatbots like ChatGPT and Claude. If preference data is genuinely as efficient as direct reward data, as this paper's regret bound suggests, that undercuts the assumption that RLHF is a statistically wasteful workaround for not having real reward signals.
The claims rest on theoretical bounds and benchmark comparisons against baselines, not a live run on an actual language model, so the real test is whether the math holds up once human raters start disagreeing with each other.