A new paper puts a number on how much preference data it takes to train a reward model that won't get gamed.
Researchers studied how AI reward models - the proxy scorers used to align language models with human preferences - degrade under heavy optimization, a failure mode known as reward hacking. They found performance scales as the square root of the smaller of two quantities: the log of the number of training comparisons, or the policy's KL-divergence budget relative to a reference model. The team tested the law using a 70B parameter model to generate gold-standard feedback and smaller 0.6B to 4B models trained as proxies, and the fit explained 97 to 99 percent of the variance across model sizes, noise levels, and optimization methods including best-of-N sampling and policy tilting. The upshot: once optimization pressure exceeds what the training data can support, more pressure does not help and can actively hurt.
This matters because teams building aligned models constantly trade off between collecting more comparison data and optimizing harder against the reward model they already have. The law suggests data collection has rapidly diminishing returns - since performance tracks the log of comparisons, doubling your data budget barely moves the needle once you have a few thousand labeled pairs. The real constraint, this framework suggests, is the divergence budget, not the size of the dataset.
The researchers frame the whole exercise as a selection problem dressed up in alignment jargon - picking the best option from noisy, roughly Gaussian signals - a less glamorous explanation for reward hacking than most papers offer, but a more falsifiable one.