Researchers have built a way to account for bias in the AI judges that now help train chatbots, though they admit the bias itself doesn't go away.
Training large language models on human preferences normally means paying people to compare pairs of responses, which is slow and expensive. To cut costs, researchers have started supplementing that process with AI judges, whose feedback is fed into active learning systems that pick the most informative comparisons to study. The problem is that AI judges don't always agree with the human population a model is meant to serve, and that mismatch can get worse once the system starts actively selecting which comparisons to use. The paper proposes Nuisance-Adjusted Optimal Design, a selection method built on the Frank-Wolfe optimization algorithm, and tests it on Chatbot Arena data across 17 judges and 15 budget configurations.
The method doesn't erase judge bias. It adjusts for it mathematically so the bias doesn't masquerade as useful signal, and the paper shows that ignoring this distinction can reverse the usual advantage of smarter comparison selection. In testing, the adjusted approach cut a benchmark regret measure by 29.1% and improved predictions of real human preferences on new data.
It's a reminder that using AI to grade AI doesn't remove the grading problem, it just moves it into the math.