New research shows AI judges can rank chatbot answers correctly and still get the real-world verdict wildly wrong.
Researchers built an audit called O*NET-BENCH, drawing on an existing survey of 45,796 worker ratings on whether AI responses meet job requirements. From that pool they tested 33 pre-existing judge configurations spanning six model families, but only against a holdout of 4,501 test ratings. Twenty-five of those configurations hit a tie-aware pair accuracy of at least 0.60 - decent agreement on which response is better - though a bare-bones TF-IDF keyword baseline nearly matched the best judge. The real trouble showed up when judges tried to estimate how many responses were actually acceptable: their guesses ranged from 3.0% to 97.9%, while occupation-matched workers put the true figure at 61.1%.
Ranking two answers in the right order is not the same skill as estimating an acceptance rate, and this audit is notable for catching judges acing the first while failing the second. That gap matters for anyone using an LLM judge to decide if a model is ready to assist or replace workers in a given occupation - the ranking can look fine right up until the rate you report is off by dozens of points. Calibration scrubbed out most of the average bias, but tuned scores still explained at most 8.5% of the variance in individual worker ratings, and giving judges a little real worker data to lean on brought only small gains.
If your evaluation pipeline reports high judge agreement and calls it validated, this is the fine print that says otherwise.