A new evaluation method called OTROPE lets cheap AI judges size up language models almost as reliably as expensive ones, and sometimes better, once you pool several of them.
Researchers built OTROPE as an off-policy evaluation technique: it grades a new target LLM using old, human-labeled data collected on a different behavior model, rather than running fresh live tests that are costly and risky. It works even without access to a model's internal likelihoods, using optimal transport to align labeled examples from the old model with unlabeled examples from the new one in a shared semantic space, then combining corrected residuals with a backup predictor. The researchers say this fixes a specific failure mode: standard evaluators break down when the new model's behavior drifts too far from the one the labels were collected on. In tests on synthetic and real LLM evaluation tasks, OTROPE beat baseline methods, and ensembles of weaker LLM evaluators built on top of it were able to approach, and sometimes surpass, the accuracy of a single stronger evaluator.
Evaluating LLMs safely and cheaply is a bottleneck for anyone shipping models faster than they can afford to test them live. If a group of cheap judges can match or beat one expensive judge, that changes the economics of quality control, letting smaller teams grade models without paying for premium evaluators or risking a live rollout just to get labeled data.
This is one arXiv paper, not yet peer reviewed or adopted anywhere, so read sometimes surpass as an early, narrow result worth tracking rather than a settled fact.