A new technique lets a panel of cheap, open-weight AI models catch hallucinations almost as well as expensive frontier ones.
Researchers built RAIM (Robust Aggregation of Inexpensive Models), which combines ten small open-weight language models, 4 to 9 billion parameters each, drawn from different model families, to judge whether AI-generated text stays faithful to its source - the job companies currently pay frontier models like Claude Sonnet to do. The method stacks the panel's outputs with a statistical model and includes a built-in check that flags when the aggregate is actually trustworthy enough to use. Tested across eight faithfulness benchmarks, the panel kept a median 93% of Claude Sonnet's agreement score and gave up only 2.9 points of accuracy on average, while running at roughly a sixty-fourth of the price. The only upfront cost is calibrating the panel on 50 to 100 labeled examples per task.
Hallucination checking is turning into infrastructure - something every AI product needs running constantly in the background, not a one-time test. At frontier-model prices, that kind of nonstop monitoring gets expensive fast, so a panel that holds onto most of the accuracy for a fraction of the cost changes the math on who can afford to run it. It is not a clean win everywhere: the panel beat the frontier judge on one benchmark and clearly lost on three, so the gains depend on the task.
In other words, cheap judges will not dethrone the best one for every case - but for the repetitive, high-volume work of flagging made-up facts, boring and inexpensive is often good enough.