AI/ ai agents · llm evaluation · ai research · benchmarking

Fast AI Decision Models Underperform and So Did the Study's Own Math

A paired evaluation of two fast AI decision models finds one dramatically outperforms the other, and uncovers errors in its own earlier accuracy claims.

A new benchmark pits two fast AI decision models against each other, then checks its own math.

Researchers evaluated two 'System-1' models built to make small, fast decisions inside AI agent systems, like picking which tool to use or flagging suspicious input, without calling a full language model each time. Across 11 decision tests built from 18 public datasets, the hosted model, Jev, beat the open-weight model, Laya, on 9 of them, by margins of 10.8 to 46 percentage points. Both models failed at routing tasks to the right underlying model, scoring no better than chance, and tied on judging whether retrieved documents were relevant. Laya also proved brittle: it flipped 30% of its answers when answer choices were simply reordered, and its accuracy collapsed to 31% when choosing among 50 similar tool options, compared with 98% for Jev on clearer cases.

The more useful finding might be the self-audit. Reviewing their own numbers, the researchers found three analysis errors and one design confound that had inflated earlier deployment claims. A missing pre-screening cost turned a reported 23.9% savings into an actual 4.3%. A gate-accuracy figure got reported as overall pipeline quality, conflating a 58% result with a 98% one. Error thresholds were tuned on the same data used to test them, missing held-out targets by up to 17%. Separately, a 'channel effect' made injection-detection false positives look worse than they were, an effect that vanished once tested with content native to the real deployment channel.

Benchmarks that survive their authors' own fact-checking are rare. This one earns points for finding its own mistakes before a reviewer, or a deployment, did.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →