AI/ ai-evaluation · ai-agents · benchmarking · thompson-sampling

Researchers Find a Cheaper Way to Catch AI Agent Failures

A new sampling method finds 86 percent of an AI agent's worst failures using just 50 test trials, versus 25 percent for standard uniform testing.

A new testing method finds the AI agent failures that matter most without wasting trials on scenarios that never break.

Researchers built a statistical policy, based on a technique called Thompson Sampling, that decides which test scenarios to rerun by weighing a scenario's risk profile, a fixed score for how much failure would cost, and failures already observed. They evaluated it offline against 70 airline-booking scenarios from the tau-bench benchmark and 824 recorded trials. At the smallest budget tested, just 50 trials, about 6% of the full trial corpus, the policy caught 86% of the most damaging failures an exhaustive search would find, versus 25% for simply spreading trials evenly across scenarios. It found 3.5 times more high-impact failures for the same spend and cut trials wasted on scenarios that never fail from 34% to under 3%.

That gap matters because most agent benchmarks still test a read-only lookup as often as an irreversible payment, even though the payment failure is the one that actually costs money or trust. Prioritizing tests this way is most useful exactly when budgets are tightest, which is the situation most teams shipping agents with real-world tool access are actually in.

The catch: this is offline replay on one airline-booking benchmark, not a live test against agents behaving unpredictably in production, so the real-world payoff is still unproven.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →