AI/ llm-routing · ai-benchmarks · agentic-ai · coding-agents

New benchmark judges AI routers inside live agent tasks

A new benchmark scores AI model routers on multi-step agent tasks, not just single prompts, using deterministic checks instead of LLM judges.

A new benchmark judges AI routers on whether they can swap in cheaper models mid-task without wrecking the outcome, not just on a single prompt.

TwinRouterBench, detailed in a paper posted to arXiv, targets routing systems used inside agentic tools like coding assistants and research agents, where one user request can trigger dozens of model calls. Its static track offers 970 router-visible snapshots pulled from mid-task points across five existing benchmarks: SWE-bench, BFCL, mtRAG, QMSum, and PinchBench. Each snapshot comes pre-labeled with the cheapest model tier verified to still finish the job, using a downgrade-and-cascade protocol, so scoring is plain arithmetic rather than another LLM grading the routing decision. A separate dynamic track runs routers live against a harness built on SWE-bench Verified's full 500 cases, and the paper reports results on a 100-case held-out slice kept separate from the static track's training data, tracking both task success and actual API spend.

That step-level focus is the real contribution. Existing router benchmarks mostly score one-shot prompts, which says nothing about whether a downgrade three calls into a coding agent's trajectory quietly breaks the whole task. As more products wire together long chains of model calls, cost-cutting by routing to cheaper models only works if someone can prove the cheaper model does not sink the outcome five steps later.

The benchmark does not grade any specific commercial router, so the harder test is whether today's routing products would survive contact with it.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →