A new study says most commercial LLM routers are worse than flipping a coin.
Researchers tested six commercial routers across 14 configurations on a benchmark spanning eight task categories. None beat a baseline that randomly picks between two well-chosen models at the same cost. Some routers underperformed that random baseline by more than 10 percentage points. The researchers trace the gap to four recurring habits: ignoring how hard a query actually is, favoring short queries over long ones, matching queries to models by surface wording instead of difficulty, and leaning on rosters of models that do not add real value.
Routers are sold as a way to cut inference costs by sending easy questions to cheap models and hard ones to expensive models. This research suggests the economics behind that pitch are shakier than vendors claim, and that the standard way routers get scored rewards exactly the shortcuts that make them fail in practice.
The researchers built their own two-model router that avoids these traps, and it still only barely edges out random selection, a reminder that a well-chosen pair of models leaves little room for cleverness.