A new routing framework decides which AI models to call by picking teammates, not just the top scorer.
Researchers describe FlexRouter, a system that selects a set of large language models for a given task based on how well they complement each other, not just how good each one looks in isolation. Most current routers rank models independently and pick the top performers, which often means stacking similar models that fail on the same inputs. FlexRouter instead uses determinantal point processes, a statistical tool for picking diverse subsets, to estimate the odds that at least one chosen model gets the answer right. It trains on failure patterns rather than requiring a labeled ideal set of models, and at inference time it greedily adds models only as long as each one meaningfully improves the group's coverage. On the RouterEval benchmark, the approach reportedly beat existing baselines on both familiar and unfamiliar tasks while using fewer redundant model calls.
The idea matters because production AI pipelines already run multiple models and let a verifier or a human pick the best output, so redundancy is a real cost, not a hypothetical one. A router that can tell two similar models apart from two models that fail differently could cut inference spend without sacrificing accuracy, which is the actual bottleneck in agentic and multi-model systems.
Whether this beats simply running every model and letting a verifier sort it out is the open question the benchmark results do not fully settle.