A new paper claims you can pretrain a single LLM router once and reuse it across any model lineup, no retraining required.
The system, called RouteFM, learns to size up anonymous candidate models from their behavior rather than binding decisions to specific model identities. It is trained across many different routing setups so a frozen version of the router can adapt to a brand new pool of models just from a handful of in-context examples. On MMR-Bench, a benchmark held out of that pretraining, RouteFM beat the strongest baseline by 2.23 quality points using only eight observations per candidate model. The paper does not define what quality points measure beyond aggregate response quality on MMR-Bench, and it never reports the baseline's own absolute score, only the margin RouteFM won by.
Most routers deployed today are fitted to one specific workload and one specific set of models, then need retraining whenever either changes, which is a real tax on teams juggling multiple LLM vendors. If a routing capability really does generalize across domains, modalities, candidate pools, and context budgets the way this paper claims, that reframes routing as a one-time investment rather than ongoing maintenance, the same shift foundation models forced on other narrow, task-specific systems.
That claim rests on a single held-out benchmark and an unexplained scoring scale. The paper, arXiv:2609.37362, lists no author names or institution, only a GitHub organization, LAMDA-Model-Reuse, hosting the code. Worth a second look once someone outside that project replicates the comparison.