A new paper shows how to make AI request routing cheaper without sacrificing much accuracy.
The problem: zero-shot classifiers that route a user's request to the right specialized AI task normally have to check that request against every category in a taxonomy, one by one. With 60 categories, that's 60 checks per request, and the cost climbs with every category you add. The researchers' fix is a two-stage pipeline: a small ModernBERT model scores all 60 categories in a single pass and narrows them to a short list, then a larger DeBERTa-v3 model reranks only that short list instead of the full set. The two models train each other in rounds, with the bigger model's judgments steadily sharpening the smaller one's shortlists. The best resulting small model agreed with the full-size reranker 77.5% of the time on a 200-example test, and its top-16 shortlist captured 91 to 100 percent of what the full process would have decided anyway.
The sharper angle is who the paper is arguing with. It explicitly positions itself against commercial "System-1" routing classifiers, naming TypeSafe AI's Jev and the open-source Laya project, and points out that both have training methods that are either undocumented or reinforcement-learning-based. Routing is becoming invisible plumbing for a lot of AI products, and most of that plumbing is opaque. The paper's other finding is a useful technical warning for anyone building similar systems: truncating a teacher model's scores to a shortlist and zeroing out the rest, a common shortcut, quietly breaks a standard distillation technique rather than just approximating it.
Worth noting: the headline numbers come from a 200-example evaluation set, and the authors say a fuller evaluation is still in progress. Promising lab result, not a finished product.