A new research system called ORACLE routes AI agent tasks to the right model without the usual accuracy hit.
ORACLE sits on top of whatever model-selection policy a company already uses. It classifies each incoming task - a coding job versus a general chat query, say - and assigns a verifier built for that specific task type. Earlier routers leaned on one fixed verifier for everything, which worked fine for uniform workloads but broke down once teams mixed coding agents with conversational ones. ORACLE also delays verifier feedback for simultaneous requests, pulling that check out of the critical path so a busy system does not slow to a crawl. A companion scheduler called DISC reserves each task's peak memory footprint up front and reroutes work to a backup model only when the time saved outweighs the accuracy lost.
Enterprises running mixed fleets of cheap and expensive models have mostly tuned routing with static, one-size-fits-all checks. On SWE-bench, tau2-bench, and Terminal-Bench 2.0, the researchers found that mismatch costs real accuracy - and fixing it with task-aware verification lifted the accuracy-cost frontier by up to 7 percentage points, with DISC adding up to 1.8x more throughput.
It is a research paper, not a shipped product, so the real test is whether any orchestration vendor bothers to bolt this onto a live fleet.