AI/ ai · multi-agent-systems · llm-routing · research

Researchers cut multi-model AI costs by over 3x

JOVE splits AI reasoning tasks across multiple language models and selectively double-checks answers, cutting cost and latency by at least 3.17x in testing.

A new framework called JOVE splits complex AI reasoning queries into smaller tasks, routes each one to whichever language model looks cheapest and most capable, and only pays for a second opinion on the handful of answers worth checking.

Researchers describe JOVE as a system for breaking a query into a directed task graph, assigning each subtask to one of several LLMs, and asynchronously verifying select outputs instead of checking everything. The system solves a fresh optimization problem for every query, weighing execution cost against the value of learning which models are actually good at which jobs. That learning updates over time through online feedback, with a bonus built in for exploring uncertain model-task pairings rather than always picking the safe bet. Tested on four reasoning benchmarks, JOVE matched the accuracy of standard inference setups while cutting average cost and latency by at least 3.17 times.

The interesting part is not the speedup. It is the admission that running a model tells you nothing about whether its answer is right, so someone has to pay for verification, and that spending should go where it teaches the system the most. That is a more disciplined version of the throw-it-at-several-models-and-vote trick that multi-agent tools already lean on.

Still, this is one arXiv paper measured on benchmark tasks, not production workloads with messy tool calls and ambiguous ground truth. Whether the savings survive contact with a real agent pipeline is the open question.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →