Researchers have built a system called Scepsy that packs AI agent workloads onto GPU clusters more efficiently, squeezing out more throughput without the usual latency hit.
Agentic workflows chain together multiple large language models and tools, and their execution paths branch, fan out, or loop depending on the input, so total run time is hard to predict. Because these workflows typically call on more LLMs than a team has GPUs to run them on, GPUs end up oversubscribed. Scepsy's fix is to profile each LLM's share of total execution time, which the researchers found stays fairly consistent even when overall latency doesn't. It combines those profiles into what it calls an Aggregate LLM Pipeline, then searches for the best mix of fractional GPU shares, parallelism settings, and replica counts before placing the result on the cluster.
On the workflows tested, Scepsy delivered up to 2.5x higher throughput before hitting capacity and cut latency by as much as 3.3x compared to systems that tune each LLM in isolation or rely on engineers guessing at allocations. That's the real cost problem with agents in production: teams either overprovision GPUs to be safe or underprovision and eat the latency, because nobody has a good way to model a whole pipeline at once. A scheduler that treats the agent pipeline as the unit of optimization, rather than each model call, is the more sensible default as these systems move past single-prompt demos.
It's still a research paper testing itself against its own benchmarks, not a tool you can install today, so treat the multiples as a ceiling rather than a guarantee.