AI/ llm-inference · edge-computing · scheduling · ai-research

Researchers Speed Up Split LLM Inference With Smarter Batching

DySCo, a new edge-cloud runtime, batches mismatched LLM requests at a shared depth, lifting throughput up to 275% over naive scheduling in tests.

A new scheduler lets phones and cloud servers split a single AI model's work without stalling between them.

Researchers built DySCo, a runtime for edge-cloud LLM inference that splits a model across resource-constrained devices and cloud GPUs. It keeps each device's key-value cache local and adds dyForward, a layer-range executor that runs chunks of a model's layers from shards already resident in memory, skipping weight reloads. When multiple edge devices hit the cloud at once, its depth-synchronized batching (DSB) advances mismatched requests to a shared depth and batches their common remaining computation. Tested across different devices, two model families, and both local and wide-area networks at an average of eight concurrent requests, DSB improved throughput by 275% over basic first-in-first-out queuing, 48% over exact-match batching, and 79% over round-robin interleaving.

The real contribution here isn't a faster model, it's an accounting of where split inference actually loses time. The researchers found that the idle gaps created by shuttling computation between edge and cloud add up to 25ms of extra latency per decoding step - before any queuing delay is even counted. That reframes the bottleneck in on-device AI assistants: it's not raw GPU throughput, it's the handoff tax between devices with mismatched split points.

It's a scheduling paper, not a new model - but scheduling is quietly where most inference speed gets won or lost.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →