A new serving system called FluidPD reassigns GPU workers between tasks on the fly, aiming to stop large language models from blowing their latency targets without buying more hardware.
Researchers split LLM serving into two phases: prefill, which processes the incoming prompt, and decode, which generates the response token by token. Most systems lock in a fixed ratio of workers for each phase, plus request routing on top of it. The problem is that real traffic does not respect that ratio. Demand swings between short bursts and sustained shifts in the prefill-to-decode balance, so a setup that is correctly tuned right now can end up starving one phase an hour later, even while other GPUs sit idle. FluidPD adds two mechanisms to fix that: FluidToken offloads a bounded slice of prefill work to decode workers when they have slack, and FluidRole reassigns entire workers between prefill and decode roles in place, without reloading the model or restarting the serving engine. Both decisions run off lightweight pressure indices meant to catch strain before it turns into a missed latency target.
On production Azure trace workloads, FluidPD reportedly improved SLO attainment by up to 94.6 percentage points over a static SGLang baseline. That matters because the usual fix for phase imbalance is autoscaling, which reacts slowly and needs spare GPUs sitting around. FluidPD's pitch is getting better latency reliability from the same fleet, which counts for more every quarter as inference costs, not training costs, become the dominant expense for companies running LLMs at scale.
Worth flagging: this is an arXiv preprint, not a deployed system, and a 94.6 percentage point jump against a single static baseline is the kind of number that tends to shrink once anyone tries to reproduce it against their own tuned setup.