A new inference scheduler for AI models slashes worst-case latency on one GPU setup, then falls apart as soon as you scale up.
Researchers built Decode-Latency Feedback Prefill (DLFP), a model-free controller for serving large language models that targets a specific bottleneck: when a new long prompt arrives, processing it ("prefill") can stall responses already being generated ("decode") for other users. DLFP only throttles prefill work that overlaps active decoding, using the delay observed in the previous scheduling cycle as feedback to resize the next prefill chunk. Tested in the vLLM serving framework on a single A100 80GB GPU running Qwen3-0.6B, it cut P99 inter-token latency by an average of 27.7% across three 100-request trials, with no change in output quality and no failures. That gain came at a cost: mean P99 time to first token rose 34.8%, though it stayed within the declared service-level target.
The more interesting part is the generalization failure the researchers report. The same controller delivered no benefit on larger models, Qwen3-8B and Qwen3-32B, or on a two-GPU tensor-parallel setup, because its feedback signal is only a proxy for scheduler timing, not actual GPU completion time. That is a useful, honest negative result in a field that rarely publishes what did not work, and it points directly at what a production-ready version needs: a controller timed to actual GPU iteration completion rather than scheduler-call intervals.
That's an unusually candid limitation section, and it matters more than the headline latency number ever will.