A new speculative-decoding technique called DScale squeezes more speed out of large language model serving without touching the model itself.
Researchers behind DScale built a lightweight add-on for block-diffusion speculative decoding, a method where a small drafter model guesses several tokens ahead and a verifier checks them in one pass. Instead of retraining or recalibrating the drafter, DScale adds a separate 112,000-parameter predictor, reshapes verification into path-aware tiles to cut wasted computation, and dynamically allocates how much verification budget each request gets. Tested on Nvidia A100-40GB GPUs running Qwen3-8B and Qwen3-4B across four datasets at concurrency levels of 8 to 32, the method posted geometric-mean throughput gains of 43.9% and 48.8% over a system called DFlash, 22.2% and 37.7% over DSpark, and 24.4% and 32.0% over Domino. On the GSM8K math benchmark, full decode-step time dropped 30.8% to 52.5% compared with DFlash.
Speculative decoding is already the standard trick for making LLM inference cheaper, but it tends to break down as more requests pile up at once, wasting compute on padding and rejected guesses. DScale's pitch is that it fixes that specifically at higher concurrency, the exact condition real production servers run under, without requiring a new drafter model for every target model.
The gains are measured against DFlash, DSpark, and Domino on one GPU generation, in a paper the authors wrote themselves, not verified by outside benchmarks or confirmed on newer H100 or B200 hardware.