AI/ ai · llm-inference · speculative-decoding · research

DRelay Speeds AI Text Generation by Fixing Bad Guesses Early

A new technique called DRelay lets AI language models correct early prediction errors mid-stream, speeding up text generation by up to 17 percent in tests.

Researchers have found a way to make AI text generation faster by letting models fix their own early mistakes instead of throwing away good guesses that come after them.

Speculative decoding speeds up large language models by using a small draft model to guess several tokens at once, which the big model then checks in parallel. The catch: a guess only counts if everything before it in the sequence was also right. One wrong token early in the batch, and every correct guess after it gets discarded too. DRelay, a new method described in a recent paper, adds a step that reads signals from the whole draft batch and swaps out that early wrong guess before the big model verifies anything, so good predictions further down the line get a chance to survive.

That's a different lever than most speedup work in this area, which focuses on making the draft model guess better in the first place rather than repairing its mistakes afterward. A repair step attacks the specific bottleneck, how many tokens survive per batch, that caps every parallel-drafting method regardless of how good the draft model is.

In tests across eight benchmarks on an H800 GPU, DRelay beat three existing methods, DFlash, Domino, and DSpark, by 8 to 17 percent in end-to-end speed under SGLang serving. Solid numbers, though the kind of lab benchmark gain that has a habit of shrinking once it meets a real production serving stack.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →