A new technique lets AI language models skip their usual word-by-word generation and produce finished text in a single pass.
The paper describes Gumbel Straight Flow (GSF), a method that distills a pretrained autoregressive language model into a separate "flow map" model. The researchers found that the way autoregressive models convert random Gumbel noise into token sequences already traces straight, non-intersecting paths from noise to finished text. GSF trains a new model to follow those paths directly, with its output speed calibrated against the original model's own predictions. Across pretraining and downstream benchmarks, the authors report that GSF beats existing few-step text generation baselines.
Standard language models generate one token at a time, which is the real reason chatbot replies stream out gradually and cost compute for every step. Collapsing that process into one or a handful of steps attacks the slowdown directly, without requiring a smaller or less capable model. GSF joins a growing body of research, including diffusion-style language models from other labs, aimed at breaking autoregressive generation's token-by-token habit.
The paper offers no concrete latency or cost numbers, only comparisons to other few-step methods, so treat "outperforms" here as a benchmark claim rather than a shipped speedup.