[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-researchers-speed-up-ai-speculative-decoding-by-27-percent":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},8580,"researchers-speed-up-ai-speculative-decoding-by-27-percent","Researchers Speed Up AI Speculative Decoding by 27 Percent","A new unpublished arXiv preprint proposes DSpine, a drafting method that speeds parallel token prediction for AI chatbots without peer review yet.","A new preprint claims a tweak to how AI models draft text answers could speed up chatbot responses by more than a quarter.\n\nThe paper, posted to arXiv on September 30, 2026 as arXiv:2609.36173v1, is an unpublished preprint that has not been peer-reviewed. It proposes a system called DSpine, which speeds up speculative decoding, a technique where an AI model drafts several possible next words at once instead of one at a time. Prior methods like DFlash generate parallel drafts but only check consistency between them after the fact, which can shorten how many tokens actually get accepted. DSpine instead threads predictions between neighboring positions at every layer of the network, so each token's guess can inform the next one earlier in the process.\n\nSpeculative decoding is one of the main levers companies pull to make large language models cheaper and faster to run, since generating text one token at a time is slow. The paper's authors report DSpine lifted the average accepted draft length on Qwen3-8B from 3.77 to 4.82 tokens (a 27.8% jump) and delivered 23.3% higher throughput than DFlash in serving tests using the SGLang framework.\n\nThose are the authors' own numbers on their own benchmarks, tested on two mid-size open models, not an independent audit, so treat the percentages as a claim rather than a verdict until other labs try to reproduce them.","[\"ai\",\"speculative-decoding\",\"llm-inference\",\"research\"]","2026-09-30T04:00:00.000Z","2026-09-30T12:40:27.743Z","2026-09-30T12:40:33.707Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Add explicit sourcing—name the arXiv paper, its posting date\u002FID, and that it's an unpublished (non-peer-reviewed) preprint—since the draft states all figures without ever attributing them to a source.","resolved","ai",[30,32,33,34],"speculative-decoding","llm-inference","research",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.36173",0,{"sections":41},[42,45,49,53,58,63,68,73,78,82,87,92,97,102],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",5105,{"name":46,"slug":47,"count":48,"latest_published_at":18},"Security","security",785,{"name":50,"slug":51,"count":52,"latest_published_at":18},"Policy","policy",417,{"name":54,"slug":55,"count":56,"latest_published_at":57},"Deals","deals",284,"2026-09-29T21:00:00.000Z",{"name":59,"slug":60,"count":61,"latest_published_at":62},"Hardware","hardware",194,"2026-09-29T13:16:04.000Z",{"name":64,"slug":65,"count":66,"latest_published_at":67},"Science","science",154,"2026-09-28T13:19:18.000Z",{"name":69,"slug":70,"count":71,"latest_published_at":72},"Consumer Tech","consumer-tech",142,"2026-09-29T18:38:03.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":77},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":79,"slug":80,"count":81,"latest_published_at":18},"Dev Tools","dev-tools",90,{"name":83,"slug":84,"count":85,"latest_published_at":86},"Startups","startups",83,"2026-09-29T21:51:36.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"General","general",49,"2026-09-28T16:44:57.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]