[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-tree-based-speculative-decoding-boosts-deepseek-v4-speed-up-to-185":10,"sections":45},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":34,"tags":35,"sources":40,"feedback":44,"feedback_at":22,"cost_usd":44,"total_tokens":44},7376,"tree-based-speculative-decoding-boosts-deepseek-v4-speed-up-to-185","Tree-Based Speculative Decoding Boosts DeepSeek-V4 Speed Up to 18.5%","A new tree-structured decoding method speeds up DeepSeek-V4 inference by up to 18.5 percent, with gains shrinking at the smallest speculation budgets.","Speculative decoding just got a tree upgrade for DeepSeek-V4, and it actually pays off in practice, not just on paper.\n\nResearchers adapted tree-structured speculative decoding, a technique that keeps multiple candidate token sequences alive from a shared starting point, to the DeepSeek-V4-Flash pipeline. The catch: DeepSeek-V4's compressed attention system squeezes context into a shared state, which breaks when different branches need different states. The team fixed this with three additions - branch-aware verification, temporary state isolation, and a refresh step for accepted paths - so branches stop corrupting each other's compressed state. Tested across draft budgets of 5 to 8, batch sizes from 1 to 64, and three benchmarks (GSM8K, MBPP, ShareGPT), the tree method accepted more tokens per step than a matched linear baseline in every setting, and improved throughput in nearly all configurations, by up to 18.5 percent.\n\nThat 18.5 percent isn't evenly distributed. Gains grow with the speculation budget and matter most on unpredictable workloads at small-to-medium batch sizes, which is exactly where inference cost tends to add up fastest for anyone serving these models at scale. At the smallest budget, the speedup is close to a rounding error.\n\nThere's also a plateau: past a certain budget, throughput stops climbing even though the accepted token length keeps rising, meaning extra correct guesses stop translating into wall-clock savings. That's a useful check against reading \"more accepted tokens\" as \"always faster.\" This is still an arXiv paper, not a shipped feature - the real test is whether inference providers bother wiring it into their serving stacks.","[\"speculative-decoding\",\"deepseek\",\"llm-inference\",\"ai-research\"]","2026-09-23T04:00:00.000Z","2026-09-23T10:21:43.712Z","2026-09-23T10:21:48.154Z","published",null,[24,30],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"The dek claims 'cutting inference latency by up to 18.5 percent,' but the source and article body only measure a throughput improvement of up to 18.5% (tokens processed per unit time), not a latency reduction — fix the dek to accurately describe the throughput gain rather than substituting the word 'latency.'","resolved",{"id":31,"reviewer":26,"round":32,"reason":33,"status":29},"editor-r2",2,"The dek now correctly describes throughput (resolving editor-r1), but the headline states the gain flatly as 'by 18.5%' while the dek and body correctly qualify it as 'up to 18.5 percent' with gains marginal at the smallest budget — fix the headline to say 'up to 18.5%' so it doesn't overstate a typical result.","ai",[36,37,38,39],"speculative-decoding","deepseek","llm-inference","ai-research",[41],{"name":42,"url":43},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.24698",0,{"sections":46},[47,50,54,58,63,67,72,77,82,87,92,97,102,107],{"name":48,"slug":34,"count":49,"latest_published_at":18},"AI",4344,{"name":51,"slug":52,"count":53,"latest_published_at":18},"Security","security",713,{"name":55,"slug":56,"count":57,"latest_published_at":18},"Policy","policy",370,{"name":59,"slug":60,"count":61,"latest_published_at":62},"Deals","deals",206,"2026-09-23T09:43:46.000Z",{"name":64,"slug":65,"count":66,"latest_published_at":18},"Hardware","hardware",169,{"name":68,"slug":69,"count":70,"latest_published_at":71},"Science","science",134,"2026-09-23T09:00:00.000Z",{"name":73,"slug":74,"count":75,"latest_published_at":76},"Consumer Tech","consumer-tech",110,"2026-09-22T20:00:00.000Z",{"name":78,"slug":79,"count":80,"latest_published_at":81},"Software","software",81,"2026-09-23T09:56:13.000Z",{"name":83,"slug":84,"count":85,"latest_published_at":86},"Dev Tools","dev-tools",79,"2026-09-22T22:21:13.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"Startups","startups",65,"2026-09-22T22:06:48.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"Gaming","gaming",45,"2026-09-22T15:35:06.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"General","general",43,"2026-09-21T23:48:56.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"Reviews","reviews",27,"2026-09-22T13:00:00.000Z",{"name":108,"slug":109,"count":110,"latest_published_at":111},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]