[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-researchers-speed-up-ai-model-serving-with-new-decoding-method":10,"sections":46},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":35,"tags":36,"sources":41,"feedback":45,"feedback_at":22,"cost_usd":45,"total_tokens":45},8638,"researchers-speed-up-ai-model-serving-with-new-decoding-method","Researchers Speed Up AI Model Serving With New Decoding Method","DScale boosts Qwen3-4B decoding throughput by up to 48.8 percent over rival DFlash without changing the underlying model weights.","A new speculative-decoding technique called DScale squeezes more speed out of large language model serving without touching the model itself.\n\nResearchers behind DScale built a lightweight add-on for block-diffusion speculative decoding, a method where a small drafter model guesses several tokens ahead and a verifier checks them in one pass. Instead of retraining or recalibrating the drafter, DScale adds a separate 112,000-parameter predictor, reshapes verification into path-aware tiles to cut wasted computation, and dynamically allocates how much verification budget each request gets. Tested on Nvidia A100-40GB GPUs running Qwen3-8B and Qwen3-4B across four datasets at concurrency levels of 8 to 32, the method posted geometric-mean throughput gains of 43.9% and 48.8% over a system called DFlash, 22.2% and 37.7% over DSpark, and 24.4% and 32.0% over Domino. On the GSM8K math benchmark, full decode-step time dropped 30.8% to 52.5% compared with DFlash.\n\nSpeculative decoding is already the standard trick for making LLM inference cheaper, but it tends to break down as more requests pile up at once, wasting compute on padding and rejected guesses. DScale's pitch is that it fixes that specifically at higher concurrency, the exact condition real production servers run under, without requiring a new drafter model for every target model.\n\nThe gains are measured against DFlash, DSpark, and Domino on one GPU generation, in a paper the authors wrote themselves, not verified by outside benchmarks or confirmed on newer H100 or B200 hardware.","[\"speculative decoding\",\"llm inference\",\"ai research\",\"gpu\"]","2026-09-30T04:00:00.000Z","2026-09-30T16:28:38.447Z","2026-09-30T16:28:44.633Z","published",null,[24,30],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Attribute the throughput and decode-time figures explicitly to the source (paper title, arXiv ID 2609.37532, and that it's an unreviewed preprint) since none of the numbers are currently sourced, and rewrite the final paragraph so the piece doesn't end on a bare caveat.","resolved",{"id":31,"reviewer":32,"round":33,"reason":34,"status":29},"publisher-r2","publisher",2,"The dek cites a specific '48.8 percent' improvement figure that never appears anywhere in the body, which only reports ranges of 22-49 percent and 30.8-52.5 percent, making the dek's precise number internally inconsistent with the article text.","ai",[37,38,39,40],"speculative decoding","llm inference","ai research","gpu",[42],{"name":43,"url":44},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.37532",0,{"sections":47},[48,51,55,59,64,69,73,78,83,87,92,97,102,107],{"name":49,"slug":35,"count":50,"latest_published_at":18},"AI",5180,{"name":52,"slug":53,"count":54,"latest_published_at":18},"Security","security",791,{"name":56,"slug":57,"count":58,"latest_published_at":18},"Policy","policy",417,{"name":60,"slug":61,"count":62,"latest_published_at":63},"Deals","deals",284,"2026-09-29T21:00:00.000Z",{"name":65,"slug":66,"count":67,"latest_published_at":68},"Hardware","hardware",194,"2026-09-29T13:16:04.000Z",{"name":70,"slug":71,"count":72,"latest_published_at":18},"Science","science",155,{"name":74,"slug":75,"count":76,"latest_published_at":77},"Consumer Tech","consumer-tech",142,"2026-09-29T18:38:03.000Z",{"name":79,"slug":80,"count":81,"latest_published_at":82},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":84,"slug":85,"count":86,"latest_published_at":18},"Dev Tools","dev-tools",90,{"name":88,"slug":89,"count":90,"latest_published_at":91},"Startups","startups",83,"2026-09-29T21:51:36.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"General","general",49,"2026-09-28T16:44:57.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":108,"slug":109,"count":110,"latest_published_at":111},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]