[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-longspark-keeps-ai-speculative-decoding-fast-as-context-grows":10,"sections":40},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":35,"feedback":39,"feedback_at":22,"cost_usd":39,"total_tokens":39},8614,"longspark-keeps-ai-speculative-decoding-fast-as-context-grows","LongSpark Keeps AI Speculative Decoding Fast as Context Grows","A new drafter design keeps speculative decoding's cost flat as context grows, fixing an inefficiency that undercuts long-context speedups.","A new paper proposes a fix for one of speculative decoding's quiet cost problems: the helper model that's supposed to speed things up gets more expensive the longer your prompt gets.\n\nThe work, posted to arXiv and not yet peer reviewed, describes a system called LongSpark. Speculative decoding speeds up text generation by having a small *drafter* model guess several tokens ahead, which the full model then checks in a single pass instead of generating token by token. The problem: as conversations or documents grow, today's best drafters have to carry a growing memory of the entire prefix, so they slow down too - quietly eating into the speed gain they exist to provide. LongSpark's drafter instead pulls fixed-size snapshots from the main model's own verification step, so its cost stays flat no matter how long the input gets.\n\nThat matters because long-context work - RAG pipelines, coding assistants scanning whole codebases, hours-long chat sessions - is exactly where speculative decoding's edge has been quietly disappearing, by the paper's own account of the problem. A drafter whose cost doesn't scale with context is a real lever for anyone running inference at scale, not just a benchmark trick, if it holds up in production serving stacks outside the authors' own tests.\n\nSpeculative decoding only became a mainstream inference trick in the last couple of years, and most efficiency gains since then have come from tuning drafter architecture rather than questioning what a drafter needs to remember at all - which is the more interesting move here than the raw speed numbers.","[\"ai\",\"llm-inference\",\"speculative-decoding\",\"dev-tools\"]","2026-09-30T04:00:00.000Z","2026-09-30T15:02:46.132Z","2026-09-30T15:02:52.087Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Move the single-source caveat earlier (e.g., into the paragraph introducing the paper) and rewrite the final paragraph so the piece doesn't end on a bare caveat sentence — close with a substantive point instead of an unfinished-feeling disclaimer.","resolved","ai",[30,32,33,34],"llm-inference","speculative-decoding","dev-tools",[36],{"name":37,"url":38},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.37029",0,{"sections":41},[42,45,49,53,58,63,68,73,78,81,86,91,96,101],{"name":43,"slug":30,"count":44,"latest_published_at":18},"AI",5135,{"name":46,"slug":47,"count":48,"latest_published_at":18},"Security","security",788,{"name":50,"slug":51,"count":52,"latest_published_at":18},"Policy","policy",417,{"name":54,"slug":55,"count":56,"latest_published_at":57},"Deals","deals",284,"2026-09-29T21:00:00.000Z",{"name":59,"slug":60,"count":61,"latest_published_at":62},"Hardware","hardware",194,"2026-09-29T13:16:04.000Z",{"name":64,"slug":65,"count":66,"latest_published_at":67},"Science","science",154,"2026-09-28T13:19:18.000Z",{"name":69,"slug":70,"count":71,"latest_published_at":72},"Consumer Tech","consumer-tech",142,"2026-09-29T18:38:03.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":77},"Software","software",91,"2026-09-25T20:55:00.000Z",{"name":79,"slug":34,"count":80,"latest_published_at":18},"Dev Tools",90,{"name":82,"slug":83,"count":84,"latest_published_at":85},"Startups","startups",83,"2026-09-29T21:51:36.000Z",{"name":87,"slug":88,"count":89,"latest_published_at":90},"General","general",49,"2026-09-28T16:44:57.000Z",{"name":92,"slug":93,"count":94,"latest_published_at":95},"Gaming","gaming",48,"2026-09-25T18:35:21.000Z",{"name":97,"slug":98,"count":99,"latest_published_at":100},"Reviews","reviews",31,"2026-09-28T14:31:34.000Z",{"name":102,"slug":103,"count":104,"latest_published_at":105},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]