[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-edge-ai-servers-get-a-speed-trick-to-split-llm-users-in-two":10,"sections":35},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":30,"feedback":34,"feedback_at":22,"cost_usd":34,"total_tokens":34},7430,"edge-ai-servers-get-a-speed-trick-to-split-llm-users-in-two","Edge AI Servers Get a Speed Trick to Split LLM Users in Two","A new scheduling framework called BALANCE runs slow, memory-light and fast, memory-heavy LLM inference side by side on the same edge server to serve more users.","Running AI chatbots on nearby cell towers instead of distant data centers is faster, but a new paper shows edge servers have been picking one bad tradeoff or another.\n\nResearchers describe two ways to run large language model inference at the network edge. Autoregressive decoding generates one token at a time, which is slow but memory-cheap. Speculative decoding speeds things up by having a small model draft several tokens for the big model to check at once, but that requires loading a second model into memory. The new framework, called BALANCE, runs both modes simultaneously on one edge server, hosting a small and a large model together, and assigns each user to whichever mode fits their latency needs and the server's available memory. Because deciding the optimal mix is computationally NP-hard, the team built a faster approximate algorithm that splits the problem in two and still guarantees a solution within a known distance of optimal.\n\nThis matters because edge computing only works if it can serve lots of users on hardware with real memory limits, not the effectively unlimited pool of a cloud data center. Picking one decoding strategy for everyone means either wasting memory speculative decoding needs or making users wait through slow token-by-token generation. Letting a server run both at once, and shifting the mix based on demand, is a scheduling problem more than a modeling one - which is a more mundane but more immediately deployable fix than training a smaller model.\n\nThe paper reports throughput gains over using either method alone, though as with most inference-scheduling research, the real test is whether telecom operators building 5G and 6G edge infrastructure adopt something like it rather than shipping a simpler, less optimal system.","[\"edge computing\",\"llm inference\",\"speculative decoding\",\"ai infrastructure\"]","2026-09-23T04:00:00.000Z","2026-09-23T13:06:10.687Z","2026-09-23T13:06:16.679Z","published",null,[],"ai",[26,27,28,29],"edge computing","llm inference","speculative decoding","ai infrastructure",[31],{"name":32,"url":33},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.05926",0,{"sections":36},[37,41,45,50,55,60,65,70,75,80,85,90,95,100],{"name":38,"slug":24,"count":39,"latest_published_at":40},"AI",4347,"2026-09-23T12:00:00.000Z",{"name":42,"slug":43,"count":44,"latest_published_at":18},"Security","security",713,{"name":46,"slug":47,"count":48,"latest_published_at":49},"Policy","policy",371,"2026-09-23T12:00:43.000Z",{"name":51,"slug":52,"count":53,"latest_published_at":54},"Deals","deals",211,"2026-09-23T13:00:46.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Hardware","hardware",170,"2026-09-23T11:59:23.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Science","science",134,"2026-09-23T09:00:00.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":69},"Consumer Tech","consumer-tech",110,"2026-09-22T20:00:00.000Z",{"name":71,"slug":72,"count":73,"latest_published_at":74},"Software","software",81,"2026-09-23T09:56:13.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":79},"Dev Tools","dev-tools",79,"2026-09-22T22:21:13.000Z",{"name":81,"slug":82,"count":83,"latest_published_at":84},"Startups","startups",65,"2026-09-22T22:06:48.000Z",{"name":86,"slug":87,"count":88,"latest_published_at":89},"Gaming","gaming",45,"2026-09-22T15:35:06.000Z",{"name":91,"slug":92,"count":93,"latest_published_at":94},"General","general",43,"2026-09-21T23:48:56.000Z",{"name":96,"slug":97,"count":98,"latest_published_at":99},"Reviews","reviews",27,"2026-09-22T13:00:00.000Z",{"name":101,"slug":102,"count":103,"latest_published_at":104},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]