[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-study-maps-where-agentic-ai-workloads-actually-bottleneck":10,"sections":35},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":24,"tags":25,"sources":30,"feedback":34,"feedback_at":22,"cost_usd":34,"total_tokens":34},5351,"study-maps-where-agentic-ai-workloads-actually-bottleneck","Study Maps Where Agentic AI Workloads Actually Bottleneck","A new benchmark suite shows that in AI agent systems, sandboxes and memory, not the language model itself, are often the real drag on speed and cost.","Turns out the slowest part of your AI agent might not be the AI.\n\nA new paper introduces AgentSysBench, a benchmark and measurement toolkit built to study ten representative agentic applications, the kind of systems where an LLM juggles tools, external environments, and state that persists across long sessions. Testing across controlled deployments and real production traces, the researchers found that non-LLM components, like sandboxes and retrieval systems, dominate latency in half the applications tested. Sandbox memory alone peaked at 28 GB per session, and task latencies varied by up to 32x depending on whether a step was GPU-bound inference or CPU-bound execution. Production sessions also sat idle, holding state for minutes or hours between active steps, while a \"control-plane tax\" of auxiliary LLM calls and tool-schema overhead ate into useful compute.\n\nThis matters because most serving infrastructure is still built for a simpler world: one model, one prompt, one response. Agentic workloads break that assumption in ways that plain GPU scaling can't fix, since the bottleneck might be a slow web fetch or an idle sandbox eating memory, not the model itself. The paper's fixes are not speculative either. Task-aware serving cut latency by 29-40%, and caching redundant search and web-fetch calls, common across sessions, removed over a third of duplicate queries.\n\nIt's a useful reminder that agent hype has outpaced agent plumbing. The industry has spent two years optimizing token throughput; this paper argues the next bottleneck is everything downstream of the model, and nobody's built the caching layer for it yet.","[\"ai-agents\",\"llm-serving\",\"benchmarking\",\"infrastructure\"]","2026-08-18T04:00:00.000Z","2026-08-18T16:31:43.434Z","2026-08-18T16:31:55.333Z","published",null,[],"ai",[26,27,28,29],"ai-agents","llm-serving","benchmarking","infrastructure",[31],{"name":32,"url":33},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.15127",0,{"sections":36},[37,41,45,50,55,60,65,70,75,79,84,89,94,99],{"name":38,"slug":24,"count":39,"latest_published_at":40},"AI",3293,"2026-08-20T04:00:00.000Z",{"name":42,"slug":43,"count":44,"latest_published_at":40},"Security","security",435,{"name":46,"slug":47,"count":48,"latest_published_at":49},"Policy","policy",210,"2026-08-19T09:32:27.000Z",{"name":51,"slug":52,"count":53,"latest_published_at":54},"Deals","deals",179,"2026-06-29T20:02:07.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Hardware","hardware",140,"2026-08-19T18:25:42.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":64},"Consumer Tech","consumer-tech",95,"2026-08-18T16:05:00.000Z",{"name":66,"slug":67,"count":68,"latest_published_at":69},"Science","science",90,"2026-08-19T18:41:02.000Z",{"name":71,"slug":72,"count":73,"latest_published_at":74},"Software","software",73,"2026-08-18T07:51:50.000Z",{"name":76,"slug":77,"count":78,"latest_published_at":18},"Dev Tools","dev-tools",69,{"name":80,"slug":81,"count":82,"latest_published_at":83},"Startups","startups",47,"2026-08-19T19:13:46.000Z",{"name":85,"slug":86,"count":87,"latest_published_at":88},"Gaming","gaming",41,"2026-07-09T04:00:00.000Z",{"name":90,"slug":91,"count":92,"latest_published_at":93},"General","general",33,"2026-08-18T22:18:13.000Z",{"name":95,"slug":96,"count":97,"latest_published_at":98},"Reviews","reviews",20,"2026-06-24T12:00:01.000Z",{"name":100,"slug":101,"count":102,"latest_published_at":103},"How-To","how-to",6,"2026-06-16T09:00:00.000Z"]