A new academic benchmark puts a number on something device makers have mostly argued by vibes: how much you give up when an AI agent runs locally instead of in the cloud.
Researchers released AgBench, an open benchmark suite for evaluating agentic AI systems, the kind that plan, call tools, and execute multi-step tasks, across local-only, hybrid, and cloud-only setups on personal devices. The team ran agentic workloads across these architectures and collected more than 162 million data points on task success, latency, cloud API cost, and how much sensitive data left the device. Local-only execution can handle many agent tasks, but it loses to cloud-only execution on success rate and speed, and the gap widens as more tasks run at once. Hybrid setups, which split work between device and cloud, can boost success rates, but their cost and privacy exposure depend entirely on how that work gets divided.
This is a direct challenge to the current pitch for on-device AI assistants, which implies local processing is a free upgrade: more private, equally capable, no catch. AgBench's data says the catch is real: you trade cloud API costs and data exposure for lower success rates and higher latency, and no single architecture wins on every metric at once. For developers and device makers, the lesson is that deployment choices need to match the actual task and hardware, not a blanket claim that local or cloud is simply better.
That is a duller message than "runs entirely on your phone," but it is closer to the truth than most on-device AI marketing slides let on.