AI/ ai · gui-agents · benchmarks · cross-device

Benchmark Exposes Cross-Device Gaps in GUI AI Agents

A new arXiv paper introduces JarvisGUI, a benchmark showing GUI AI agents struggle across Android, Windows, and Ubuntu devices.

A new benchmark suggests today's GUI AI agents fall apart the moment a task moves from your phone to your laptop.

The benchmark, called JarvisGUI and described in a paper posted to arXiv, tests AI agents on workflows that require juggling multiple devices and operating systems, such as moving a file from an Android phone to a Windows desktop or carrying context over to an Ubuntu machine. Instead of relying on a fixed set of tasks, JarvisGUI treats every task as an input-output transformation and automatically generates new multi-step, cross-device combinations, so agents cannot simply memorize a static test set. The paper's authors ran current open-source GUI agents through these virtual, multi-OS environments and found they consistently lost track of shared state, misjudged context when switching platforms, and broke down on tasks that chained several dependent steps.

Most GUI-agent benchmarks test one device at a time, which is not how people actually use their gadgets. A screenshot taken on a phone often ends up pasted into a spreadsheet on a laptop, and JarvisGUI is one of the first evaluations built to catch an agent that fumbles exactly that handoff.

Until an agent can remember what happened on your phone once you switch to your laptop, the pitch of an AI assistant that runs your whole workflow stays a demo, not a product.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →