Researchers have a new trick for mobile AI agents that forget how they solved yesterday's problem: reuse only the parts of memory that still apply, not the whole record.
DeltaReplay tackles a real limitation in memory-augmented GUI agents, the software that taps through apps to book flights, fill forms, or navigate settings menus. These agents store past successful run-throughs and replay them for similar tasks, but repeating a task rarely means an exact match — a new task might share three of five steps with an old record, use different parameters, or match nothing at all. DeltaReplay stores past runs as a graph of screens and actions, then checks each step against the new task and current screen to decide whether to follow it exactly, swap in new parameters, or hand control back to the underlying agent. On the AndroidWorld and SPA-Bench benchmarks, that step-by-step judgment raised task success rates by up to 10.3 and 25.0 percentage points over a baseline agent using the same model.
That's a meaningful gain for a problem that's quietly expensive: forcing an agent to follow irrelevant memory misleads it, while discarding anything that doesn't match perfectly throws away useful experience. Mobile GUI agents are the backbone of a lot of current agentic-AI demos, and their real-world reliability depends on handling tasks that are slightly different every time, not replaying a fixed script.
Benchmarks aren't production apps, and a 25-point jump on SPA-Bench says as much about that test suite's variety as about how DeltaReplay would hold up against, say, a banking app that redesigns its layout every quarter.