AI/ ai · on-device ai · memory · mobile

New Benchmark Tests AI Memory Across a Year of Phone Use

MobileMem is a new framework that tests whether on-device AI agents can remember, reason about, and adapt to a full year of a user's phone activity.

Researchers have built a benchmark for testing whether on-device AI assistants can actually remember your life, not just retrieve facts about it.

MobileMem, introduced in a newly published research paper, is a framework and benchmark grounded in a year's worth of simulated mobile experiences. It uses what the researchers call a knowledge-grounded synthesis pipeline to generate coherent, temporally consistent trajectories from user-app sessions, rather than relying on scattered logs. The benchmark covers both text and multimodal settings, testing multi-hop and temporal reasoning, knowledge updating, and inference of a user's implicit preferences. The stated goal is an assistant that remembers the past, understands the present, and adapts as a user's habits change.

Most memory benchmarks for AI agents test whether a model can retrieve one isolated fact from a conversation log. MobileMem instead treats memory as something built from lived experience over time, which is closer to what an on-device assistant would actually need to be useful day to day. That distinction matters as phone makers and AI labs keep promising assistants that know you, not just answer you.

Whether any assistant currently clears that bar is a separate question. Right now MobileMem is a way to measure the gap, not proof anyone has closed it.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →