A new benchmark called VibeLifeBench just tested whether AI agents can act like real personal assistants over weeks, not minutes. The short answer: they can't, not yet.
Researchers built 200 long-horizon tasks spanning ten everyday-life domains, each unfolding as a scripted, multi-week timeline inside a simulated world of 22 mock services. The world keeps moving on its own clock, and many changes happen silently, so an agent has to actively check back in to notice them rather than wait to be told. Tasks are graded by fine-grained checks that look only at what the agent actually left behind: whether it hit the end state, whether it acted in time, and whether it respected constraints that were never explicitly stated. Seven frontier models were tested, and every single one scored low.
This matters because most AI agent benchmarks today test something much easier: a single self-contained request in a static environment, answered and forgotten. Real assistance looks nothing like that. It means holding a plan together for weeks, deciding on your own when to check in, when to act, and when to stay quiet, and catching changes nobody bothered to announce. That's a different skill than answering a well-formed prompt well.
The gap the researchers found isn't a rounding error to be closed with a slightly bigger model. It's a reminder that "agentic AI" demos, tidy multi-step tasks completed in one sitting, are still a long way from the messier job of actually running someone's life in the background.