A new benchmark says today's AI agents are nowhere near ready to act as your hands and eyes in the real world.
Researchers built EgoBench, a benchmark of 1,590 tasks grounded in first-person video across five everyday scenarios. Each task forces an agent to combine visual perception with multi-step tool use, while a simulated user pushes back with naturalistic feedback mid-task. The team tested eight video-capable multimodal AI agents across three interaction modes and scored them with a strict pass/fail validation system rather than looser similarity checks. The best agent cleared only 34.95% of tasks on average.
That number matters because these are exactly the skills companies are racing to ship in phone assistants and smart glasses: see what's happening, decide what tool to call, and adjust when a person interrupts. A one-third success rate suggests the gap between demo reels and dependable autonomy is still wide, particularly once reasoning has to survive contact with a real, talkative user.
Call it a reality check for the egocentric-AI hype cycle: strapping a camera to an agent doesn't make it capable, it just makes its failures easier to watch.