AI/ ai agents · benchmarks · multimodal ai · egocentric video

New Benchmark Finds AI Agents Struggle With Tool Use

EgoBench tested eight video AI agents on real-world tool tasks, and the best one succeeded only 34.95% of the time.

A new benchmark says today's AI agents are nowhere near ready to act as your hands and eyes in the real world.

Researchers built EgoBench, a benchmark of 1,590 tasks grounded in first-person video across five everyday scenarios. Each task forces an agent to combine visual perception with multi-step tool use, while a simulated user pushes back with naturalistic feedback mid-task. The team tested eight video-capable multimodal AI agents across three interaction modes and scored them with a strict pass/fail validation system rather than looser similarity checks. The best agent cleared only 34.95% of tasks on average.

That number matters because these are exactly the skills companies are racing to ship in phone assistants and smart glasses: see what's happening, decide what tool to call, and adjust when a person interrupts. A one-third success rate suggests the gap between demo reels and dependable autonomy is still wide, particularly once reasoning has to survive contact with a real, talkative user.

Call it a reality check for the egocentric-AI hype cycle: strapping a camera to an agent doesn't make it capable, it just makes its failures easier to watch.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →