Researchers built a benchmark that turns cooking videos into simulated kitchens where AI agents can actually mess up a recipe.
Ego2World takes annotated egocentric cooking footage and compiles it into executable planning environments. A compiler maps each recorded step and object to symbolic action rules, persistent world states, and task conditions, so an agent's chosen actions get executed and checked against outcomes rather than just scored against the original video. The system tracks world state and an agent's belief about that state separately, letting researchers isolate how memory and observation choices affect planning. Testing six planners on 105 tasks, the researchers found that actions the system accepted as valid often still left the task unfinished, and execution traces could tell apart a run that stalled, one that partially finished, and one that ran to completion without hitting the goal.
This closes a specific gap between passively studying human-recorded video and testing actual decision-making: agents are graded on the consequences of their own choices under incomplete information, which looks a lot more like how a kitchen robot or assistant would really have to operate. A paired study using Qwen-Plus also complicates the instinct to just bolt memory onto a model: persistent belief tracking cut unnecessary visual lookups by over 90% and made actions a few percentage points more valid, but burned more tokens and didn't make agents any more likely to actually finish the job.
In other words, remembering what you've already seen in the kitchen makes you a more efficient cook, not necessarily a more successful one.