Researchers have built a benchmark specifically to catch robot AI models forgetting things they just saw.
The new benchmark, called MIKASA-Robo-VLA, gives vision-language-action (VLA) models - the AI systems that watch a camera feed, read an instruction, and decide how a robot arm should move - 90 manipulation tasks to complete. Eighty of those tasks hide a visual cue the robot needs to remember after it disappears from view; the other 10 keep the cue visible throughout, as a control group. The project expands an earlier 32-task suite called MIKASA-Robo, adds language instructions across the board, and ships 22,500 example robot trajectories covering 10 different types of memory challenge. In 28 of the 70 tasks with a timed information gap, that gap outlasts the 16-frame lookback window of the widest-context VLA model the researchers surveyed.
That's the real finding here: most VLA models only glance at the last few camera frames, so they're built to forget. When the researchers fine-tuned a baseline model called pi-0.5 on 14 tasks - fed only current images and the robot's own joint position, no memory module, no history - it succeeded on just 21.1% of tasks on average.
The gap matters because real warehouses and kitchens are full of cues that vanish - a door that closes, an object that gets covered - and a robot that can't remember them is a robot you can't trust to work alone.