A new robot-training method skips numeric reward hacking by turning demonstration videos into formal logic instead of a single score.
Researchers built Video2STL, a framework that watches an observation-only video, no robot needed, no action labels, and has a vision-language model translate it into a Signal Temporal Logic specification: a formal, checkable description of what happens and when. The VLM works out the task structure, then real robot trajectories fill in the numerical thresholds and timing bounds. Short-horizon parts of the spec give the robot dense, step-by-step rewards, while a separate long-horizon check grants one-time credit for staying on a valid path toward the goal. Because the format doesn't care what's doing the moving, the team also fed it videos of humans and animals and had a robot learn from those directly.
Most video-to-robot approaches either compress a video into a single similarity score or have an AI write reward code by hand. Both are hard to audit and easy to game. Making the logic explicit and inspectable is the real selling point: you can read the temporal spec and know exactly what behavior is being rewarded, instead of trusting a black-box number.
On four manipulation tasks the method reported 85.8% success-once versus 81.5% for a standard dense-reward baseline and 65.0% for a reward-code baseline called Text2Reward, and on quadruped walking it hit 100% success across a range of speeds. Solid numbers, but it is one paper's benchmark on its own tasks, and formal-logic approaches to robotics have a long history of working beautifully in the lab and buckling on messy real-world floors.