A new robot-control framework called URAI roughly triples task success by letting AI models write code instead of narrating every single movement.
Researchers built URAI (Universal Robot-Agent Interface), which pairs a programming agent that writes reusable robot tools with an execution agent that calls them. Each tool call runs a full motion - a reach, a grasp, a retreat - locally before handing control back to the model, so the AI decides what to do between actions instead of inside them. Across five RoboDojo simulation tasks with four different frozen execution agents, this setup pushed aggregate success from 18% under direct fingertip control to 53%. Scripting the same tool library in advance as one fixed program, rather than letting a model choose after each call, only reached 24%, versus 56% for two agents deciding step by step.
That gap is the real finding here: neither extreme works well. Reasoning through every joint angle is slow and token-hungry, but handing a program full control loses the model's judgment when a task goes sideways. Splitting the labor - code for motion, models for decisions - also made three of four agents 1.3-1.5 times faster with 1.5-1.7 times fewer output tokens, though one model, DeepSeek-V4-Flash, saw almost no cost change, so the efficiency gains aren't automatic across every architecture.
The tools also persisted across episodes without retraining the underlying model, which is the durable, reusable piece robotics has chased since code-as-policy approaches first got popular; seven real-world dual-arm tests, including cloth folding and tic-tac-toe, are a useful sanity check but still a long way from proof this holds up outside the lab.