A new survey puts a name and a map on the scattered field of robots that use AI foundation models to see, think, and move: Generative Physical Artificial Intelligence.
The paper, posted to arXiv on September 17, 2026, sorts current robot-AI research into five overlapping categories. Robot Foundation Models transfer skills across different machines. Vision-Language Action models handle perception and control in a single system. Large Behavior Models generate human-like movement, Diffusion Policy Models produce smoother multi-step actions, and World Foundation Models simulate physics to generate training data. The authors show these pieces feed each other: World Foundation Models generate the training data that Vision-Language Action and Diffusion Policy models learn from, while Robot Foundation Models let those learned skills move from one robot body to another.
This matters because robotics has spent years chasing a GPT moment - a single model that generalizes the way large language models did for text. The survey's honest answer is that no such model exists yet; instead the field is stitching together five specialist approaches and hoping they compose well together. That's a more useful map for engineers than another company's demo reel, even if it makes for a less exciting headline.
The paper covers examples from self-driving cars to hospital robots and industrial and humanoid systems, and it flags data-efficient learning, sim-to-real transfer, edge-compatible hardware, and safety frameworks as unsolved problems - academic-speak for still years away from your local hospital floor.