A new AI system builds physics simulations from nothing but a text prompt, then checks its own work and fixes the mistakes.
Researchers built Text2Sim, an agentic pipeline that converts a plain-text request into a working, editable physics simulation, including assets, layout, physical parameters, motion, and rendering. It runs on the Genesis simulation engine and splits the job across a hierarchy of AI agents: a Planner that breaks down the request, specialized Writer agents that generate the code and assets, and an independent Critic that reviews the output. To fix errors, the system draws on a library of Debug Cards, compact troubleshooting notes distilled from graphics expert demonstrations, which guide execution-based repairs. The team evaluated it on 42 held-out prompts covering rigid bodies, articulated mechanisms, deformable objects, and cloth.
Building simulations by hand is slow, fiddly work: every asset, parameter, and motion has to be coded and debugged separately. Text2Sim beat four existing state-of-the-art systems on both automatic physical and visual scoring, and in blind user studies significantly more participants preferred its output over the baselines. The researchers plan to release the code, the Debug Card library, and a dataset pairing text prompts with finished simulations, which could make it easier for other teams to build training data without hand-coding every scene.
The real test is whether text-in, working-simulation-out holds up on requests messier than the curated 42-prompt benchmark the team used to grade itself.