AI agents still can't run a CAD program end to end, according to a new benchmark.
Researchers posted a paper on arXiv (arXiv:2609.37686), an unreviewed preprint published September 30, 2026, describing EngiWorld, a benchmark of 1,301 expert-curated tasks spanning six engineering domains: CAD, CAE, CAM, BIM, EDA, and 3D visualization. The tasks span 26 professional software platforms with both GUI and CLI interfaces, and cover six task types, from picking the right software to open-ended design work. The team built a verifier suite that checks whether an agent's output is geometrically valid, physically feasible, and rule-compliant, then scores partial credit against a design spec instead of just pass or fail. Seven frontier models were tested; the best scored an EngiScore of only 44.3, and multi-software attempts succeeded just 3.6% of the time.
Most AI benchmarks measure chat or code generation, tasks with a keyboard and a text box. Engineering software demands geometry, physics, and file formats that don't forgive a rounding error, so a 44.3 score is a reminder that "agentic" still means something narrower in the physical-design world than in the browser.
Until an agent can open a CAD file, adjust a tolerance, and hand a clean file to CAM without a human checking its work, professional engineering shops have little to fear from automation.