AI can now generate a slick-looking web dashboard in seconds. Getting it to actually respond when you click a button is a different story.
A new benchmark called WebUIProof tests AI-generated web interfaces the way a real user would: by clicking buttons, dragging sliders, and checking whether the interface actually does what it is supposed to do, inside a headless browser. Researchers built the system to cover two task categories: everyday interfaces like dashboards and interactive tools, and harder 3D cases like particle systems, galaxy simulations, and physics dynamics. They ran eight commercial LLMs through it and found the interfaces frequently broke on interaction tests even when the page rendered without errors, with 3D simulations failing most often. The team then used the same pass-fail signal to train smaller open models, Qwen2.5 14B and MIMO 7B, with reinforcement learning, and reported fewer build failures and better functional completion as a result.
That gap between looking right and working right is the real story. Most code-generation benchmarks still grade on compile success or a screenshot, which rewards AI for producing something that resembles a working app rather than one that is. WebUIProof's interaction-test approach is a more honest yardstick, and the finding that 3D interfaces fail hardest should temper some of the recent hype around AI demos of particle systems and physics sims that make the rounds online.
A pretty screenshot has always been the easiest thing for AI to fake. Clicking the button is where the bluff gets called.