A new benchmark finally tests whether AI coding agents can get embedded firmware actually working, not just write code that looks plausible.
The benchmark, described in a paper posted to arXiv, gives AI agents a plain-text engineering brief, a constrained workspace, and a simulated ESP32 board to build against. Agents have to turn requirements into working code, test it themselves, and keep iterating until the device behaves correctly, rather than producing one static file and stopping. The suite covers five embedded-control tasks across four feedback setups, ranging from one-shot generation to CI-style pass-fail signals to detailed oracle feedback. Researchers ran seven GPT-family and Qwen-family model configurations through the full suite three times each, for 420 total runs.
The results push back on the idea that any current model has embedded coding solved. gpt-5.4 topped the leaderboard but still did not saturate the benchmark, meaning it kept failing some tasks regardless of how much feedback it received. Qwen3.5-27B was the strongest model you could plausibly run locally, while smaller local models lost accuracy fast and used feedback far less efficiently as iterations piled up.
Embedded firmware has always been the place where confident-sounding code meets physical reality: timing, sensors, hardware state. A benchmark that forces agents to prove closed-loop correctness, not just pass a code review, is a reasonable way to find out which models actually hold up there.