AI/ ai agents · benchmarks · simulink · engineering tools

Best AI Agent Scores 43 Out of 100 on Engineering Benchmark

SimuVerity tests AI agents on 101 Simulink engineering tasks, and even the top performer barely clears 40 percent on real engineering requirements.

A new benchmark says AI agents are nowhere close to generating engineering-grade simulation models.

Researchers built SimuVerity, a benchmark of 101 text-to-executable Simulink model-generation tasks spanning ten engineering domains. Each task comes with an executable-system profile that defines the engineering spec, plus four families of native simulation scenarios. A hierarchical evaluator checks whether a model actually gets delivered, runs natively, and clears a qualification bar, then scores qualified models across six dimensions: accuracy, output quality, mechanistic fidelity, control and causal integrity, operating-domain robustness, and dynamic response. The team ran six agent systems through it, and the best one managed an overall score of just 42.86 out of 100.

Most existing Simulink benchmarks just check whether a model compiles or resembles a reference - that says almost nothing about whether it would survive a real engineering review. SimuVerity's results point to two separate failure points: agents struggle to produce an implementation that even qualifies as real, and then struggle again to meet the multi-dimensional requirements once it does qualify. Some models that score well still come out visually disordered, a reminder that a passing grade doesn't mean the output is ready to hand off.

If 42.86 out of 100 is the current ceiling, engineering teams should treat AI-generated Simulink models as drafts, not deliverables.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →