AI/ llms · hardware-verification · chip-design · ai-agents

Survey Charts How LLMs Are Entering Hardware Verification

A sweeping review finds LLMs excel at reasoning and orchestration in chip verification, but checking tool acceptance isn't the same as proving correctness.

A new survey takes stock of how large language models are being wired into the unglamorous but critical work of verifying chip designs before they become silicon.

The paper reviews work across SystemVerilog assertion generation, testbench and stimulus generation, bug localization and design repair, model checking, equivalence checking, and SAT/SMT solver optimization. It sorts the field by methodology, verification goal, tool interaction, benchmark, and evaluation criteria, covering both inference-time tricks like prompting, retrieval, and structured reasoning, and training-time adaptation of the models themselves. The authors also flag a newer category: agentic workflows where an LLM orchestrates verification tools rather than just generating one artifact.

The useful finding is also the cautionary one. LLMs work best as the reasoning and search layer sitting on top of simulators, solvers, and formal engines that provide the actual proof. But passing those tools' checks is not the same as verifying a design does what its specification says - an assertion or test can satisfy a checker while still missing the intended behavior. The survey names semantic alignment between specs and verification evidence, generalizing to unfamiliar designs, and rigorous cost-and-robustness evaluation as the open problems.

It is the hardware-verification version of a complaint already familiar from AI coding tools: a model that generates code which passes its own tests is not the same as a model that understands the task. In chip design, where a missed edge case ships in silicon rather than a patchable binary, that gap is a lot more expensive to find out about.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →