A new compiler promises to make AI agents actually obey the rules they're given, instead of just promising to.
NOMOS is a four-pass compiler that turns a plain-English policy document into a deterministic gate sitting between an LLM agent and the tools it can call. Its key trick is static verification: checking extracted rules against a tool's own schema, with no theorem prover, solver, or LLM judge involved. That step alone repaired or rejected 37% of candidate rules in an airline scenario and 13% in a retail scenario - rules that, left alone, would have blocked the very actions they were meant to allow, or tried to read arguments a tool doesn't have. On the tau-squared-bench benchmark, the compiled gate cut rule violations among state-changing calls from 66.3% to 2.6% in airline tests and from 30.8% to 6.9% in retail. Once compiled, each decision takes microseconds, because the agent never calls an LLM at run time to check itself; compiling the policy in the first place does use a model, a 26B open-weight gemma model running on-premise, and that setup performed about as well as rules written by hand or compiled with a frontier model.
This matters because the usual alternatives are both bad: hand-written rule lists break as soon as an edge case shows up, and asking an LLM to grade every action is slow, costly, and prone to the same reasoning failures you built it to catch. The static-checking step is the real finding here - a cheap, boring sanity check that catches broken policy before it ships, instead of after it blocks a paying customer's refund.
On a separate attack benchmark, AgentDojo, the compiled gate reached zero successful attacks on banking scenarios, folding nine attack types into three structural rules, while staying at or below 3.6% on three other domains - the remaining failures mostly came from goals with no tool call to govern in the first place. Tidy result, but also a reminder: a compiler is only as good as the English it's compiling.