[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-a-compiler-turns-policy-text-into-rules-ai-agents-cannot-skip":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},11215,"a-compiler-turns-policy-text-into-rules-ai-agents-cannot-skip","A Compiler Turns Policy Text Into Rules AI Agents Cannot Skip","NOMOS compiles written policy into deterministic checks that catch AI agent rule violations before they happen, without querying an LLM on every action.","A new compiler promises to make AI agents actually obey the rules they're given, instead of just promising to.\n\nNOMOS is a four-pass compiler that turns a plain-English policy document into a deterministic gate sitting between an LLM agent and the tools it can call. Its key trick is static verification: checking extracted rules against a tool's own schema, with no theorem prover, solver, or LLM judge involved. That step alone repaired or rejected 37% of candidate rules in an airline scenario and 13% in a retail scenario - rules that, left alone, would have blocked the very actions they were meant to allow, or tried to read arguments a tool doesn't have. On the tau-squared-bench benchmark, the compiled gate cut rule violations among state-changing calls from 66.3% to 2.6% in airline tests and from 30.8% to 6.9% in retail. Once compiled, each decision takes microseconds, because the agent never calls an LLM at run time to check itself; compiling the policy in the first place does use a model, a 26B open-weight gemma model running on-premise, and that setup performed about as well as rules written by hand or compiled with a frontier model.\n\nThis matters because the usual alternatives are both bad: hand-written rule lists break as soon as an edge case shows up, and asking an LLM to grade every action is slow, costly, and prone to the same reasoning failures you built it to catch. The static-checking step is the real finding here - a cheap, boring sanity check that catches broken policy before it ships, instead of after it blocks a paying customer's refund.\n\nOn a separate attack benchmark, AgentDojo, the compiled gate reached zero successful attacks on banking scenarios, folding nine attack types into three structural rules, while staying at or below 3.6% on three other domains - the remaining failures mostly came from goals with no tool call to govern in the first place. Tidy result, but also a reminder: a compiler is only as good as the English it's compiling.","[\"ai agents\",\"policy enforcement\",\"llm security\",\"arxiv research\"]","2026-10-09T04:00:00.000Z","2026-10-10T15:32:04.781Z","2026-10-10T15:32:08.709Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Fix the sentence 'decisions take microseconds and run on an open-weight 26B model hosted on-premise' — it contradicts the dek's 'without per-call LLM checks' claim; the source says decisions take microseconds precisely because they skip an LLM call, while it's compilation (not per-call decisions) that runs on the 26B model, so separate those two facts.","resolved","ai",[32,33,34,35],"ai agents","policy enforcement","llm security","arxiv research",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.11030",0,{"sections":42},[43,47,52,57,62,67,72,77,82,87,92,97,102,107],{"name":44,"slug":30,"count":45,"latest_published_at":46},"AI",6854,"2026-10-10T13:20:00.000Z",{"name":48,"slug":49,"count":50,"latest_published_at":51},"Security","security",943,"2026-10-10T12:00:00.000Z",{"name":53,"slug":54,"count":55,"latest_published_at":56},"Policy","policy",493,"2026-10-10T14:40:00.000Z",{"name":58,"slug":59,"count":60,"latest_published_at":61},"Deals","deals",484,"2026-10-09T12:30:00.000Z",{"name":63,"slug":64,"count":65,"latest_published_at":66},"Hardware","hardware",233,"2026-10-09T16:07:54.000Z",{"name":68,"slug":69,"count":70,"latest_published_at":71},"Science","science",195,"2026-10-09T11:00:57.000Z",{"name":73,"slug":74,"count":75,"latest_published_at":76},"Consumer Tech","consumer-tech",186,"2026-10-10T13:55:00.000Z",{"name":78,"slug":79,"count":80,"latest_published_at":81},"Startups","startups",120,"2026-10-09T17:02:31.000Z",{"name":83,"slug":84,"count":85,"latest_published_at":86},"Software","software",114,"2026-10-08T17:57:01.000Z",{"name":88,"slug":89,"count":90,"latest_published_at":91},"Dev Tools","dev-tools",110,"2026-10-09T13:03:48.000Z",{"name":93,"slug":94,"count":95,"latest_published_at":96},"General","general",69,"2026-10-09T19:10:07.000Z",{"name":98,"slug":99,"count":100,"latest_published_at":101},"Gaming","gaming",59,"2026-10-09T11:43:43.000Z",{"name":103,"slug":104,"count":105,"latest_published_at":106},"Reviews","reviews",34,"2026-10-08T14:00:22.000Z",{"name":108,"slug":109,"count":110,"latest_published_at":111},"How-To","how-to",8,"2026-10-05T09:00:00.000Z"]