AI/ ai · legal-tech · tax-law · neuro-symbolic

New AI System Turns Tax Law Into Verified Code

Researchers built a neuro-symbolic pipeline that hit 100% accuracy on held-out tax returns, far outpacing a lone large language model's 66%.

A new AI framework can turn the labyrinth of tax law into code that actually checks out.

Researchers built Law&Order, a system that feeds tax forms and filing instructions through a large language model, then verifies every resulting calculation cell against the legal text it's supposed to implement. When a cell fails that check, the system automatically repairs it using real filings from OpenTaxSolver, an open-source tax-prep tool, and loops until the logic matches. The team then tested the finished programs on 51 held-out returns from TaxCalcBench, a benchmark the system never saw during training or repair. The best standalone large language model managed just 66% accuracy on the same task; Law&Order hit 100%, at both the cell level and the full-form level.

That gap matters because tax law is a brutal test case for AI: arithmetic, branching logic, recursion, and tabular lookups, nested inside intentionally precise legal language. An LLM alone treats this like prose and guesses. Pairing it with symbolic verification, checking that generated code actually computes what the statute says, turns guesswork into something closer to a proof. That is a meaningfully different claim than most document-understanding hype, and one with obvious uses beyond the IRS: benefits eligibility, immigration forms, and other rules-as-code systems government agencies have struggled to automate reliably.

Still, 51 returns is a small test set, and this one was scoped tightly around a single, well-structured domain. The real test for autoformalization is whether the same recipe holds up on messier, less tabular statutes.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →