Researchers have built a benchmark that scores AI systems on whether they can predict a policy's ripple effects, not just its headline costs and benefits.
The benchmark, described in a new arXiv paper, covers 96 named public-policy cases across eight domains, split evenly between four action types: implement, modify, pilot, and block. Each case is tagged with source-linked variables tracking benefit, capture, gaming, burden shift, instability, uncertainty, irreversibility, distributional risk, and implementation capacity. A runner regenerates every method output and aggregate score straight from the case table, and the simulator being tested never sees the expert's actual policy choice. The team's side-effect simulator posted a mean policy-effect quality score of 0.945, ahead of a risk-register baseline at 0.838 and a causal-loop baseline at 0.879.
Most policy-evaluation tools estimate direct costs and benefits and assume the surrounding institutions stay put. This benchmark instead rewards models for catching second-order effects - regulatory capture, gaming of new rules, compliance theater - that show up after implementation, which is where a lot of real-world policy analysis quietly fails.
The gains show up in side-effect recall, not in picking the single best policy action, so treat this as a sharper diagnostic tool rather than a system that will tell you what to actually do.