AI/ ai · policy · benchmarks · research

New Benchmark Tests Whether AI Sees Policy Side Effects

A new 96-case benchmark finds AI policy simulators score higher when they model second-order effects like capture and gaming, not just direct costs.

Researchers have built a benchmark that scores AI systems on whether they can predict a policy's ripple effects, not just its headline costs and benefits.

The benchmark, described in a new arXiv paper, covers 96 named public-policy cases across eight domains, split evenly between four action types: implement, modify, pilot, and block. Each case is tagged with source-linked variables tracking benefit, capture, gaming, burden shift, instability, uncertainty, irreversibility, distributional risk, and implementation capacity. A runner regenerates every method output and aggregate score straight from the case table, and the simulator being tested never sees the expert's actual policy choice. The team's side-effect simulator posted a mean policy-effect quality score of 0.945, ahead of a risk-register baseline at 0.838 and a causal-loop baseline at 0.879.

Most policy-evaluation tools estimate direct costs and benefits and assume the surrounding institutions stay put. This benchmark instead rewards models for catching second-order effects - regulatory capture, gaming of new rules, compliance theater - that show up after implementation, which is where a lot of real-world policy analysis quietly fails.

The gains show up in side-effect recall, not in picking the single best policy action, so treat this as a sharper diagnostic tool rather than a system that will tell you what to actually do.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →