Policy/ ai governance · inference time ai · compute policy · arxiv research

Researchers Map 20 Ways to Police AI After It's Trained

A new taxonomy scores 20 inference-time oversight mechanisms and finds none pass muster against a well-resourced, state-level deployer.

Most AI governance rules still stop at the training run, but a new paper argues the real leverage point is shifting downstream, to the moment a model actually answers a prompt.

Researchers built a feasibility taxonomy of 20 inference-time governance mechanisms spanning monitoring, verification, and enforcement, each scored on a four-point readiness scale using evidence from four vendors. They then tested the taxonomy against a two-dimensional adversary model covering three capability tiers and four adversary roles, and mapped the mechanisms to four governance scenarios: domestic regulation, multilateral coordination, industry self-regulation, and compute-marketplace oversight. Fifteen of the 20 mechanisms already have commercial technical substrates running in production, though how governance-ready and tamper-resistant they are varies a lot. A second reviewer's independent readiness ratings agreed closely with the authors' own, at a quadratic-weighted Cohen's kappa of 0.74.

Today's frontier-AI rules, from compute-reporting thresholds to export controls, are built around the training run as the unit of regulation. But capability increasingly shows up later in the pipeline, through inference-time scaling, agent scaffolding, and models compressed to run on ordinary hardware, all of which slip past a training-focused checkpoint. The paper's sharpest finding undercuts optimism about catching up: none of the 20 mechanisms rate as adequate against a high-capability, state-level deployer, and fine-tuning strips out the model-internal safeguards entirely, though controls sitting outside the model itself can survive that.

In other words, having a technical option on paper is not the same as having a regulator's tool that works against a well funded adversary, a gap policymakers leaning on inference-time fixes should not overlook.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →