Researchers just published a checklist for catching an AI system before it goes rogue.
A paper posted to arXiv proposes a framework of behavioral indicators meant to signal when an AI system may be progressing toward catastrophic capabilities. The authors borrow methodology from cybersecurity and national-security threat monitoring, building out metrics, indicators, and thresholds across multiple dimensions of AI capability and behavior. The stated goal is to give researchers and policymakers a shared, evidence-based way to track warning signs instead of relying on ad hoc judgment calls. The paper positions itself as a monitoring protocol, not a forecast of when or whether such risks will actually show up.
That framing matters because "AI could be dangerous" has mostly stayed a vague, unfalsifiable worry. Borrowing intrusion-detection logic - define the indicators, set the thresholds, watch for them - is an attempt to make the risk trackable rather than just debatable. If labs and regulators actually used a shared framework like this, it would give them a common bar for deciding when a system's behavior warrants intervention, instead of each lab quietly setting its own.
The catch is the same one that dogs every self-reported metric: a framework only works if the organizations being monitored actually publish the numbers, and AI labs have not exactly built a track record of consistent disclosure.