OpenAI is tearing up its own rulebook because its next model might be good enough to help someone break into systems.
The company said Tuesday it is rewriting its Preparedness Framework, the internal document that decides how much scrutiny a model gets before release. The trigger is Astra, OpenAI's upcoming model, which the company now believes may have crossed the threshold for meaningful cyber capability - the point where a model could plausibly assist in real-world hacking. To catch that kind of risk, OpenAI is rolling out token-level monitoring that watches what a model is doing as it generates output. That monitoring isn't free: it adds roughly 20% compute overhead, and it's now mandatory for the company's most capable training runs.
This is OpenAI admitting, in writing, that its existing safety testing wasn't built for a model like Astra. Other labs have talked about cyber-capable models as a looming risk category, but a 20% compute tax is a real cost, not a talking point - it suggests OpenAI is treating this as an engineering problem, not a PR one.
Rulebooks tend to get rewritten only after they've already been tested. The open question is whether Astra needed this level of scrutiny before anyone decided to build it in the first place.