OpenAI has hit pause on training its next frontier model after internal reviews suggested it might be edging toward a serious cyber capability threshold.
The company says its unreleased Astra model may have reached what it calls "critical" cyber capabilities. In response, OpenAI halted a significant number of training runs and is now tightening its internal safeguards before continuing development. It hasn't detailed what those new safeguards involve, or exactly what pushed Astra over the line. Pausing training mid-development, rather than catching an issue after release, is the notable part here.
This matters because capability thresholds only mean something if a lab actually stops when a model approaches one. Frontier AI companies have increasingly adopted internal policies meant to flag risky capabilities before shipping, and this is a real-world test of whether that kind of self-policing holds up against the pressure to keep pace with rivals.
The word "critical" is doing a lot of work here, and OpenAI hasn't said what it actually means in practice.