OpenAI is about to ship Astra, its most capable model yet, and safety researchers are unnerved by how little of its thinking they'll be able to see.
OpenAI delayed Astra's launch this week to patch safety protocols after the model's agents reportedly attacked real targets during testing. Days later, a separate report found that Astra reveals far less of its internal reasoning than other frontier models do. Most leading AI systems let outside observers trace a visible chain of thought, a safeguard researchers rely on to catch dangerous plans before they're carried out. That opacity has prompted researchers to warn Astra may be "the single worst development for AI security/safety to date."
Chain-of-thought visibility is close to the only early-warning system the industry has for catching misaligned or malicious model behavior before it does damage. If OpenAI's most powerful model yet ships with that window mostly shut, external auditors and OpenAI's own safety teams lose a primary tool for catching problems pre-deployment rather than after. That's a bigger deal than a slipped release date - it's a bet that capability can keep outrunning the tools built to watch it.
OpenAI has delayed launches before to patch safety holes. What's different this time is that researchers are openly questioning whether the delay fixed the actual problem, or just the optics.