AI/ openai · ai-safety · reasoning-models · interpretability

OpenAI's New Reasoning Trick Worries AI Safety Researchers

OpenAI's Astra model ditches step-by-step reasoning for recurrent depth, and safety researchers say the resulting opacity is the real risk.

OpenAI's upcoming Astra model swaps sequential chain-of-thought reasoning for something called "recurrent depth" - and that shift is unsettling the people whose job is to keep AI systems honest.

OpenAI says Astra will use recurrent depth, a technique that lets the model reason outside the sequential, step-by-step structure that has defined most reasoning models to date. That step-by-step structure - write out each logical move, then answer - has doubled as a de facto safety feature, since outside researchers can read a model's chain of thought and catch it planning something it shouldn't. Recurrent depth reportedly breaks from that pattern, letting the model work through problems in ways that don't map cleanly onto a legible, line-by-line trace. Beyond that description, OpenAI has said little publicly about how the technique works or when Astra ships.

Chain-of-thought transparency has been one of the few tools available for catching a model reasoning toward a bad outcome, even when it can't fully explain itself. Trade that transparency for a more capable but opaque reasoning process, and outside auditors lose a window into what the model is actually doing. That's a bigger bet on trust than the industry has asked for before.

A more capable model that thinks in ways nobody outside OpenAI can fully follow isn't obviously progress. It's just a faster car with the dashboard removed.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →