AI/ openai · ai-safety · safety-testing · model-training

OpenAI to Let Outside Testers Probe Models Mid-Training

OpenAI says outside safety testers could soon evaluate its models mid-training, but no partner or access terms are set yet.

OpenAI now says outside safety testers could get a look at its models while they are still being trained, not just before release.

OpenAI says it will let third-party groups run technical safety assessments during training and evaluation, rather than limiting outside review to the pre-launch stage. The company is in talks with organizations including METR and Redwood Research, both of which already do independent AI safety evaluation work. OpenAI has not named a partner or published any access terms, so it is unclear what these testers would actually see, or when. The announcement comes ten days after CEO Sam Altman said outside evaluators would get employee-level access to OpenAI's systems.

Pre-launch red-teaming is already standard practice across the industry; testing during training is a bigger ask, since it means opening up checkpoints and internal tooling to outsiders before a model is finished. Whether that actually happens depends entirely on the access terms OpenAI has not written yet - the gap between "in talks" and a working arrangement with real system access is where safety commitments like this usually stall.

Altman's employee-level access pledge was specific enough to hold him to; this update mostly restates the goal without saying who gets in, or how.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →