Two of AI's biggest commercial rivals have published findings from a mutual audit of each other's models.
OpenAI and Anthropic ran a cross-lab safety evaluation testing each other's systems across a range of risk categories: misalignment, instruction following, hallucinations, and jailbreaking resistance. The two companies shared the results, calling it a first of its kind. The collaboration carries a structural advantage internal red-teaming can't replicate — each lab has genuine competitive incentive to find real problems in the other's work. What the tests actually revealed, and whether the full methodology is open for outside scrutiny, remains unclear from what's been published.
The deeper significance is structural. AI labs have long evaluated their own models, setting the standards, running the tests, and interpreting the results — a setup regulators and critics have repeatedly flagged as inadequate. A peer with competitive skin in the game is a harder skeptic than an in-house safety team, and that matters as governments push for more credible AI accountability.
Still, OpenAI and Anthropic share a mutual interest in framing AI risk as something responsible companies can self-manage. Peer review between rivals beats self-review, but it isn't independent oversight — and the distinction is worth keeping.