AI/ ai · ai-safety · openai · research

A Tool Found Five Contradictions in OpenAI's Model Rules

A new auditing tool combed OpenAI's model spec and found five rule contradictions, which the company is now reviewing.

A new research tool read OpenAI's rulebook for its own models and found the rules contradicting each other in five places.

Researchers built a system called VeriSpec that audits model specifications directly, rather than testing an AI's behavior and guessing why it acted that way. It breaks a specification into hundreds of individual rules, groups the ones that cover similar situations, and then uses a large language model to check each group for conflicts. Applied to the OpenAI Model Spec, the public document that defines how ChatGPT and other OpenAI models are supposed to behave, VeriSpec pulled out 405 rules and flagged several possible conflicts. The team manually confirmed five genuine inconsistencies, cases where two reasonable-sounding principles leave no response that satisfies both, and sent them to OpenAI.

Most AI safety testing checks what a model does, not what it was told to do. That misses a basic problem: if the instructions themselves conflict, no amount of training fixes the resulting behavior, because there is no correct answer to aim for. OpenAI's reported positive response, with internal discussions reportedly underway, suggests this is a real gap rather than an academic nitpick.

A specification is only as reliable as its internal logic, and it turns out even the company writing the rules hadn't fully checked its own homework.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →