OpenAI has trained models specifically to critique AI-written summaries, and the results suggest humans catch far more flaws when they have AI help.
Researchers built "critique-writing" models tasked with identifying errors in summaries produced by other AI systems. Human evaluators shown those critiques found flaws much more often than evaluators working without them. A notable pattern emerged: larger models improved at critique-writing faster than they improved at summary-writing, meaning scale disproportionately benefits the reviewer role over the producer role. OpenAI frames this as evidence that AI systems can meaningfully assist human supervision of other AI systems on difficult tasks.
This lands squarely in the scalable oversight problem - how do humans stay meaningfully in control of AI as those systems tackle tasks too complex or voluminous for anyone to evaluate directly? If AI critics can reliably surface the errors that humans would otherwise miss, that extends supervision without requiring humans to match AI capabilities themselves. It is a rare case where a lab is publishing work that openly acknowledges the limits of unaided human review.
The obvious caveat: a critique model is itself an AI system, which means its blind spots become your blind spots. If the model systematically misses a category of error, any human relying on it will too - and may have less reason to go looking.