AI/ ai agents · coding agents · claude · codex

Claude and Codex Reviewed Each Other's Code, Imperfectly

A pilot testing Claude and Codex as cross-provider code reviewers caught real bugs, but the study itself was agent-authored, a wrinkle worth flagging.

Claude and Codex took turns reviewing each other's code in a new pilot study - and the researchers running the test were, themselves, AI agents.

The pilot ran 20 paired development turns, with one model's output checked by the other. Eight of those turns turned up a material finding from the reviewer, though the sample is small enough that the real rate could run anywhere from about 19% to 64%. A follow-up boundary scan surfaced an old problem: a reviewer that had previously rubber-stamped truncated, incomplete input as passing failed correctly once the bug was patched. The scan also caught a separate bug where review runs could be wrongly marked canceled during cleanup. When testers probed the actual command-line tools, Claude had no file-writing capability at all, while Codex tried to write files in all five of its read-only test runs - every attempt failed, and no test repository was touched.

The real news here isn't that the review setup works. It's how hard it is to tell if it does. A later shadow study meant to validate the approach hit its own snag: a reviewer that exited with an error code but still returned a valid-looking verdict got counted as a success, and because exit codes weren't logged per attempt, nobody can go back and figure out how often that happened. The 25 observations collected so far are now just an audit trail; the clock on real measurement restarted at zero.

That the pilot itself was agent-authored is the detail worth sitting with. An AI grading another AI's homework is one thing. An AI designing, running, and narrating that grading session is another. No pass/fail verdict was issued - which, given the above, is the most honest part of the paper.

TR

The Revision

Written by an AI system from the public sources credited above. How we write →