[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"branding":3,"analytics":7,"article-claude-and-codex-reviewed-each-others-code-imperfectly":10,"sections":41},{"siteName":4,"siteTagline":5,"publisherName":4,"contactEmail":6},"The Revision","Tech news, decoded.","editor@therevision.news",{"gaMeasurementId":8,"adsenseClientId":9},"G-ZW2MV82GYR","ca-pub-8533917693782264",{"article":11},{"id":12,"slug":13,"title":14,"dek":15,"body_md":16,"tags_json":17,"published_at":18,"created_at":19,"updated_at":20,"status":21,"review_note":22,"review_notes":23,"image_url":22,"persona_id":22,"persona_name":22,"section":30,"tags":31,"sources":36,"feedback":40,"feedback_at":22,"cost_usd":40,"total_tokens":40},10968,"claude-and-codex-reviewed-each-others-code-imperfectly","Claude and Codex Reviewed Each Other's Code, Imperfectly","A pilot testing Claude and Codex as cross-provider code reviewers caught real bugs, but the study itself was agent-authored, a wrinkle worth flagging.","Claude and Codex took turns reviewing each other's code in a new pilot study - and the researchers running the test were, themselves, AI agents.\n\nThe pilot ran 20 paired development turns, with one model's output checked by the other. Eight of those turns turned up a material finding from the reviewer, though the sample is small enough that the real rate could run anywhere from about 19% to 64%. A follow-up boundary scan surfaced an old problem: a reviewer that had previously rubber-stamped truncated, incomplete input as passing failed correctly once the bug was patched. The scan also caught a separate bug where review runs could be wrongly marked canceled during cleanup. When testers probed the actual command-line tools, Claude had no file-writing capability at all, while Codex tried to write files in all five of its read-only test runs - every attempt failed, and no test repository was touched.\n\nThe real news here isn't that the review setup works. It's how hard it is to tell if it does. A later shadow study meant to validate the approach hit its own snag: a reviewer that exited with an error code but still returned a valid-looking verdict got counted as a success, and because exit codes weren't logged per attempt, nobody can go back and figure out how often that happened. The 25 observations collected so far are now just an audit trail; the clock on real measurement restarted at zero.\n\nThat the pilot itself was agent-authored is the detail worth sitting with. An AI grading another AI's homework is one thing. An AI designing, running, and narrating that grading session is another. No pass\u002Ffail verdict was issued - which, given the above, is the most honest part of the paper.","[\"ai agents\",\"coding agents\",\"claude\",\"codex\"]","2026-10-09T04:00:00.000Z","2026-10-09T23:54:43.269Z","2026-10-09T23:54:48.983Z","published",null,[24],{"id":25,"reviewer":26,"round":27,"reason":28,"status":29},"editor-r1","editor",1,"Restore the concrete product names from the source (Claude, Codex) instead of generic 'one agent'\u002F'another agent,' and disclose that the pilot itself is described in the source as 'agent-authored' — that's a material caveat for reader skepticism that's currently missing.","resolved","ai",[32,33,34,35],"ai agents","coding agents","claude","codex",[37],{"name":38,"url":39},"arXiv cs.AI","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.10961",0,{"sections":42},[43,46,50,55,60,64,68,73,78,83,88,93,98,103],{"name":44,"slug":30,"count":45,"latest_published_at":18},"AI",6707,{"name":47,"slug":48,"count":49,"latest_published_at":18},"Security","security",931,{"name":51,"slug":52,"count":53,"latest_published_at":54},"Policy","policy",486,"2026-10-08T22:40:11.000Z",{"name":56,"slug":57,"count":58,"latest_published_at":59},"Deals","deals",474,"2026-10-08T22:00:00.000Z",{"name":61,"slug":62,"count":63,"latest_published_at":18},"Hardware","hardware",231,{"name":65,"slug":66,"count":67,"latest_published_at":18},"Science","science",192,{"name":69,"slug":70,"count":71,"latest_published_at":72},"Consumer Tech","consumer-tech",181,"2026-10-08T23:26:35.000Z",{"name":74,"slug":75,"count":76,"latest_published_at":77},"Startups","startups",117,"2026-10-08T16:45:00.000Z",{"name":79,"slug":80,"count":81,"latest_published_at":82},"Software","software",114,"2026-10-08T17:57:01.000Z",{"name":84,"slug":85,"count":86,"latest_published_at":87},"Dev Tools","dev-tools",105,"2026-10-07T16:59:11.000Z",{"name":89,"slug":90,"count":91,"latest_published_at":92},"General","general",66,"2026-10-09T04:46:11.000Z",{"name":94,"slug":95,"count":96,"latest_published_at":97},"Gaming","gaming",58,"2026-10-08T20:08:45.000Z",{"name":99,"slug":100,"count":101,"latest_published_at":102},"Reviews","reviews",34,"2026-10-08T14:00:22.000Z",{"name":104,"slug":105,"count":106,"latest_published_at":107},"How-To","how-to",8,"2026-10-05T09:00:00.000Z"]