A new paper tests a common assumption in AI tooling: that having a different model check your model's work catches more mistakes.
The author ran 900 review sessions across 30 artifacts seeded with 150 planted errors, using three reviewer models from two different developers and ten review setups. A top-tier model reviewing another model's output scored no significantly better F1 than that same top model reviewing its own output in a fresh session. The two approaches did catch different errors, overlapping only 41.2% of the time. Pairing one same-model review with one cross-model review flagged more planted errors than two same-model reviews (56.7% versus 42.7%), but it did not clearly beat two reviews from the top-tier model alone, so the study cannot say whether the gain comes from using a second model or just from using a stronger one. A cheaper cross-model reviewer performed no better than same-model review.
For engineering teams, this cuts against the pipeline design that treats adding a second vendor's model as a reliability upgrade. The real lever here appears to be reviewer capability and withheld context, not model diversity for its own sake. That is a cheaper, less vendor-locked way to improve code review than paying for two API bills.
Treat the numbers carefully. This is one author's self-run experiment with 30 artifacts, a handful of reviewer models, and 14 failed calls and one baseline run excluded after an audit. The data, scripts, and artifacts are not public; they are available only on request, so nobody else has checked this yet.