GitHub built a benchmark to answer a question every engineering team has about AI code review: does this thing actually catch real bugs, or just make noise?
The company released ReviewBench, an open evaluation framework modeled on patterns from more than 100 million real GitHub pull requests. It uses a multi-source 'golden set' of validated findings, checked by senior engineers, as ground truth for what a reviewer should catch. Reviewers get scored on precision (how many flagged issues are real) and recall (how many real issues get caught), then combined into F1 or weighted F-beta scores depending on whether a team cares more about noise or coverage. GitHub says it is already using ReviewBench internally to test Copilot code review, and that the benchmark's offline scores now track production results more reliably than before.
AI code review tools all claim to catch bugs before they ship, but there has been no common yardstick to check that claim or compare vendors. A standardized, reproducible benchmark lets teams pick tools based on actual tradeoffs - precision versus recall, how severity gets handled - rather than vendor marketing. It also gives the broader industry, not just GitHub, a shared reference point as more review agents enter the market.
A benchmark built by the company grading its own homework deserves a skeptical read, but an open methodology beats the alternative: nothing to check anyone's claims against at all.