GitHub launches ReviewBench for AI code review
Research preview measures missed issues and noisy findings.

GitHub announced ReviewBench in research preview on October 5. The benchmark tests AI code review agents on 219 pull requests from 187 public repositories across 19 programming languages.
Its reference findings combine human reviews, code analysis and multiple models. Claude Sonnet 5 judges whether findings are valid.
Two scoring methods matter. Grounded scores compare agents against known issues. Augmented scores also credit valid discoveries missing from that reference set. GitHub uses grounded recall for comparisons because augmented recall changes with each agentโs findings.
Senior engineers who had not built the dataset agreed with its labels 96.6% of the time, GitHub reports.
Teams should check how results vary by language, issue severity and the balance between finding more problems and avoiding noise. ByteForward has not independently tested the benchmark.
Illustrative archival photograph of two programmers sharing a desktop computer in September 2011. Pair Programming by Calqui. Photograph taken on 30 September 2011. Via Wikimedia Commons. Licensed under Creative Commons Attribution ShareAlike 3.0. Converted to WebP. This image rendition is shared under the same license.



