Accessibility Adjustments

Use these optional tools to adjust reading and display preferences. These tools cannot resolve every accessibility barrier. Please contact the website owner if you need assistance.

  • Text adjustments
  • Content scaling 100%
  • Font size 100%
  • Line height 100%
  • Letter spacing 100%
  • Colour adjustments
  • Orientation adjustments

GitHub launches ReviewBench for AI code review

Research preview measures missed issues and noisy findings.

Listen to this article

GitHub announced ReviewBench in research preview on October 5. The benchmark tests AI code review agents on 219 pull requests from 187 public repositories across 19 programming languages.

Its reference findings combine human reviews, code analysis and multiple models. Claude Sonnet 5 judges whether findings are valid.

Two scoring methods matter. Grounded scores compare agents against known issues. Augmented scores also credit valid discoveries missing from that reference set. GitHub uses grounded recall for comparisons because augmented recall changes with each agentโ€™s findings.

Senior engineers who had not built the dataset agreed with its labels 96.6% of the time, GitHub reports.

Teams should check how results vary by language, issue severity and the balance between finding more problems and avoiding noise. ByteForward has not independently tested the benchmark.

Illustrative archival photograph of two programmers sharing a desktop computer in September 2011. Pair Programming by Calqui. Photograph taken on 30 September 2011. Via Wikimedia Commons. Licensed under Creative Commons Attribution ShareAlike 3.0. Converted to WebP. This image rendition is shared under the same license.

Maya Chen
Maya Chen

Maya Chen is focused on covering AI models, research, and the evidence behind new capabilities. Maya follows model launches, benchmarks, open weights, and scientific uses of AI with one question in mind. What changed, and how would we know? The voice is curious and exacting, with a soft spot for elegant technical ideas and little patience for a leaderboard without context.

Leave a Reply

Your email address will not be published. Required fields are marked *

Gravatar profile