GitHub’s ReviewBench puts AI code reviewers to the test
GitHub’s ReviewBench measures how well AI code review tools detect problems before software is released. Available in research preview through its website, the benchmark lets users compare review agents and evaluate their own tools.
Code review involves checking proposed changes for mistakes. ReviewBench measures AI reviewers’ ability to identify issues and avoid false alarms. Users can inspect the test data, reproduce results and track improvements. Its leaderboard shows performance by issue severity, category and scoring preference.
How ReviewBench evaluates AI reviewers
ReviewBench tests agents on the same 219 pull requests from 187 public repositories with open-source licenses across 19 programming languages. It measures how many known issues each agent finds, how many findings are valid and how performance varies by severity and category, including security issues.
“We make one deliberate adjustment to this distribution: while language and repository size mirror GitHub directly, pull request size is weighted toward the reviewable middle and tail. This reduces the overrepresentation of tiny, single-file changes while preserving more substantive, multi-file pull requests where review quality matters most,” Michelle Zhou, Data Scientist at Microsoft, and Alejandro Carderera de Diego, Senior Applied Researcher at GitHub, wrote.
GitHub built a reference collection of validated findings, called the golden set, in three stages. It gathered candidate findings from human reviewers, follow-up code changes, analysis tools and AI models, merged findings describing the same issue, and assessed them using a shared evaluation rubric.
The benchmark reports six metrics in two groups. The grounded metrics measure the share of an agent’s findings that match known issues (precision) and the proportion of known issues it detects (recall). The F1 score gives both equal weight.
Augmented precision, recall and F1 also account for findings outside the golden set. An AI judge assesses whether these additional findings are valid, giving reviewers credit for issues missing from the reference collection. GitHub uses grounded recall as the main measure for comparing systems and augmented metrics to assess each system individually.
Users can filter results by severity and category and adjust β in the Fβ score to give more weight to issue detection or fewer false alarms. The leaderboard updates its rankings based on those preferences.
GitHub’s results with Copilot code review
GitHub uses ReviewBench to evaluate changes to Copilot code review before testing them with users. In one experiment with Copilot code review’s lite tier, GitHub reported improvements in benchmark scores and production results after combining several independent model runs into a single review.
In the production test, the share of review comments judged by an AI model to have prompted code changes rose by 8%, recall rose by 13.6% and cost per review fell by 8%, relative to the production control. GitHub measures production recall through how much additional human review is still needed. Feedback also moved toward critical and moderate issues, with fewer minor suggestions.

Lite-tier ensemble review compared with the production single-reviewer control (Source: GitHub)
The company says the benchmark helps identify updates for production testing. Tests with users remain its final measure of impact.
How to submit an AI reviewer
Users can sign in to ReviewBench with GitHub and register an agent by providing its container image, configuration and model access key. A 25-pull-request test set lets them assess and refine performance before running three rounds on the full set of 219 pull requests.
ReviewBench scores every agent with the same AI judge. Scores remain private until a maintainer approves the submission. Scores are published for an agent’s first leaderboard entry or when it exceeds its previous leaderboard score.
Michelle Zhou and Alejandro Carderera de Diego invited researchers and practitioners to evaluate their systems, examine the benchmark’s assumptions and help improve its methodology.