GitHub Launches ReviewBench to Score AI Code Review Agents
GitHub released ReviewBench on 5 October 2026, an open benchmark that scores AI code review agents on 219 public pull requests against a human-validated golden set.
PromptCrates Editorial
Staff Writer

GitHub on Monday, 5 October 2026, released ReviewBench, an open benchmark for AI code review agents built from 219 public pull requests across 19 programming languages, together with a public leaderboard and a self-serve flow for vendors to submit their own runs. The benchmark is a research preview, and GitHub says its corpus was modeled on an analysis of 103.9 million real pull requests on the platform.
The GitHub blog post by Michelle Zhou and Alejandro Carderera de Diego argues that existing code review benchmarks force tradeoffs between label quality, coverage and realism. GitHub says ReviewBench is meant to address that gap, and that it has also made the company's offline testing of Copilot code review better at anticipating production results.
How GitHub built the ReviewBench dataset
The 219 pull requests come from 187 public open source repositories, with language and repository-size distributions that GitHub says closely match the platform overall. The team made one deliberate adjustment: pull request size is weighted toward the reviewable middle and tail, which reduces the share of tiny single-file changes and keeps more multi-file changes where review quality matters.
Ground truth comes from several sources. GitHub says it gathered candidate findings from human reviewers, issues inferred from authors' follow-up commits, static analysis tools and multiple frontier LLMs, then semantically merged duplicates so that agreement between sources would not inflate the golden set. A finding counts as a true positive only if it is true, relevant and non-trivial under a shared rubric, and every finding is labeled for severity (critical, medium or low) and for a category such as correctness, security or testing.
The grader is itself a model. GitHub says it uses Claude Sonnet 5 as the LLM judge and publishes both the rubric and the judge it uses. Before release, senior engineers who had not helped build the dataset independently re-labeled every ground-truth finding, and their true or false positive calls matched ReviewBench 96.6% of the time.
Grounded and augmented scores for code review
ReviewBench reports two families of metrics. Grounded precision, recall and F1 use only the existing golden-set labels, which gives a like-for-like comparison. Augmented metrics also judge findings that do not match the golden set, so an agent can earn credit for valid issues that no source had surfaced. Because each agent's new discoveries change its own denominator, GitHub uses grounded recall as the headline cross-system comparison and treats augmented scores as a per-system diagnostic.
Users can also tune the ranking. The ReviewBench site ranks agents by grounded F1 by default, while its Precise and Thorough filters switch to F0.5 or F2 scores, which weight precision or recall more heavily. Results can also be sliced by severity and category.
What the first ReviewBench leaderboard shows
The opening leaderboard puts GitHub's own product on top. In the site's leaderboard data, Copilot Code Review in its Balanced configuration ranks first, with grounded precision of about 88% and grounded recall of about 26%. Devin AI ranks second and Qodo third, followed by OpenAI's Codex running GPT-5.6 Sol, then Copilot Code Review in its Lite configuration and Codex running GPT-5.6 Luna, then Cubic and Greptile, with Cursor's reviewer further down the table.
Recall trails precision for every entry. Grounded precision across the listed configurations falls between roughly 84% and 90%, but none of them finds more than about a quarter of the known issues, according to the same data. Put another way, when these agents flag a problem it is usually real, but most known issues go unflagged. The site also reports how long each agent takes to review one pull request, averaged across three runs for most entries (Cubic and Greptile were run once each), though duration does not affect any score.
The site spells out how those entries were produced. Its notes say the initial results were run by the ReviewBench team using each vendor's publicly available code review product, that vendors did not run or verify them, and that the products may have changed since testing. Run dates vary by entry, and some go back to June. The site also states that ReviewBench is developed by GitHub, which makes Copilot code review, and that it is not sponsored or endorsed by the other vendors listed.
Why the benchmark's production signal matters
GitHub's main claim is that ReviewBench predicts what happens in production. In a recent experiment with a multi-model ensemble review for Copilot's lite tier, the benchmark predicted higher precision, recall and comment volume at lower cost. The later A/B test moved the same way, GitHub says: addressed rate, its online stand-in for precision, rose 8.0%, recall rose 13.6%, comment volume rose 61% and cost per review fell 8.0%. ReviewBench predicted a 227% rise in critical comments, against 262% measured online.
Vendors can now test themselves. Teams sign in with GitHub, register an agent with a container image, a configuration and their own model key, try a 25-pull-request test set, then run the full 219-pull-request set three times, scored by the same judge. Scores stay private until a maintainer approves them, and they appear on the leaderboard only if they beat the agent's current leaderboard score or are its first entry.
PromptCrates has previously covered questions about how AI benchmark scores are produced, including how OpenAI's Astra AGI score came from a harness and how DeepMind ran a double-blind frontier model evaluation. ReviewBench publishes its dataset, rubric, judge prompt and runner, so outside teams can rerun GitHub's ranking. For teams wiring review agents into workflows such as Slack Code's coding-agent channels, the site cautions that results reflect this benchmark only and may not predict performance on their own code.


