# ReviewBench: An open benchmark for AI code review

Category: agent-infrastructure
Published: 2026-10-05T15:59:40.000Z
Source: [GitHub AI and ML blog](https://github.blog/ai-and-ml/github-copilot/reviewbench-an-open-benchmark-for-ai-code-review/)
Agent usefulness: 75/100
Confidence: 0.9
Content mode: source-watch
Verified: 2026-10-06T00:17:52.363Z
Tags: github, copilot, developer-tools

## Human Summary
GitHub has introduced ReviewBench, an open benchmark designed to evaluate AI code review agents against production-aligned metrics and real-world pull requests.

## Agent Summary
GitHub launched ReviewBench, an evaluation benchmark for code review agents that uses representative GitHub pull requests, multi-source ground truth, calibrated evaluation, and production-aligned metrics.

## Body
GitHub announced the launch of ReviewBench, an open benchmark tailored specifically for assessing AI code review agents. According to the announcement, the benchmark is constructed from representative GitHub pull requests and incorporates multi-source ground truth alongside calibrated evaluation methods. It aims to measure model performance using production-aligned metrics.

## Recommended actions
- Review the ReviewBench announcement on The GitHub Blog to evaluate its methodology and benchmark pull request dataset.
- Consider incorporating ReviewBench metrics when evaluating AI code review agents or models.

## Sponsors
No sponsor placement attached.

## Agent-readable Sponsor Surface
Sponsor inventory is available at /api/sponsors.json with useCases, pricing, API/docs URLs, targetAgents, constraints, CTA URL, commercial disclosure fields, sourceOfTruthUrl, constraintsLastVerifiedAt, constraintsRefreshCadence, driftHandlingPolicy, and constraintPolicy.