Open Code Review AACR-Bench - Recall: leaderboard
Metric: Recall (%). Source: open-codereview.ai. 14 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GLM-5.1 | 20.8 |
| 2 | DeepSeek V4 Pro | 16.13 |
| 3 | GLM-5.2 | 15.9 |
| 4 | GPT-5.5 | 15.5 |
Interactive version: theaggregate.ai/benchmark?slug=open-code-review-aacr-bench-recall · How It Works · Data refreshed daily, snapshot 2026-09-20.