CoCoReviewBench - Grounding: leaderboard
Metric: Category-level grounding score (how explicitly a comment points to a part of the paper and the issue in it) on the 1-5 judge scale, printed as the difference from the human reference reviews (3.75, leave-one-out: one human review scored against the others), so a positive value is better than the human reviewers and the range is -2.75 to +1.25; AI reviews of 1,300 ICLR and NeurIPS papers sampled from CoCoReviewBench (100 per year, original score distribution kept), each split into atomic opinions and compared category by category with the filtered human opinions by a GPT-5-Mini judge at temperature 0, general models prompted with the AI Scientist review prompt at temperature 0.6; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 18 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.2 | 0.92 |
| 2 | GPT-5 Mini | 0.77 |
| 3 | Gemini 3 Flash | 0.69 |
| 4 | Gemini 3 Pro | 0.69 |
| 5 | QwQ-32B | 0.58 |
| 6 | Qwen 3 32B (Thinking) | 0.43 |
| 7 | Nemotron 3 Nano 30B A3B (Reasoning) | 0.14 |
| 8 | Qwen 2.5 7B | -0.3 |
| 9 | Qwen 3 8B (Thinking) | -0.32 |
| 10 | Qwen 3 8B (Non-reasoning) | -0.4 |
| 11 | Llama 3.1 8B | -0.45 |
| 12 | Llama 3.3 70B | -0.63 |
Interactive version: theaggregate.ai/benchmark?slug=cocoreviewbench-grounding · How It Works · Data refreshed daily, snapshot 2026-10-07.