CoCoReviewBench - Paper Level: leaderboard
Metric: Paper-level review score (the same judge prompt applied once to all categories of the review together) on the 1-5 judge scale, printed as the difference from the human reference reviews (3.61, leave-one-out: one human review scored against the others), so a positive value is better than the human reviewers and the range is -2.61 to +1.39; AI reviews of 1,300 ICLR and NeurIPS papers sampled from CoCoReviewBench (100 per year, original score distribution kept), each split into atomic opinions and compared category by category with the filtered human opinions by a GPT-5-Mini judge at temperature 0, general models prompted with the AI Scientist review prompt at temperature 0.6; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 18 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.2 | 0.9 |
| 2 | GPT-5 Mini | 0.81 |
| 3 | Gemini 3 Pro | 0.5 |
| 4 | Gemini 3 Flash | 0.44 |
| 5 | QwQ-32B | 0.37 |
| 6 | Qwen 3 32B (Thinking) | 0.21 |
| 7 | Nemotron 3 Nano 30B A3B (Reasoning) | 0.19 |
| 8 | Qwen 2.5 7B | -0.4 |
| 9 | Llama 3.3 70B | -0.53 |
| 10 | Qwen 3 8B (Thinking) | -0.56 |
| 11 | Qwen 3 8B (Non-reasoning) | -0.57 |
| 12 | Llama 3.1 8B | -0.68 |
Interactive version: theaggregate.ai/benchmark?slug=cocoreviewbench-paper-level · How It Works · Data refreshed daily, snapshot 2026-10-07.