CoCoReviewBench - Thoroughness: leaderboard
Metric: Category-level thoroughness score (how completely the comments cover the human opinions of the category) on the 1-5 judge scale, printed as the difference from the human reference reviews (2.37, leave-one-out: one human review scored against the others), so a positive value is better than the human reviewers and the range is -1.37 to +2.63; AI reviews of 1,300 ICLR and NeurIPS papers sampled from CoCoReviewBench (100 per year, original score distribution kept), each split into atomic opinions and compared category by category with the filtered human opinions by a GPT-5-Mini judge at temperature 0, general models prompted with the AI Scientist review prompt at temperature 0.6; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 18 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.2 | 0.64 |
| 2 | GPT-5 Mini | 0.58 |
| 3 | Gemini 3 Pro | 0.16 |
| 4 | Gemini 3 Flash | 0.13 |
| 5 | QwQ-32B | 0.13 |
| 6 | Qwen 3 32B (Thinking) | 0.12 |
| 7 | Nemotron 3 Nano 30B A3B (Reasoning) | 0.12 |
| 8 | Qwen 2.5 7B | -0.06 |
| 9 | Qwen 3 8B (Non-reasoning) | -0.1 |
| 10 | Qwen 3 8B (Thinking) | -0.13 |
| 11 | Llama 3.1 8B | -0.22 |
| 12 | Llama 3.3 70B | -0.23 |
Interactive version: theaggregate.ai/benchmark?slug=cocoreviewbench-thoroughness · How It Works · Data refreshed daily, snapshot 2026-10-07.