CoCoReviewBench - Category Completeness: leaderboard

Metric: Category completeness (%): the average number of review categories the AI review covers per paper as a share of the union of categories all human reviewers of the paper cover; AI reviews of 1,300 ICLR and NeurIPS papers sampled from CoCoReviewBench (100 per year, original score distribution kept), each split into atomic opinions and compared category by category with the filtered human opinions by a GPT-5-Mini judge at temperature 0, general models prompted with the AI Scientist review prompt at temperature 0.6; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 18 models tracked.

Top models

#ModelScore
1GPT-5 Mini89.97
2GPT-5.284.49
3Nemotron 3 Nano 30B A3B (Reasoning)82.93
4QwQ-32B79.83
5Qwen 3 8B (Thinking)75.73
6Qwen 3 32B (Thinking)73.17
7Qwen 2.5 7B72.4
8Qwen 3 8B (Non-reasoning)72.28
9Gemini 3 Flash68.67
10Gemini 3 Pro67.69
11Llama 3.3 70B65.49
12Llama 3.1 8B62.07

Interactive version: theaggregate.ai/benchmark?slug=cocoreviewbench-category-completeness · How It Works · Data refreshed daily, snapshot 2026-10-07.