CoCoReviewBench - Verifiability: leaderboard

Metric: Category-level verifiability score (whether a claim can be checked from its reasoning, common knowledge or references) on the 1-5 judge scale, printed as the difference from the human reference reviews (2.38, leave-one-out: one human review scored against the others), so a positive value is better than the human reviewers and the range is -1.38 to +2.62; AI reviews of 1,300 ICLR and NeurIPS papers sampled from CoCoReviewBench (100 per year, original score distribution kept), each split into atomic opinions and compared category by category with the filtered human opinions by a GPT-5-Mini judge at temperature 0, general models prompted with the AI Scientist review prompt at temperature 0.6; higher is better. Source: arxiv.org. Saturation forecast: Around August 2027. 18 models tracked.

Top models

#ModelScore
1GPT-5.20.78
2GPT-5 Mini0.53
3Gemini 3 Pro0.34
4Gemini 3 Flash0.14
5QwQ-32B0.02
6Qwen 3 32B (Thinking)-0.15
7Nemotron 3 Nano 30B A3B (Reasoning)-0.29
8Qwen 3 8B (Thinking)-0.45
9Qwen 3 8B (Non-reasoning)-0.58
10Qwen 2.5 7B-0.59
11Llama 3.1 8B-0.66
12Llama 3.3 70B-0.78

Interactive version: theaggregate.ai/benchmark?slug=cocoreviewbench-verifiability · How It Works · Data refreshed daily, snapshot 2026-10-07.