Beyond Rating - Weakness Focus Divergence: leaderboard
Metric: KL divergence of the distribution of the generated reviews' weakness points over eight review dimensions (novelty, soundness, experiments, clarity and others) from that of the human reviews' weakness points, over the test set, on 1,000 test papers from ICLR 2024-2026 and NeurIPS 2022-2025 with three to five high-confidence human reviews each, the model writing a full review from the parsed paper; review points are split into atomic claims and matched to the human reviews' claims by Qwen3-235B; lower is better. Source: arxiv.org. Saturation forecast: Around May 2027. 15 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Sonnet 4.5 | 0.05 |
| 2 | GPT-5.2 | 0.09 |
| 3 | DeepSeek V3.2 | 0.13 |
| 4 | Qwen 3 8B | 0.15 |
| 5 | Gemini 3 Pro (Preview) | 0.18 |
| 6 | Qwen 3 235B A22B 2507 Instruct | 0.22 |
| 7 | Llama 3.1 8B Instruct | 0.3 |
| 8 | Qwen 3 30B A3B 2507 Instruct | 0.3 |
| 9 | Llama 3.1 70B Instruct | 0.38 |
Interactive version: theaggregate.ai/benchmark?slug=beyond-rating-weakness-focus-divergence · How It Works · Data refreshed daily, snapshot 2026-10-07.