Beyond Rating - Weakness Focus Divergence: leaderboard

Metric: KL divergence of the distribution of the generated reviews' weakness points over eight review dimensions (novelty, soundness, experiments, clarity and others) from that of the human reviews' weakness points, over the test set, on 1,000 test papers from ICLR 2024-2026 and NeurIPS 2022-2025 with three to five high-confidence human reviews each, the model writing a full review from the parsed paper; review points are split into atomic claims and matched to the human reviews' claims by Qwen3-235B; lower is better. Source: arxiv.org. Saturation forecast: Around May 2027. 15 models tracked.

Top models

#ModelScore
1Claude Sonnet 4.50.05
2GPT-5.20.09
3DeepSeek V3.20.13
4Qwen 3 8B0.15
5Gemini 3 Pro (Preview)0.18
6Qwen 3 235B A22B 2507 Instruct0.22
7Llama 3.1 8B Instruct0.3
8Qwen 3 30B A3B 2507 Instruct0.3
9Llama 3.1 70B Instruct0.38

Interactive version: theaggregate.ai/benchmark?slug=beyond-rating-weakness-focus-divergence · How It Works · Data refreshed daily, snapshot 2026-10-07.