Beyond Rating - Weakness F1: leaderboard
Metric: F1 (from 0-1, times 100) combining point-wise precision (share of the generated review's weakness points matched by any human reviewer's weakness point) with weakness Max-Recall, on 1,000 test papers from ICLR 2024-2026 and NeurIPS 2022-2025 with three to five high-confidence human reviews each, the model writing a full review from the parsed paper; review points are split into atomic claims and matched to the human reviews' claims by Qwen3-235B; higher is better. Source: arxiv.org. Saturation forecast: Around 2033. 15 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | DeepSeek V3.2 | 31 |
| 2 | Claude Sonnet 4.5 | 29 |
| 3 | Qwen 3 235B A22B 2507 Instruct | 28 |
| 4 | Gemini 3 Pro (Preview) | 26 |
| 5 | Qwen 3 30B A3B 2507 Instruct | 26 |
| 6 | GPT-5.2 | 25 |
| 7 | Qwen 3 8B | 24 |
| 8 | Llama 3.1 70B Instruct | 18 |
| 9 | Llama 3.1 8B Instruct | 17 |
Interactive version: theaggregate.ai/benchmark?slug=beyond-rating-weakness-f1 · How It Works · Data refreshed daily, snapshot 2026-10-07.