Beyond Rating - Strength Recall: leaderboard

Metric: Strength Max-Recall (from 0-1, times 100): share of a single human reviewer's strength points that the generated review's strength points cover, taking the best-covered reviewer per paper, on 1,000 test papers from ICLR 2024-2026 and NeurIPS 2022-2025 with three to five high-confidence human reviews each, the model writing a full review from the parsed paper; review points are split into atomic claims and matched to the human reviews' claims by Qwen3-235B; higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 15 models tracked.

Top models

#ModelScore
1Qwen 3 30B A3B 2507 Instruct53
2Qwen 3 235B A22B 2507 Instruct53
3DeepSeek V3.251
4Claude Sonnet 4.548
5Qwen 3 8B46
6Gemini 3 Pro (Preview)42
7Llama 3.1 8B Instruct39
8Llama 3.1 70B Instruct39
9GPT-5.238

Interactive version: theaggregate.ai/benchmark?slug=beyond-rating-strength-recall · How It Works · Data refreshed daily, snapshot 2026-10-07.