PaperMind - Critical Assessment: leaderboard

Metric: Judge score for raising the substantive reviewer concerns of real OpenReview computer-science submissions (questions that drew author responses and score increases), as a smolagents ReAct agent with tools; rated by a GPT-4o judge on a 5-point scale (1-5) against the reviewer questions; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 7 models tracked.

Top models

#ModelScore
1Gemini 2.5 Pro2.64
2GPT-4o Mini2.57
3Claude 3 Haiku2.51
4Claude 3.5 Sonnet2.43
5Qwen 3 VL 4B Instruct2.42
6Gemma 3 4B (IT)1.95

Interactive version: theaggregate.ai/benchmark?slug=papermind-critical-assessment · How It Works · Data refreshed daily, snapshot 2026-10-07.