SPUR - Quantitative Reasoning: leaderboard

Metric: Accuracy (%) on SPUR quantitative reasoning multiple-choice questions about multi-panel biomedical experimental figures from PubMed Central papers (questions drafted by GPT-4o, kept only when GPT-4o failed them in at least six of ten text-only attempts, then expert-reviewed): the questions ask the model to verify inter-group differences, effect sizes and significance by numerical comparison and calculation; higher is better. Source: arxiv.org. Saturation forecast: Around 2031. 20 models tracked.

Top models

#ModelScore
1Claude 3.7 Sonnet (Thinking)59.96
2Gemini 3 Pro (Preview)58.9
3GLM-4.5V58.48
4Gemini 2.5 Pro (Preview 06-05)57.94
5Llama 4 Maverick57.02
6Ministral-3-14B-Instruct-251256.81
7GPT-5.156.36
8O4 Mini (High)56.36
9Qwen 3 VL 30B A3B (Thinking)54.17
10Seed-1.653.89
11Ministral-3-8B-Instruct-251253.53
12Qwen 2.5 VL 72B Instruct52.51
13Qwen 3 VL 30B A3B Instruct51.41
14Gemma 3 27B (IT)51.38
15InternVL3-78B51.06

Interactive version: theaggregate.ai/benchmark?slug=spur-quantitative-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-07.