SPUR - Qualitative Reasoning: leaderboard

Metric: Accuracy (%) on SPUR qualitative reasoning multiple-choice questions about multi-panel biomedical experimental figures from PubMed Central papers (questions drafted by GPT-4o, kept only when GPT-4o failed them in at least six of ten text-only attempts, then expert-reviewed): the questions ask the model to combine visual evidence, domain knowledge and experimental design to interpret biological significance; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 20 models tracked.

Top models

#ModelScore
1Gemini 3 Pro (Preview)90.31
2Claude 3.7 Sonnet (Thinking)87.58
3Gemini 2.5 Pro (Preview 06-05)86.54
4GPT-5.186.52
5Llama 4 Maverick84.64
6O4 Mini (High)84.33
7GLM-4.5V80.94
8Seed-1.680.31
9InternVL3-78B75.24
10Qwen 3 VL 30B A3B (Thinking)75.16
11Grok 4.1 Fast73.44
12Qwen 2.5 VL 72B Instruct73.1
13Mistral Small 3.172.82
14Ministral-3-14B-Instruct-251272.5
15Ministral-3-8B-Instruct-251270.85

Interactive version: theaggregate.ai/benchmark?slug=spur-qualitative-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-07.