SPUR - Trend Analysis: leaderboard

Metric: Accuracy (%) on SPUR the 1,357 trend analysis multiple-choice questions about multi-panel biomedical experimental figures from PubMed Central papers (questions drafted by GPT-4o, kept only when GPT-4o failed them in at least six of ten text-only attempts, then expert-reviewed): the questions ask the model to interpret directional changes across panels of the same type; higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 20 models tracked.

Top models

#ModelScore
1Ministral-3-14B-Instruct-251257.79
2Ministral-3-8B-Instruct-251257.49
3GLM-4.5V55.71
4Qwen 3 VL 30B A3B (Thinking)54.85
5Gemini 2.5 Pro (Preview 06-05)53.3
6Llama 4 Maverick51.88
7Qwen 2.5 VL 72B Instruct51.87
8Mistral Small 3.151.82
9Claude 3.7 Sonnet (Thinking)51.3
10GPT-5.151.18
11Gemini 3 Pro (Preview)51.04
12Qwen 3 VL 30B A3B Instruct50.41
13InternVL3-78B49.52
14Gemma 3 27B (IT)48.59
15O4 Mini (High)48.37

Interactive version: theaggregate.ai/benchmark?slug=spur-trend-analysis · How It Works · Data refreshed daily, snapshot 2026-10-07.