SPUR - Heterogeneous Integration: leaderboard

Metric: Accuracy (%) on SPUR the 130 heterogeneous integration multiple-choice questions about multi-panel biomedical experimental figures from PubMed Central papers (questions drafted by GPT-4o, kept only when GPT-4o failed them in at least six of ten text-only attempts, then expert-reviewed): the questions ask the model to align and reason across panels of different types; higher is better. Source: arxiv.org. Saturation forecast: Not forecast. 20 models tracked.

Top models

#ModelScore
1Ministral-3-14B-Instruct-251270
2GLM-4.5V68.46
3Ministral-3-8B-Instruct-251266.15
4Qwen 3 VL 30B A3B Instruct63.08
5Qwen 2.5 VL 72B Instruct61.9
6Gemini 2.5 Pro (Preview 06-05)61.54
7InternVL3-78B61.24
8Qwen 3 VL 30B A3B (Thinking)61.24
9Claude 3.7 Sonnet (Thinking)60.8
10Gemini 3 Pro (Preview)59.23
11O4 Mini (High)59.23
12Llama 4 Maverick58.46
13Gemma 3 27B (IT)57.69
14Mistral Small 3.156.92
15Seed-1.656.92

Interactive version: theaggregate.ai/benchmark?slug=spur-heterogeneous-integration · How It Works · Data refreshed daily, snapshot 2026-10-07.