SpecVQA - English Descriptive (L0): leaderboard

Metric: Accuracy (0-1, times 100) on the English descriptive (L0) questions; expert-revised question-answer pairs on 620 spectra (NMR, IR, XRD, Raman, MS, UV-Vis, XPS) from published papers; each answer judged correct or not by GPT-o4-mini against the reference, numeric answers within a 5 percent tolerance; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 20 models tracked.

Top models

#ModelScore
1Gemini 3 Flash (Preview)77.82
2Gemini 3 Pro (Preview)77.18
3Gemini 2.5 Pro76.45
4Gemini 2.5 Flash73.95
5O4 Mini (2025-04-16)71.44
6GPT-5 (High)70.17
7O3 (2025-04-16)69.53
8GPT-569.23
9GPT-5 (Low)68.25
10GPT-5.1 (Non-reasoning)67.76
11GPT-5.2 (Non-reasoning)67.76
12Claude Sonnet 4.5 (Thinking)61.48
13Claude Sonnet 4 (Thinking)59.47
14Qwen 3 VL 8B (Thinking)58.64
15Qwen 3 VL 8B Instruct57.21

Interactive version: theaggregate.ai/benchmark?slug=specvqa-english-descriptive-l0 · How It Works · Data refreshed daily, snapshot 2026-10-07.