SpecVQA: leaderboard

Metric: Accuracy (0-1, times 100), the unweighted mean of the English and Chinese descriptive and reasoning splits; expert-revised question-answer pairs on 620 spectra (NMR, IR, XRD, Raman, MS, UV-Vis, XPS) from published papers; each answer judged correct or not by GPT-o4-mini against the reference, numeric answers within a 5 percent tolerance; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 20 models tracked.

Top models

#ModelScore
1Gemini 3 Flash (Preview)78.72
2Gemini 3 Pro (Preview)77.87
3Gemini 2.5 Pro76.15
4Gemini 2.5 Flash72.74
5GPT-5 (High)70.59
6O4 Mini (2025-04-16)70.54
7GPT-570.39
8O3 (2025-04-16)70.31
9GPT-5 (Low)69.61
10GPT-5.1 (Non-reasoning)65.78
11GPT-5.2 (Non-reasoning)65.4
12Claude Sonnet 4.5 (Thinking)59.41
13Claude Sonnet 4 (Thinking)56.88
14Qwen 3 VL 8B (Thinking)56.31
15Doubao-Seed-1.6 (Thinking)56.02

Interactive version: theaggregate.ai/benchmark?slug=specvqa · How It Works · Data refreshed daily, snapshot 2026-10-07.