SpecVQA - Chinese Descriptive (L0): leaderboard

Metric: Accuracy (0-1, times 100) on the Chinese descriptive (L0) questions; expert-revised question-answer pairs on 620 spectra (NMR, IR, XRD, Raman, MS, UV-Vis, XPS) from published papers; each answer judged correct or not by GPT-o4-mini against the reference, numeric answers within a 5 percent tolerance; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 20 models tracked.

Top models

#ModelScore
1Gemini 3 Flash (Preview)80.47
2Gemini 3 Pro (Preview)79.39
3Gemini 2.5 Pro77.58
4Gemini 2.5 Flash74.73
5O4 Mini (2025-04-16)72.37
6GPT-571.54
7GPT-5 (High)71.15
8GPT-5 (Low)70.66
9O3 (2025-04-16)70.27
10GPT-5.1 (Non-reasoning)68.99
11GPT-5.2 (Non-reasoning)68.99
12Claude Sonnet 4.5 (Thinking)63.15
13Claude Sonnet 4 (Thinking)59.57
14Qwen 3 VL 8B Instruct59.27
15Qwen 3 VL 8B (Thinking)58.05

Interactive version: theaggregate.ai/benchmark?slug=specvqa-chinese-descriptive-l0 · How It Works · Data refreshed daily, snapshot 2026-10-07.