SpecVQA - Chinese Reasoning (L1): leaderboard

Metric: Accuracy (0-1, times 100) on the Chinese reasoning (L1) questions; expert-revised question-answer pairs on 620 spectra (NMR, IR, XRD, Raman, MS, UV-Vis, XPS) from published papers; each answer judged correct or not by GPT-o4-mini against the reference, numeric answers within a 5 percent tolerance; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 20 models tracked.

Top models

#ModelScore
1Gemini 3 Flash (Preview)79
2Gemini 3 Pro (Preview)77.31
3Gemini 2.5 Pro75.71
4Gemini 2.5 Flash71.85
5GPT-5 (High)71.47
6GPT-571
7O3 (2025-04-16)70.43
8O4 Mini (2025-04-16)69.96
9GPT-5 (Low)69.87
10GPT-5.1 (Non-reasoning)62.9
11GPT-5.2 (Non-reasoning)61.58
12Claude Sonnet 4.5 (Thinking)57.82
13Claude Sonnet 4 (Thinking)55.65
14Qwen 3 VL 8B (Thinking)55.08
15Doubao-Seed-1.6 (Thinking)54.52

Interactive version: theaggregate.ai/benchmark?slug=specvqa-chinese-reasoning-l1 · How It Works · Data refreshed daily, snapshot 2026-10-07.