MolQuest (Static): leaderboard

Metric: Structure accuracy (%): the share of cases whose predicted SMILES is canonically identical to the ground truth, in the static setting (all available spectral data given at once, one answer), on the 530 MolQuest structure-elucidation cases (small molecules from post-2025 chemistry papers' supporting information), temperature 0; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 12 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3 Pro52.08#77
2Gemini 3 Flash51.13#93
3Gemini 2.5 Pro30.19#145
4Claude Opus 4.528.49#79
5Kimi K2 (Thinking)20.57#236 (Kimi K2)
6DeepSeek V3.2 (Thinking)20.38#198 (DeepSeek V3.2)
7Claude Sonnet 4.517.55#138
8Claude Haiku 4.59.62#271
9GPT-5.27.36#105
10DeepSeek V3.16.79#260
11DeepSeek V3.2 (Non-reasoning)5.66#198 (DeepSeek V3.2)
12Qwen 3 Max4.72#201

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=molquest-static · How It Works · Data refreshed daily, snapshot 2026-10-11.