MolQuest: leaderboard

Metric: Structure accuracy (%): the share of cases whose predicted SMILES is canonically identical to the ground truth, in the interactive agent setting (the model requests simulated experiments one at a time: molecular weight and formula, 1H/13C/19F/31P NMR, IR, MS, melting point, TLC, optical rotation, with no limit on rounds, and submits a final SMILES), on the 530 MolQuest structure-elucidation cases (small molecules from post-2025 chemistry papers' supporting information), temperature 0; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 12 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3 Flash51.51#93
2Gemini 3 Pro48.3#77
3Claude Opus 4.525.66#79
4Gemini 2.5 Pro22.08#145
5Claude Sonnet 4.518.11#138
6DeepSeek V3.2 (Thinking)16.6#198 (DeepSeek V3.2)
7Qwen 3 Max15.28#201
8GPT-5.211.7#105
9Claude Haiku 4.511.51#271
10Kimi K2 (Thinking)11.32#236 (Kimi K2)
11DeepSeek V3.2 (Non-reasoning)11.32#198 (DeepSeek V3.2)
12DeepSeek V3.17.36#260

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=molquest · How It Works · Data refreshed daily, snapshot 2026-10-11.