MolQuest (Static): leaderboard
Metric: Structure accuracy (%): the share of cases whose predicted SMILES is canonically identical to the ground truth, in the static setting (all available spectral data given at once, one answer), on the 530 MolQuest structure-elucidation cases (small molecules from post-2025 chemistry papers' supporting information), temperature 0; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 12 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Gemini 3 Pro | 52.08 | #77 |
| 2 | Gemini 3 Flash | 51.13 | #93 |
| 3 | Gemini 2.5 Pro | 30.19 | #145 |
| 4 | Claude Opus 4.5 | 28.49 | #79 |
| 5 | Kimi K2 (Thinking) | 20.57 | #236 (Kimi K2) |
| 6 | DeepSeek V3.2 (Thinking) | 20.38 | #198 (DeepSeek V3.2) |
| 7 | Claude Sonnet 4.5 | 17.55 | #138 |
| 8 | Claude Haiku 4.5 | 9.62 | #271 |
| 9 | GPT-5.2 | 7.36 | #105 |
| 10 | DeepSeek V3.1 | 6.79 | #260 |
| 11 | DeepSeek V3.2 (Non-reasoning) | 5.66 | #198 (DeepSeek V3.2) |
| 12 | Qwen 3 Max | 4.72 | #201 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=molquest-static · How It Works · Data refreshed daily, snapshot 2026-10-11.