MolViBench - Level 1 Remember and Understand: leaderboard
Metric: Pass@1 (%) on the 75 Level 1 tasks (Remember and Understand), direct generation in one pass with no execution feedback, RDKit molecular coding tasks answered by generated Python code, official APIs, temperature 0, at most 4,096 output tokens; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.6 | 78.7 |
| 2 | Claude Opus 4.6 (Thinking) | 78.7 |
| 3 | Gemini 3 Pro | 76 |
| 4 | GPT-5.3 Codex | 76 |
| 5 | GPT-5.2 Codex | 76 |
| 6 | Kimi K2.5 | 74.7 |
| 7 | DeepSeek V3.2 (Thinking) | 69.3 |
| 8 | DeepSeek V3.2 (Non-reasoning) | 68 |
| 9 | MiniMax-M2.5 | 64 |
Interactive version: theaggregate.ai/benchmark?slug=molvibench-level-1-remember-and-understand · How It Works · Data refreshed daily, snapshot 2026-10-07.