MolViBench - Level 4 Analyze and Evaluate: leaderboard
Metric: Pass@1 (%) on the 75 Level 4 tasks (Analyze and Evaluate), direct generation in one pass with no execution feedback, RDKit molecular coding tasks answered by generated Python code, official APIs, temperature 0, at most 4,096 output tokens; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 9 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Opus 4.6 | 24 |
| 2 | Claude Opus 4.6 (Thinking) | 24 |
| 3 | Gemini 3 Pro | 21.3 |
| 4 | MiniMax-M2.5 | 14.7 |
| 5 | DeepSeek V3.2 (Thinking) | 13.3 |
| 6 | GPT-5.2 Codex | 9.3 |
| 7 | Kimi K2.5 | 6.7 |
| 8 | DeepSeek V3.2 (Non-reasoning) | 6.7 |
| 9 | GPT-5.3 Codex | 5.3 |
Interactive version: theaggregate.ai/benchmark?slug=molvibench-level-4-analyze-and-evaluate · How It Works · Data refreshed daily, snapshot 2026-10-07.