MolViBench - Level 3 Analyze: leaderboard

Metric: Pass@1 (%) on the 72 Level 3 tasks (Analyze), direct generation in one pass with no execution feedback, RDKit molecular coding tasks answered by generated Python code, official APIs, temperature 0, at most 4,096 output tokens; higher is better. Source: arxiv.org. Saturation forecast: Around 2030. 9 models tracked.

Top models

#ModelScore
1Claude Opus 4.6 (Thinking)30.6
2Claude Opus 4.627.8
3MiniMax-M2.526.4
4GPT-5.2 Codex23.6
5Gemini 3 Pro19.4
6DeepSeek V3.2 (Thinking)19.4
7GPT-5.3 Codex16.7
8DeepSeek V3.2 (Non-reasoning)15.3
9Kimi K2.56.9

Interactive version: theaggregate.ai/benchmark?slug=molvibench-level-3-analyze · How It Works · Data refreshed daily, snapshot 2026-10-07.