MolViBench - Level 4 Analyze and Evaluate: leaderboard

Metric: Pass@1 (%) on the 75 Level 4 tasks (Analyze and Evaluate), direct generation in one pass with no execution feedback, RDKit molecular coding tasks answered by generated Python code, official APIs, temperature 0, at most 4,096 output tokens; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 9 models tracked.

Top models

#ModelScore
1Claude Opus 4.624
2Claude Opus 4.6 (Thinking)24
3Gemini 3 Pro21.3
4MiniMax-M2.514.7
5DeepSeek V3.2 (Thinking)13.3
6GPT-5.2 Codex9.3
7Kimi K2.56.7
8DeepSeek V3.2 (Non-reasoning)6.7
9GPT-5.3 Codex5.3

Interactive version: theaggregate.ai/benchmark?slug=molvibench-level-4-analyze-and-evaluate · How It Works · Data refreshed daily, snapshot 2026-10-07.