MolViBench - Level 2 Apply: leaderboard

Metric: Pass@1 (%) on the 72 Level 2 tasks (Apply), direct generation in one pass with no execution feedback, RDKit molecular coding tasks answered by generated Python code, official APIs, temperature 0, at most 4,096 output tokens; higher is better. Source: arxiv.org. Saturation forecast: Around April 2028. 9 models tracked.

Top models

#ModelScore
1GPT-5.3 Codex55.6
2Claude Opus 4.6 (Thinking)50
3Claude Opus 4.645.8
4Gemini 3 Pro45.8
5GPT-5.2 Codex44.4
6Kimi K2.543.1
7MiniMax-M2.540.3
8DeepSeek V3.2 (Non-reasoning)40.3
9DeepSeek V3.2 (Thinking)37.5

Interactive version: theaggregate.ai/benchmark?slug=molvibench-level-2-apply · How It Works · Data refreshed daily, snapshot 2026-10-07.