MolViBench - Level 1 Remember and Understand: leaderboard

Metric: Pass@1 (%) on the 75 Level 1 tasks (Remember and Understand), direct generation in one pass with no execution feedback, RDKit molecular coding tasks answered by generated Python code, official APIs, temperature 0, at most 4,096 output tokens; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 9 models tracked.

Top models

#ModelScore
1Claude Opus 4.678.7
2Claude Opus 4.6 (Thinking)78.7
3Gemini 3 Pro76
4GPT-5.3 Codex76
5GPT-5.2 Codex76
6Kimi K2.574.7
7DeepSeek V3.2 (Thinking)69.3
8DeepSeek V3.2 (Non-reasoning)68
9MiniMax-M2.564

Interactive version: theaggregate.ai/benchmark?slug=molvibench-level-1-remember-and-understand · How It Works · Data refreshed daily, snapshot 2026-10-07.