MolViBench (Chinese Prompts): leaderboard

Metric: Pass@1 (%) over all 358 tasks with the task prompts in Chinese, direct generation, RDKit molecular coding tasks answered by generated Python code, official APIs, temperature 0, at most 4,096 output tokens; higher is better. Source: arxiv.org. Saturation forecast: Around November 2027. 5 models tracked.

Top models

#ModelScore
1Claude Opus 4.638.3
2Gemini 3 Pro37.7
3GPT-5.3 Codex36.9
4DeepSeek V3.2 (Non-reasoning)33.2
5DeepSeek V3.2 (Thinking)29.3

Interactive version: theaggregate.ai/benchmark?slug=molvibench-chinese-prompts · How It Works · Data refreshed daily, snapshot 2026-10-07.