MolViBench (Agent Collaboration): leaderboard

Metric: Pass@1 (%) over all 358 tasks under Agent Collaboration: a Coder and a Tester instance of the same model exchange test plans and PASS or FAIL feedback for up to 3 rounds, RDKit molecular coding tasks answered by generated Python code, official APIs, temperature 0, at most 4,096 output tokens; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 4 models tracked.

Top models

#ModelScore
1Gemini 3 Pro36
2Claude Opus 4.634.1
3DeepSeek V3.2 (Thinking)34.1
4Claude Opus 4.6 (Thinking)34.1

Interactive version: theaggregate.ai/benchmark?slug=molvibench-agent-collaboration · How It Works · Data refreshed daily, snapshot 2026-10-07.