SMDD-Bench - Lead Optimization: leaderboard

Metric: Success rate (%) on the 340 SMDD-Bench lead-optimization tasks (improve ADMET objectives while holding constraints and the binding interaction), minimalist ReAct agent harness with RDKit-style tools, 8 Boltz-2 and 15 ADMET-AI oracle calls per task, temperature 1.0, at most 100 turns; success means the submitted molecule passes the task's hidden evaluator; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 7 models tracked.

Top models

#ModelScore
1GPT-5.4 (Medium)57.6
2Gemini 3.1 Pro (Preview) (Medium)55.6
3Claude Sonnet 4.653.5
4Kimi K2.5 (Thinking)43.5
5Qwen 3.5 397B A17B40
6DeepSeek V3.234.7
7MiniMax-M2.727.1

Interactive version: theaggregate.ai/benchmark?slug=smdd-bench-lead-optimization · How It Works · Data refreshed daily, snapshot 2026-10-07.