URSA - Expert-2026 Solv-2: leaderboard

Metric: Solv-2 solved targets (%; share of the 100 URSA-expert-2026 novel, not-yet-synthesized drug-like targets with stock-terminated routes proposed for those targets whose every reaction also passes the ChemCensor precedent-based plausibility check (Solv-2), best of up to 10 routes per target, URSA-minor-1.1.0-U2 configuration with ChemCensor v1.2.0 and the U2 USPTO precedent database, shared stock of 255,365 building blocks; default reasoning effort; higher is better). Source: arxiv.org. Saturation forecast: Around January 2028. 17 models tracked.

Top models

#ModelScore
1GPT-5.521
2Gemini 3.1 Pro (Preview)13
3Claude Opus 4.810
4Claude Opus 4.78
5Grok 4.17
6Claude Opus 4.65
7Kimi K2.55
8Claude Sonnet 4.63
9Claude Opus 4.53
10Qwen 3.5 397B A17B2
11GLM-52
12Claude Sonnet 4.51
13GPT-5.11
14GPT-5.40
15GPT-5.20

Interactive version: theaggregate.ai/benchmark?slug=ursa-expert-2026-solv-2 · How It Works · Data refreshed daily, snapshot 2026-09-29.