URSA - Drugs and Clinicals ChemCensor: leaderboard

Metric: Average ChemCensor plausibility score (out of 5) per reaction step along the best route per target, averaged over the 100 URSA-drugs&clinicals-2026 approved drugs and clinical candidates; URSA-minor-1.1.0-U2 configuration with ChemCensor v1.2.0 and the U2 USPTO precedent database; default reasoning effort; higher is better. Source: arxiv.org. Saturation forecast: Around August 2027. 17 models tracked.

Top models

#ModelScore
1GPT-5.52.43
2Gemini 3.1 Pro (Preview)2.42
3Grok 4.12.01
4Claude Opus 4.81.99
5Claude Opus 4.71.97
6Kimi K2.51.78
7Claude Opus 4.51.71
8Qwen 3.5 397B A17B1.7
9Claude Opus 4.61.67
10Claude Sonnet 4.51.53
11GLM-51.48
12Claude Sonnet 4.61.44
13GPT-5.41.43
14GPT-5.21.33
15DeepSeek V3.21.06

Interactive version: theaggregate.ai/benchmark?slug=ursa-drugs-and-clinicals-chemcensor · How It Works · Data refreshed daily, snapshot 2026-09-29.