PhysicianBench — leaderboard

Long-horizon physician workflow benchmark grounded in clinical records, measuring checkpoint and end-to-end task success.

Metric: Pass@1 (self-reported). Source: benchmarklist.com. Status: saturation imminent. 12 models tracked.

Top models

#ModelScore
1GPT-5.546.3
2Claude Opus 4.631.7
3Claude Opus 4.729.3
4GPT-5.427.7
5Claude Sonnet 4.623
6DeepSeek V4 Pro18.7
7MiMo-V2.5-Pro16.7
8Qwen 3.6 Plus13.7
9MiniMax-M2.78.7
10Gemini 3.1 Pro (Preview)6
11Grok 4.205.3

Interactive version: theaggregate.ai/benchmark?slug=physicianbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.