PhysicianBench: leaderboard

Long-horizon physician workflow benchmark grounded in clinical records, measuring checkpoint and end-to-end task success.

Metric: Pass@1 (self-reported). Source: benchmarklist.com. Status: saturation imminent. 12 models tracked.

Top models

#ModelScore
1GPT-5.546.3
2Claude Opus 4.631.7
3Claude Opus 4.729.3
4GPT-5.427.7
5Claude Sonnet 4.623
6DeepSeek V4 Pro18.7
7Kimi K2.617
8MiMo-V2.5-Pro16.7
9Qwen 3.6 Plus13.7
10MiniMax-M2.78.7
11Gemini 3.1 Pro (Preview)6
12Grok 4.205.3

Interactive version: theaggregate.ai/benchmark?slug=physicianbench · How It Works · Data refreshed daily, snapshot 2026-09-05.