MedScribe — leaderboard
Can models support doctors with their administrative work?.
Metric: Score (self-reported). Source: benchmarklist.com. Status: saturated. 53 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Claude Fable 5 (Max) | 88.52 |
| 2 | GPT-5.1 (High) | 88.09 |
| 3 | MiniMax-M3 | 87.25 |
| 4 | GPT-5.5 (xHigh) | 86.87 |
| 5 | Claude Opus 4.6 (Max) | 86.74 |
| 6 | Muse Spark | 85.9 |
| 7 | Claude Opus 4.8 (Max) | 85.75 |
| 8 | Claude Opus 4.5 (High) | 85.32 |
| 9 | Claude Haiku 4.5 | 85.23 |
| 10 | Claude Sonnet 4.5 | 84.52 |
| 11 | GPT-5.2 (xHigh) | 84.39 |
| 12 | GPT-5 (High) | 83.65 |
| 13 | Gemini 2.5 Flash (Thinking) | 82.98 |
| 14 | Claude Opus 4.7 (Max) | 82.95 |
| 15 | Gemini 2.5 Flash | 82.87 |
Interactive version: theaggregate.ai/benchmark?slug=medscribe · How the rankings work · Data refreshed daily, snapshot 2026-07-22.