MedScribe — leaderboard

Can models support doctors with their administrative work?.

Metric: Score (self-reported). Source: benchmarklist.com. Status: saturated. 53 models tracked.

Top models

#ModelScore
1Claude Fable 5 (Max)88.52
2GPT-5.1 (High)88.09
3MiniMax-M387.25
4GPT-5.5 (xHigh)86.87
5Claude Opus 4.6 (Max)86.74
6Muse Spark85.9
7Claude Opus 4.8 (Max)85.75
8Claude Opus 4.5 (High)85.32
9Claude Haiku 4.585.23
10Claude Sonnet 4.584.52
11GPT-5.2 (xHigh)84.39
12GPT-5 (High)83.65
13Gemini 2.5 Flash (Thinking)82.98
14Claude Opus 4.7 (Max)82.95
15Gemini 2.5 Flash82.87

Interactive version: theaggregate.ai/benchmark?slug=medscribe · How the rankings work · Data refreshed daily, snapshot 2026-07-22.