MedCode — leaderboard

MedCode evaluates model capability on healthcare & medical tasks from the linked upstream source with Score as the primary reported metric.

Metric: Score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 54 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview) (High)59.06
2Claude Fable 5 (Max)56.07
3Gemini 3 Flash (Preview) (High)55.92
4Gemini 3.5 Flash (High)55.83
5Claude Opus 4.7 (Max)54.86
6Claude Opus 4.8 (Max)53.22
7GPT-5.1 (High)52.73
8Gemini 3 Pro (High)52.2
9Muse Spark51.31
10Gemini 2.5 Pro50.59
11GPT-5.2 (xHigh)49.75
12GPT-5 (High)49.63
13Claude Opus 4.5 (High)49.16
14Claude Opus 4.6 (Max)49.13
15GPT-5.5 (xHigh)49.1

Interactive version: theaggregate.ai/benchmark?slug=medcode · How the rankings work · Data refreshed daily, snapshot 2026-07-22.