MedCode: leaderboard

ICD-10-CM diagnosis coding from de-identified discharge summaries, progress and consult notes, scored on 2,755 primary and secondary codes; Vals AI with Harvard Medical School and Protege AI.

Metric: Accuracy (%). Source: www.vals.ai. Status: saturation imminent. 90 models tracked.

Top models

#ModelScore
1Claude Opus 5 (Max)63.57
2Gemini 3.1 Pro (Preview) (High)59.06
3Claude Fable 5 (Max)56.07
4Gemini 3 Flash (Preview) (High)55.92
5Gemini 3.5 Flash (High)55.83
6Claude Opus 4.7 (Max)54.86
7Claude Fable 5.1 (Max)53.51
8Gemini 3.7 Flash (High)53.39
9Claude Opus 4.8 (Max)53.22
10Gemini 3.6 Flash (High)53.15
11GPT-5.1 (High)52.73
12Gemini 3 Pro (Preview) (High)52.2
13Muse Spark51.31
14Gemini 2.5 Pro50.59
15GPT-5.2 (xHigh)49.75

Interactive version: theaggregate.ai/benchmark?slug=medcode · How It Works · Data refreshed daily, snapshot 2026-09-05.