MedXpertQA — leaderboard

MedXpertQA evaluates model capability on healthcare & medical tasks from the linked upstream source with Score as the primary reported metric.

Metric: Score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 19 models tracked.

Top models

#ModelScore
1Gemini 3.1 Pro (Preview)80.7
2GPT-5.477.3
3Qwen 3.7 Plus71
4Qwen 3.6 Plus68.7
5Claude Opus 4.6 (Max)64.4
6Gemma 4 31B61.3
7Gemma 4 26B A4B58.1
8O149.89
9Gemma 4 12B48.7
10GPT-4o35.96
11Gemma 4 E4B28.7
12Gemini 2.0 Flash28.04
13QVQ-72B-Preview27.06
14Claude 3.5 Sonnet26.65
15Gemini 1.5 Pro26.16

Interactive version: theaggregate.ai/benchmark?slug=medxpertqa · How the rankings work · Data refreshed daily, snapshot 2026-07-22.