MedXpertQA — leaderboard
MedXpertQA evaluates model capability on healthcare & medical tasks from the linked upstream source with Score as the primary reported metric.
Metric: Score (self-reported). Source: benchmarklist.com. Status: saturation imminent. 19 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3.1 Pro (Preview) | 80.7 |
| 2 | GPT-5.4 | 77.3 |
| 3 | Qwen 3.7 Plus | 71 |
| 4 | Qwen 3.6 Plus | 68.7 |
| 5 | Claude Opus 4.6 (Max) | 64.4 |
| 6 | Gemma 4 31B | 61.3 |
| 7 | Gemma 4 26B A4B | 58.1 |
| 8 | O1 | 49.89 |
| 9 | Gemma 4 12B | 48.7 |
| 10 | GPT-4o | 35.96 |
| 11 | Gemma 4 E4B | 28.7 |
| 12 | Gemini 2.0 Flash | 28.04 |
| 13 | QVQ-72B-Preview | 27.06 |
| 14 | Claude 3.5 Sonnet | 26.65 |
| 15 | Gemini 1.5 Pro | 26.16 |
Interactive version: theaggregate.ai/benchmark?slug=medxpertqa · How the rankings work · Data refreshed daily, snapshot 2026-07-22.