HELM MedQA — leaderboard
HELM MedQA: Evaluates clinical, biomedical, medical-exam, coding, or healthcare-document reasoning.
Metric: Exact match (self-reported). Source: benchmarklist.com. Status: saturation imminent. 13 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5 | 96.82 |
| 2 | GPT-5 Mini | 95.63 |
| 3 | O4 Mini | 94.83 |
| 4 | Gemini 2.5 Pro | 93.44 |
| 5 | O3 Mini | 92.05 |
| 6 | GPT-4o | 87.67 |
| 7 | Claude 3.5 Sonnet | 86.48 |
| 8 | Claude 3.7 Sonnet | 85.69 |
| 9 | Gemini 2.0 Flash | 84.89 |
| 10 | Gemini 1.5 Pro | 76.94 |
| 11 | GPT-4o Mini | 74.95 |
Interactive version: theaggregate.ai/benchmark?slug=helm-medqa · How the rankings work · Data refreshed daily, snapshot 2026-07-22.