MedHELM — leaderboard

Clinically grounded medical evaluation suite built on HELM, covering medical reasoning, safety, fairness, robustness, and specialty-specific tasks.

Metric: Mean Win Rate (self-reported). Source: benchmarklist.com. Status: saturation imminent. 9 models tracked.

Top models

#ModelScore
1O3 Mini64.11
2Claude 3.7 Sonnet63.57
3Claude 3.5 Sonnet63.39
4GPT-4o56.96
5Gemini 2.0 Flash41.96
6GPT-4o Mini39.29
7Gemini 1.5 Pro24.11

Interactive version: theaggregate.ai/benchmark?slug=medhelm · How the rankings work · Data refreshed daily, snapshot 2026-07-22.