MedAgentBench — leaderboard

Interactive EHR-agent benchmark with physician-written tasks over healthcare data and FHIR-style clinical workflows.

Metric: Overall SR (self-reported). Source: benchmarklist.com. Status: saturation imminent. 12 models tracked.

Top models

#ModelScore
1Claude 3.5 Sonnet69.67
2GPT-4o64
3DeepSeek V362.67
4Gemini 1.5 Pro62
5GPT-4o Mini56.33
6O3 Mini51.67
7Llama 3.3 70B46.33
8Gemini 2.0 Flash38.33

Interactive version: theaggregate.ai/benchmark?slug=medagentbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.