MedAgentBench: leaderboard

Interactive EHR-agent benchmark with physician-written tasks over healthcare data and FHIR-style clinical workflows.

Metric: Overall SR (self-reported). Source: benchmarklist.com. Status: saturated. 12 models tracked.

Top models

#ModelScore
1GPT-4o64
2DeepSeek V362.67
3Gemini 1.5 Pro62
4GPT-4o Mini56.33
5O3 Mini51.67
6Llama 3.3 70B46.33
7Gemini 2.0 Flash38.33

Interactive version: theaggregate.ai/benchmark?slug=medagentbench · How It Works · Data refreshed daily, snapshot 2026-09-05.