LiveMedBench — leaderboard

Live medical benchmark with time-stamped real-world cases and after-cutoff scoring for measuring medical model robustness over time.

Metric: Overall Score (%). Source: zhilingyan.github.io. Status: saturation imminent. 38 models tracked.

Top models

#ModelScore
1GPT-5.239.23
2GPT-5.138.45
3GPT-528.58
4Grok 4.128.28
5GPT-OSS-120B25.03
6GLM-4.522.46
7Gemini 3 Flash21.67
8Gemini 3 Pro18.29
9GLM-4.617.59
10Claude 3.7 Sonnet16.99
11Gemini 2.5 Pro16.06
12Qwen 3 14B15.45
13GPT-4.113.79
14QwQ-32B13.5
15GLM-4.7 (Thinking)13.35

Interactive version: theaggregate.ai/benchmark?slug=livemedbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.