LiveMedBench — leaderboard
Live medical benchmark with time-stamped real-world cases and after-cutoff scoring for measuring medical model robustness over time.
Metric: Overall Score (%). Source: zhilingyan.github.io. Status: saturation imminent. 38 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.2 | 39.23 |
| 2 | GPT-5.1 | 38.45 |
| 3 | GPT-5 | 28.58 |
| 4 | Grok 4.1 | 28.28 |
| 5 | GPT-OSS-120B | 25.03 |
| 6 | GLM-4.5 | 22.46 |
| 7 | Gemini 3 Flash | 21.67 |
| 8 | Gemini 3 Pro | 18.29 |
| 9 | GLM-4.6 | 17.59 |
| 10 | Claude 3.7 Sonnet | 16.99 |
| 11 | Gemini 2.5 Pro | 16.06 |
| 12 | Qwen 3 14B | 15.45 |
| 13 | GPT-4.1 | 13.79 |
| 14 | QwQ-32B | 13.5 |
| 15 | GLM-4.7 (Thinking) | 13.35 |
Interactive version: theaggregate.ai/benchmark?slug=livemedbench · How the rankings work · Data refreshed daily, snapshot 2026-07-22.