EHRBench - Diagnosis: leaderboard
Metric: Accuracy (%) on the diagnosis-decision questions (complete a withheld diagnosis of an encounter), EHR-grounded multiple-choice clinical decision questions (4, 5 and 6 options) built from MIMIC-III, MIMIC-IV and PROMOTE encounter trajectories with knowledge-base verification, JSON-constrained answers, deterministic decoding; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 31 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.2 Instant | 72.02 |
| 2 | GPT-4.1 | 69.87 |
| 3 | Llama 3.3 70B Instruct | 68.35 |
| 4 | GPT-5 Chat | 68.26 |
| 5 | Qwen 3 32B | 67.97 |
| 6 | GLM-4 32B (0414) | 67.09 |
| 7 | Qwen 2.5 32B | 66.48 |
| 8 | GPT-4.1 Mini | 66.41 |
| 9 | Mistral Small 3 | 66.2 |
| 10 | GPT-5 Mini | 65.4 |
| 11 | Qwen 3 4B | 59.67 |
| 12 | GLM-4 9B (0414) | 58.36 |
| 13 | Qwen 3 8B | 58.19 |
| 14 | GPT-4.1 Nano | 58.02 |
| 15 | Yi 1.5 34B | 56.7 |
Interactive version: theaggregate.ai/benchmark?slug=ehrbench-diagnosis · How It Works · Data refreshed daily, snapshot 2026-10-07.