EHRBench - Diagnosis: leaderboard

Metric: Accuracy (%) on the diagnosis-decision questions (complete a withheld diagnosis of an encounter), EHR-grounded multiple-choice clinical decision questions (4, 5 and 6 options) built from MIMIC-III, MIMIC-IV and PROMOTE encounter trajectories with knowledge-base verification, JSON-constrained answers, deterministic decoding; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 31 models tracked.

Top models

#ModelScore
1GPT-5.2 Instant72.02
2GPT-4.169.87
3Llama 3.3 70B Instruct68.35
4GPT-5 Chat68.26
5Qwen 3 32B67.97
6GLM-4 32B (0414)67.09
7Qwen 2.5 32B66.48
8GPT-4.1 Mini66.41
9Mistral Small 366.2
10GPT-5 Mini65.4
11Qwen 3 4B59.67
12GLM-4 9B (0414)58.36
13Qwen 3 8B58.19
14GPT-4.1 Nano58.02
15Yi 1.5 34B56.7

Interactive version: theaggregate.ai/benchmark?slug=ehrbench-diagnosis · How It Works · Data refreshed daily, snapshot 2026-10-07.