EHRBench - Prognosis: leaderboard

Metric: Accuracy (%) on the prognosis-decision questions (predict a diagnosis at the next encounter), EHR-grounded multiple-choice clinical decision questions (4, 5 and 6 options) built from MIMIC-III, MIMIC-IV and PROMOTE encounter trajectories with knowledge-base verification, JSON-constrained answers, deterministic decoding; higher is better. Source: arxiv.org. Saturation forecast: Around August 2027. 31 models tracked.

Top models

#ModelScore
1GPT-5.2 Instant60.59
2GPT-5 Chat58.46
3GPT-4.158.33
4GPT-5 Mini56.79
5GPT-4.1 Mini56.05
6Qwen 3 32B55.04
7Llama 3.3 70B Instruct54.44
8Mistral Small 353.81
9GLM-4 32B (0414)53.36
10Qwen 2.5 32B51.54
11Qwen 3 8B49.93
12GPT-4.1 Nano49.39
13Qwen 3 4B48.76
14GLM-4 9B (0414)47.9
15Yi 1.5 34B47.86

Interactive version: theaggregate.ai/benchmark?slug=ehrbench-prognosis · How It Works · Data refreshed daily, snapshot 2026-10-07.