EHRBench - MIMIC-IV: leaderboard

Metric: Accuracy (%) on the questions built from MIMIC-IV records, EHR-grounded multiple-choice clinical decision questions (4, 5 and 6 options) built from MIMIC-III, MIMIC-IV and PROMOTE encounter trajectories with knowledge-base verification, JSON-constrained answers, deterministic decoding; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 31 models tracked.

Top models

#ModelScore
1GPT-5.2 Instant71.06
2GPT-4.169.77
3GPT-5 Chat69.18
4Llama 3.3 70B Instruct67.94
5GPT-4.1 Mini67.36
6Qwen 3 32B67.18
7Qwen 2.5 32B67
8GLM-4 32B (0414)66.45
9GPT-5 Mini66.05
10Mistral Small 365.36
11Qwen 3 8B62.11
12GPT-4.1 Nano61.89
13Qwen 3 4B61.73
14GLM-4 9B (0414)61.11
15Yi 1.5 34B60.07

Interactive version: theaggregate.ai/benchmark?slug=ehrbench-mimic-iv · How It Works · Data refreshed daily, snapshot 2026-10-07.