EHRBench - Treatment: leaderboard

Metric: Accuracy (%) on the treatment-decision questions, EHR-grounded multiple-choice clinical decision questions (4, 5 and 6 options) built from MIMIC-III, MIMIC-IV and PROMOTE encounter trajectories with knowledge-base verification, JSON-constrained answers, deterministic decoding; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 31 models tracked.

Top models

#ModelScore
1GPT-5 Chat80.45
2GPT-5.2 Instant80.13
3GPT-4.180.1
4Llama 3.3 70B Instruct79.05
5GPT-4.1 Mini77.9
6GLM-4 32B (0414)77.9
7Qwen 3 32B77.34
8Qwen 2.5 32B76.87
9GPT-5 Mini76.17
10Mistral Small 375.02
11Qwen 3 8B74.49
12GPT-4.1 Nano74.03
13Qwen 3 4B73.46
14GLM-4 9B (0414)72.61
15Yi 1.5 34B72.25

Interactive version: theaggregate.ai/benchmark?slug=ehrbench-treatment · How It Works · Data refreshed daily, snapshot 2026-10-07.