EHRBench - PROMOTE: leaderboard

Metric: Accuracy (%) on the questions built from PROMOTE records, EHR-grounded multiple-choice clinical decision questions (4, 5 and 6 options) built from MIMIC-III, MIMIC-IV and PROMOTE encounter trajectories with knowledge-base verification, JSON-constrained answers, deterministic decoding; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 31 models tracked.

Top models

#ModelScore
1GPT-5.2 Instant70.7
2GPT-5 Chat68.26
3GPT-4.167.87
4GLM-4 32B (0414)65.85
5Qwen 3 32B65.55
6Llama 3.3 70B Instruct65.35
7GPT-4.1 Mini64.45
8GPT-5 Mini63.88
9Mistral Small 363.38
10Qwen 2.5 32B63.05
11Qwen 3 8B58.74
12Qwen 3 4B58.39
13GPT-4.1 Nano58.28
14Yi 1.5 34B56.56
15GLM-4 9B (0414)56.41

Interactive version: theaggregate.ai/benchmark?slug=ehrbench-promote · How It Works · Data refreshed daily, snapshot 2026-10-07.