EHRNote-ChatQA: leaderboard

Metric: QA-level paired accuracy (%): the content answer and its paired evidence-grounding answer both correct, micro-averaged over all turns, on 967 multi-turn samples, each over the full set of MIMIC-IV discharge summaries of one patient (1 to 5 notes), 8,036 content and 8,036 paired evidence-grounding 5-way multiple-choice questions (chance 20), chat history kept across turns, greedy decoding where settable; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 22 models tracked.

Top models

#ModelScore
1GPT-5.487.95
2Gemini 3 Flash (Preview)87.08
3Qwen 3 Next 80B A3B Instruct79.22
4Llama 4 Scout Instruct71.34
5Qwen 3 30B A3B 2507 Instruct70.78
6GPT-5.4 Mini68.96
7MedGemma-27B-IT65.57
8DeepSeek R1 Distill Llama 70B62.22
9Qwen 3 4B 2507 Instruct55.99
10DeepSeek R1 Distill Qwen 32B53.29
11DeepSeek R1 Distill Qwen 14B47.5
12Ministral-3-3B-Instruct-251244.3
13Phi-4 Mini Instruct36.78
14Phi-3.5-mini-instruct32.53
15Ministral-3-8B-Instruct-251231.89

Interactive version: theaggregate.ai/benchmark?slug=ehrnote-chatqa · How It Works · Data refreshed daily, snapshot 2026-09-29.