EHRNote-ChatQA - All Turns Correct: leaderboard

Metric: Sample-level strict paired accuracy (%): share of the 967 patient samples in which every content and evidence-grounding turn is answered correctly, on 967 multi-turn samples, each over the full set of MIMIC-IV discharge summaries of one patient (1 to 5 notes), 8,036 content and 8,036 paired evidence-grounding 5-way multiple-choice questions (chance 20), chat history kept across turns, greedy decoding where settable; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 22 models tracked.

Top models

#ModelScore
1GPT-5.440.02
2Gemini 3 Flash (Preview)35.76
3Qwen 3 Next 80B A3B Instruct19.86
4Qwen 3 30B A3B 2507 Instruct10.03
5Llama 4 Scout Instruct9.51
6GPT-5.4 Mini8.69
7DeepSeek R1 Distill Llama 70B8.38
8MedGemma-27B-IT7.86
9DeepSeek R1 Distill Qwen 32B4.55
10Qwen 3 4B 2507 Instruct3
11Ministral-3-3B-Instruct-25121.76
12DeepSeek R1 Distill Qwen 14B1.45
13Ministral-3-8B-Instruct-25121.14
14Phi-4 Mini Instruct0.41
15Phi-3.5-mini-instruct0.31

Interactive version: theaggregate.ai/benchmark?slug=ehrnote-chatqa-all-turns-correct · How It Works · Data refreshed daily, snapshot 2026-09-29.