EHRNote-ChatQA - Evidence Grounding: leaderboard

Metric: QA-level evidence-grounding accuracy (%): choosing the minimal set of notes and section headers that supports the preceding content answer, micro-averaged over all turns, on 967 multi-turn samples, each over the full set of MIMIC-IV discharge summaries of one patient (1 to 5 notes), 8,036 content and 8,036 paired evidence-grounding 5-way multiple-choice questions (chance 20), chat history kept across turns, greedy decoding where settable; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 22 models tracked.

Top models

#ModelScore
1Gemini 3 Flash (Preview)90.1
2GPT-5.489.02
3Qwen 3 Next 80B A3B Instruct82.09
4Llama 4 Scout Instruct76
5Qwen 3 30B A3B 2507 Instruct74.9
6GPT-5.4 Mini74.05
7MedGemma-27B-IT74
8DeepSeek R1 Distill Llama 70B65.64
9Qwen 3 4B 2507 Instruct62.82
10DeepSeek R1 Distill Qwen 32B58.66
11Ministral-3-3B-Instruct-251255.46
12DeepSeek R1 Distill Qwen 14B53.94
13Ministral-3-8B-Instruct-251250.86
14Phi-4 Mini Instruct48.83
15MedGemma-4B-IT44.26

Interactive version: theaggregate.ai/benchmark?slug=ehrnote-chatqa-evidence-grounding · How It Works · Data refreshed daily, snapshot 2026-09-29.