EHRNote-ChatQA - Content: leaderboard

Metric: QA-level content-question accuracy (%), micro-averaged over all turns, on 967 multi-turn samples, each over the full set of MIMIC-IV discharge summaries of one patient (1 to 5 notes), 8,036 content and 8,036 paired evidence-grounding 5-way multiple-choice questions (chance 20), chat history kept across turns, greedy decoding where settable; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 22 models tracked.

Top models

#ModelScore
1GPT-5.498.33
2Gemini 3 Flash (Preview)96.17
3Qwen 3 Next 80B A3B Instruct95.42
4Qwen 3 30B A3B 2507 Instruct93.12
5Llama 4 Scout Instruct92.15
6GPT-5.4 Mini91.63
7Qwen 3 4B 2507 Instruct88.4
8MedGemma-27B-IT87.76
9DeepSeek R1 Distill Llama 70B82.91
10DeepSeek R1 Distill Qwen 32B78.87
11Phi-3.5-mini-instruct76.94
12DeepSeek R1 Distill Qwen 14B76.14
13Ministral-3-3B-Instruct-251275.02
14Phi-4 Mini Instruct74.81
15DeepSeek R1 Distill Llama 8B65.31

Interactive version: theaggregate.ai/benchmark?slug=ehrnote-chatqa-content · How It Works · Data refreshed daily, snapshot 2026-09-29.