EHRNote-ChatQA: leaderboard
Metric: QA-level paired accuracy (%): the content answer and its paired evidence-grounding answer both correct, micro-averaged over all turns, on 967 multi-turn samples, each over the full set of MIMIC-IV discharge summaries of one patient (1 to 5 notes), 8,036 content and 8,036 paired evidence-grounding 5-way multiple-choice questions (chance 20), chat history kept across turns, greedy decoding where settable; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 22 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.4 | 87.95 |
| 2 | Gemini 3 Flash (Preview) | 87.08 |
| 3 | Qwen 3 Next 80B A3B Instruct | 79.22 |
| 4 | Llama 4 Scout Instruct | 71.34 |
| 5 | Qwen 3 30B A3B 2507 Instruct | 70.78 |
| 6 | GPT-5.4 Mini | 68.96 |
| 7 | MedGemma-27B-IT | 65.57 |
| 8 | DeepSeek R1 Distill Llama 70B | 62.22 |
| 9 | Qwen 3 4B 2507 Instruct | 55.99 |
| 10 | DeepSeek R1 Distill Qwen 32B | 53.29 |
| 11 | DeepSeek R1 Distill Qwen 14B | 47.5 |
| 12 | Ministral-3-3B-Instruct-2512 | 44.3 |
| 13 | Phi-4 Mini Instruct | 36.78 |
| 14 | Phi-3.5-mini-instruct | 32.53 |
| 15 | Ministral-3-8B-Instruct-2512 | 31.89 |
Interactive version: theaggregate.ai/benchmark?slug=ehrnote-chatqa · How It Works · Data refreshed daily, snapshot 2026-09-29.