EHRNote-ChatQA - All Turns Correct: leaderboard
Metric: Sample-level strict paired accuracy (%): share of the 967 patient samples in which every content and evidence-grounding turn is answered correctly, on 967 multi-turn samples, each over the full set of MIMIC-IV discharge summaries of one patient (1 to 5 notes), 8,036 content and 8,036 paired evidence-grounding 5-way multiple-choice questions (chance 20), chat history kept across turns, greedy decoding where settable; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 22 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.4 | 40.02 |
| 2 | Gemini 3 Flash (Preview) | 35.76 |
| 3 | Qwen 3 Next 80B A3B Instruct | 19.86 |
| 4 | Qwen 3 30B A3B 2507 Instruct | 10.03 |
| 5 | Llama 4 Scout Instruct | 9.51 |
| 6 | GPT-5.4 Mini | 8.69 |
| 7 | DeepSeek R1 Distill Llama 70B | 8.38 |
| 8 | MedGemma-27B-IT | 7.86 |
| 9 | DeepSeek R1 Distill Qwen 32B | 4.55 |
| 10 | Qwen 3 4B 2507 Instruct | 3 |
| 11 | Ministral-3-3B-Instruct-2512 | 1.76 |
| 12 | DeepSeek R1 Distill Qwen 14B | 1.45 |
| 13 | Ministral-3-8B-Instruct-2512 | 1.14 |
| 14 | Phi-4 Mini Instruct | 0.41 |
| 15 | Phi-3.5-mini-instruct | 0.31 |
Interactive version: theaggregate.ai/benchmark?slug=ehrnote-chatqa-all-turns-correct · How It Works · Data refreshed daily, snapshot 2026-09-29.