EHRNote-ChatQA - Content: leaderboard
Metric: QA-level content-question accuracy (%), micro-averaged over all turns, on 967 multi-turn samples, each over the full set of MIMIC-IV discharge summaries of one patient (1 to 5 notes), 8,036 content and 8,036 paired evidence-grounding 5-way multiple-choice questions (chance 20), chat history kept across turns, greedy decoding where settable; higher is better. Source: arxiv.org. Saturation forecast: Estimated already saturated. 22 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | GPT-5.4 | 98.33 |
| 2 | Gemini 3 Flash (Preview) | 96.17 |
| 3 | Qwen 3 Next 80B A3B Instruct | 95.42 |
| 4 | Qwen 3 30B A3B 2507 Instruct | 93.12 |
| 5 | Llama 4 Scout Instruct | 92.15 |
| 6 | GPT-5.4 Mini | 91.63 |
| 7 | Qwen 3 4B 2507 Instruct | 88.4 |
| 8 | MedGemma-27B-IT | 87.76 |
| 9 | DeepSeek R1 Distill Llama 70B | 82.91 |
| 10 | DeepSeek R1 Distill Qwen 32B | 78.87 |
| 11 | Phi-3.5-mini-instruct | 76.94 |
| 12 | DeepSeek R1 Distill Qwen 14B | 76.14 |
| 13 | Ministral-3-3B-Instruct-2512 | 75.02 |
| 14 | Phi-4 Mini Instruct | 74.81 |
| 15 | DeepSeek R1 Distill Llama 8B | 65.31 |
Interactive version: theaggregate.ai/benchmark?slug=ehrnote-chatqa-content · How It Works · Data refreshed daily, snapshot 2026-09-29.