EHRNote-ChatQA - Evidence Grounding: leaderboard
Metric: QA-level evidence-grounding accuracy (%): choosing the minimal set of notes and section headers that supports the preceding content answer, micro-averaged over all turns, on 967 multi-turn samples, each over the full set of MIMIC-IV discharge summaries of one patient (1 to 5 notes), 8,036 content and 8,036 paired evidence-grounding 5-way multiple-choice questions (chance 20), chat history kept across turns, greedy decoding where settable; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 22 models tracked.
Top models
| # | Model | Score |
|---|---|---|
| 1 | Gemini 3 Flash (Preview) | 90.1 |
| 2 | GPT-5.4 | 89.02 |
| 3 | Qwen 3 Next 80B A3B Instruct | 82.09 |
| 4 | Llama 4 Scout Instruct | 76 |
| 5 | Qwen 3 30B A3B 2507 Instruct | 74.9 |
| 6 | GPT-5.4 Mini | 74.05 |
| 7 | MedGemma-27B-IT | 74 |
| 8 | DeepSeek R1 Distill Llama 70B | 65.64 |
| 9 | Qwen 3 4B 2507 Instruct | 62.82 |
| 10 | DeepSeek R1 Distill Qwen 32B | 58.66 |
| 11 | Ministral-3-3B-Instruct-2512 | 55.46 |
| 12 | DeepSeek R1 Distill Qwen 14B | 53.94 |
| 13 | Ministral-3-8B-Instruct-2512 | 50.86 |
| 14 | Phi-4 Mini Instruct | 48.83 |
| 15 | MedGemma-4B-IT | 44.26 |
Interactive version: theaggregate.ai/benchmark?slug=ehrnote-chatqa-evidence-grounding · How It Works · Data refreshed daily, snapshot 2026-09-29.