MR-Bench: leaderboard

Metric: Accuracy (%, times 100), mean of the medication-imputation and procedure-selection tasks; 1,000 MIMIC-IV hospital admissions, eight-option questions, temperature 0.5, an answer counted from the first or the last option letter the output names, whichever scores higher over the dataset; higher is better. Source: arxiv.org. Saturation forecast: Around February 2027. 24 models tracked.

Top models

#ModelScore
1GPT-564.1
2Gemini 3 Pro61.3
3DeepSeek V3.2 (Thinking)58.9
4Qwen 3 Max56.4
5GPT-4o50.6
6Qwen 3 8B43.6
7Qwen 3 4B 2507 (Thinking)43
8Qwen 2.5 7B Instruct40.4
9Gemma 3 4B (IT)37.1
10Llama 3 8B Instruct29.9
11Llama 3.1 8B Instruct28.5
12MedGemma-4B26.9

Interactive version: theaggregate.ai/benchmark?slug=mr-bench · How It Works · Data refreshed daily, snapshot 2026-10-07.