MR-Bench - Medication Imputation: leaderboard

Metric: Accuracy (%, times 100) at picking the masked medication among distractors that include interacting or contraindicated drugs; 1,000 MIMIC-IV hospital admissions, eight-option questions, temperature 0.5, an answer counted from the first or the last option letter the output names, whichever scores higher over the dataset; higher is better. Source: arxiv.org. Saturation forecast: Around April 2027. 24 models tracked.

Top models

#ModelScore
1Gemini 3 Pro55
2GPT-554.4
3DeepSeek V3.2 (Thinking)48.4
4Qwen 3 Max44.6
5GPT-4o36.3
6Qwen 3 4B 2507 (Thinking)27.2
7Qwen 3 8B26.9
8Gemma 3 4B (IT)23.7
9MedGemma-4B21.4
10Qwen 2.5 7B Instruct21.3
11Llama 3 8B Instruct16.6
12Llama 3.1 8B Instruct15.3

Interactive version: theaggregate.ai/benchmark?slug=mr-bench-medication-imputation · How It Works · Data refreshed daily, snapshot 2026-10-07.