MR-Bench - Procedure Selection: leaderboard

Metric: Accuracy (%, times 100) at picking the admission's procedure set against other patients' procedures; 1,000 MIMIC-IV hospital admissions, eight-option questions, temperature 0.5, an answer counted from the first or the last option letter the output names, whichever scores higher over the dataset; higher is better. Source: arxiv.org. Saturation forecast: Around January 2027. 24 models tracked.

Top models

#ModelScore
1GPT-573.8
2DeepSeek V3.2 (Thinking)69.4
3Qwen 3 Max68.2
4Gemini 3 Pro67.6
5GPT-4o64.9
6Qwen 3 8B60.2
7Qwen 2.5 7B Instruct59.5
8Qwen 3 4B 2507 (Thinking)58.8
9Gemma 3 4B (IT)50.4
10Llama 3 8B Instruct43.1
11Llama 3.1 8B Instruct41.6
12MedGemma-4B32.4

Interactive version: theaggregate.ai/benchmark?slug=mr-bench-procedure-selection · How It Works · Data refreshed daily, snapshot 2026-10-07.