MedLoCoMo - Single-Admission: leaderboard

Metric: Score (%) on the 8,946 single-admission questions: item-weighted combination of judged accuracy on answerable questions (two per admission) and abstention accuracy on unanswerable ones (one per admission); full patient history in context (100 synthetic multi-admission doctor-patient conversations from MIMIC-IV, 74.5k tokens on average), short free-text answers; answerable items judged correct or not by a Gemini 3 Flash Preview judge, adversarial unanswerable items scored by abstention; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 13 models tracked.

Top models

#ModelScore
1GPT-5.188.6
2Qwen 3.5 27B84.7
3Qwen 3.5 9B60.9
4Qwen 3.5 4B55.8
5Gemma 3 27B40.5
6Gemma 3 12B32.6
7MedGemma-4B21.8
8Lingshu-32B18.4
9Gemma 3 4B14.3

Interactive version: theaggregate.ai/benchmark?slug=medlocomo-single-admission · How It Works · Data refreshed daily, snapshot 2026-09-29.