MedLoCoMo: leaderboard

Metric: Score (%) over all 17,892 questions: item-weighted combination of judged accuracy on answerable questions (two per admission) and abstention accuracy on unanswerable ones (one per admission); full patient history in context (100 synthetic multi-admission doctor-patient conversations from MIMIC-IV, 74.5k tokens on average), short free-text answers; answerable items judged correct or not by a Gemini 3 Flash Preview judge, adversarial unanswerable items scored by abstention; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 13 models tracked.

Top models

#ModelScore
1GPT-5.172.8
2Qwen 3.5 27B68.1
3Qwen 3.5 9B51.8
4Qwen 3.5 4B48.8
5Gemma 3 27B40.3
6Gemma 3 12B35.7
7MedGemma-4B18.4
8Lingshu-32B16.5
9Gemma 3 4B12.6

Interactive version: theaggregate.ai/benchmark?slug=medlocomo · How It Works · Data refreshed daily, snapshot 2026-09-29.