MedLoCoMo - Answerable Accuracy: leaderboard

Metric: LLM-judge accuracy (%) on the answerable questions of both scopes; full patient history in context (100 synthetic multi-admission doctor-patient conversations from MIMIC-IV, 74.5k tokens on average), short free-text answers; answerable items judged correct or not by a Gemini 3 Flash Preview judge, adversarial unanswerable items scored by abstention; higher is better. Source: arxiv.org. Saturation forecast: Around December 2026. 13 models tracked.

Top models

#ModelScore
1GPT-5.164.2
2Qwen 3.5 9B58.4
3Qwen 3.5 27B56.1
4Qwen 3.5 4B50.8
5Gemma 3 27B33.9
6Gemma 3 12B33.3
7Lingshu-32B24.8
8MedGemma-4B22.7
9Gemma 3 4B18.6

Interactive version: theaggregate.ai/benchmark?slug=medlocomo-answerable-accuracy · How It Works · Data refreshed daily, snapshot 2026-09-29.