Synthetic Hospital - Summarization: leaderboard

Metric: Finding-level F1 (%; whole-patient clinical summary scored against graph-defined key findings, ontology-grounded structured prompting; public split of 200 synthetic longitudinal patients (1,268 in all, built from USMLE-style questions with ontology-grounded labels); single-turn, one prompting strategy locked per task for every model; graph-derived deterministic scoring). Source: arxiv.org. Saturation forecast: Around 2029. 10 models tracked.

Top models

#ModelScore
1Claude Opus 4.655
2Kimi K2.5 (Thinking)53.2
3GPT-5.348.9
4GLM-548.4
5Qwen 3.5 397B A17B42.5
6Llama 4 Scout40.3
7DeepSeek V3.239.9
8Gemma 3 27B38.5

Interactive version: theaggregate.ai/benchmark?slug=synthetic-hospital-summarization · How It Works · Data refreshed daily, snapshot 2026-09-26.