Synthetic Hospital - Patient Diagnosis: leaderboard

Metric: Severity-weighted F1 (%; reconstruction of the patient longitudinal problem list as ICD-10 codes with acuity, chart-neutral scoring, chain-of-thought prompting; public split of 200 synthetic longitudinal patients (1,268 in all, built from USMLE-style questions with ontology-grounded labels); single-turn, one prompting strategy locked per task for every model; graph-derived deterministic scoring). Source: arxiv.org. Saturation forecast: Around December 2026. 10 models tracked.

Top models

#ModelScore
1Kimi K2.5 (Thinking)73.2
2GPT-5.370.3
3DeepSeek V3.266
4GLM-565.7
5Qwen 3.5 397B A17B65.3
6Claude Opus 4.661.5
7Llama 4 Scout53.6
8Gemma 3 27B28.7

Interactive version: theaggregate.ai/benchmark?slug=synthetic-hospital-patient-diagnosis · How It Works · Data refreshed daily, snapshot 2026-09-26.