Synthetic Hospital - Imaging Indication: leaderboard

Metric: Concept F1 (%; inferring the clinical question behind a vague imaging order, scored on graph-linked diagnosis and finding concepts, few-shot; public split of 200 synthetic longitudinal patients (1,268 in all, built from USMLE-style questions with ontology-grounded labels); single-turn, one prompting strategy locked per task for every model; graph-derived deterministic scoring). Source: arxiv.org. Saturation forecast: Around 2031. 10 models tracked.

Top models

#ModelScore
1GPT-5.351.8
2Kimi K2.5 (Thinking)51.7
3GLM-550.5
4Qwen 3.5 397B A17B49.7
5DeepSeek V3.249.2
6Claude Opus 4.647.3
7Gemma 3 27B41.8
8Llama 4 Scout40.6

Interactive version: theaggregate.ai/benchmark?slug=synthetic-hospital-imaging-indication · How It Works · Data refreshed daily, snapshot 2026-09-26.