CURE (Clinical Diagnosis, History, Images and Findings) - Top-3: leaderboard

Metric: Open-ended diagnosis, Hit@3 in percent: share of cases whose reference diagnosis is among the model's top 3 ranked diagnoses (matched by a DeepSeek-V3.1 judge, kappa 0.86 against a physician on 50 cases), on CURE's 500 test cases (Eurorad 2025 radiology cases: clinical history, a mean of 8 images, radiologist findings with diagnostic statements removed), zero-shot, given the clinical history, images and the radiologist's written findings; higher is better. Source: arxiv.org. Saturation forecast: Around 2029. 16 models tracked.

Top models

#ModelScoreOverall rank
1GPT-5.265.8#105
2Gemini 3 Pro (Preview)65#64
3Kimi K2.563.2#139
4Gemma 3 12B (IT)60.4#655
5Qwen 3 VL 235B A22B Instruct58.6#264
6Gemma 3 4B (IT)58#971
7GLM-4.6V57.2#309
8Qwen 3 VL 235B A22B (Thinking)56#228 (Qwen 3 VL 235B A22B)
9Qwen 3 VL 32B Instruct53#276
10Qwen 3 VL 30B A3B Instruct51.6#365
11Qwen 3 VL 32B (Thinking)50.6#287 (Qwen 3 VL 32B)
12MedGemma-4B-IT50.6#842
13Qwen 3 VL 8B Instruct49.8#401
14Qwen 3 VL 30B A3B (Thinking)45#338 (Qwen 3 VL 30B A3B)
15Qwen 3 VL 8B (Thinking)43.4

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=cure-clinical-diagnosis-history-images-and-findings-top-3 · How It Works · Data refreshed daily, snapshot 2026-10-11.