MedFlowBench (Radiology): leaderboard

Metric: Strict answer-plus-evidence score (0-1, shown times 100), credit only when the answer is right and the module's evidence check passes against hidden masks and labels, averaged without weights over the three radiology modules (LUMIERE longitudinal brain MRI (139 questions: RANO response category from baseline and follow-up studies); UCSF-PDGM multi-sequence brain MRI (495 cases: tumor diagnosis); NSCLC lung PET/CT (162 cases: tumor location, pathological T and N stage, histology and grade)); the model drives 3D Slicer through the MedOpenClaw runtime with viewer-native actions only (series selection, scrolling, windowing, prealigned fusion), inspecting the complete study within a 20-round budget and returning an answer with structured evidence (Track A); higher is better. Source: arxiv.org. Saturation forecast: Around March 2028. 14 models tracked.

Top models

#ModelScoreOverall rank
1Gemini 3.1 Pro (Preview)32#54
2Gemini 3 Flash (Preview)26#78
3GPT-5.524#26
4Claude Sonnet 4.612#85
5Qwen 3.5 27B12#183
6Qwen 3.5 35B A3B11#250
7Claude Opus 4.710#45
8GPT-5.4 Mini8#202
9Qwen 3.5 4B6#470
10Qwen 3.5 9B5#363
11Gemma 3 27B (IT)4#509
12MedGemma 1.5 4B4#775

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=medflowbench-radiology · How It Works · Data refreshed daily, snapshot 2026-10-11.