MedFlowBench (Radiology): leaderboard
Metric: Strict answer-plus-evidence score (0-1, shown times 100), credit only when the answer is right and the module's evidence check passes against hidden masks and labels, averaged without weights over the three radiology modules (LUMIERE longitudinal brain MRI (139 questions: RANO response category from baseline and follow-up studies); UCSF-PDGM multi-sequence brain MRI (495 cases: tumor diagnosis); NSCLC lung PET/CT (162 cases: tumor location, pathological T and N stage, histology and grade)); the model drives 3D Slicer through the MedOpenClaw runtime with viewer-native actions only (series selection, scrolling, windowing, prealigned fusion), inspecting the complete study within a 20-round budget and returning an answer with structured evidence (Track A); higher is better. Source: arxiv.org. Saturation forecast: Around March 2028. 14 models tracked.
Top models
| # | Model | Score | Overall rank |
|---|---|---|---|
| 1 | Gemini 3.1 Pro (Preview) | 32 | #54 |
| 2 | Gemini 3 Flash (Preview) | 26 | #78 |
| 3 | GPT-5.5 | 24 | #26 |
| 4 | Claude Sonnet 4.6 | 12 | #85 |
| 5 | Qwen 3.5 27B | 12 | #183 |
| 6 | Qwen 3.5 35B A3B | 11 | #250 |
| 7 | Claude Opus 4.7 | 10 | #45 |
| 8 | GPT-5.4 Mini | 8 | #202 |
| 9 | Qwen 3.5 4B | 6 | #470 |
| 10 | Qwen 3.5 9B | 5 | #363 |
| 11 | Gemma 3 27B (IT) | 4 | #509 |
| 12 | MedGemma 1.5 4B | 4 | #775 |
No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.
Interactive version: theaggregate.ai/benchmark?slug=medflowbench-radiology · How It Works · Data refreshed daily, snapshot 2026-10-11.