MedBench v5 - CCR-Multimodal: leaderboard

Metric: Macro-average score (0-100) over the 12 tasks of the MedBench v5 Clinical Cognitive Responsiveness multimodal track (lesion detection, image classification, report OCR, visual question answering, report generation, longitudinal imaging and multimodal clinical decision support), each task normalized to 0-100; higher is better. Source: arxiv.org. Saturation forecast: Around September 2027. 10 models tracked.

Top models

#ModelScore
1Seed 2.0 Pro51.65
2Gemini 3.5 Flash50.68
3GPT-5.550.07
4Qwen 3.7 Plus49.24
5Claude Opus 4.747.98
6Kimi K2.646.4
7GLM-5.142.35
8DeepSeek V4 Pro41.18
9Grok 4.2035.09

Interactive version: theaggregate.ai/benchmark?slug=medbench-v5-ccr-multimodal · How It Works · Data refreshed daily, snapshot 2026-09-29.