MedQ-Deg - Clinical Understanding: leaderboard

Metric: Accuracy (%) on the Clinical Understanding questions of MedQ-Deg (capability dimension of the medical perception branch), pooled over the L1 (mild) and L2 (severe) degraded images, multiple-choice medical VQA items merged from OmniMedVQA, GMAI-MMBench and MedXpertQA, each image corrupted by modality-specific degradations whose severity three radiologists calibrated (24,894 QA pairs over 7 modalities and 18 degradation types), single-letter answer, temperature 1.0, mean of 3 runs; higher is better. Source: arxiv.org. Saturation forecast: Around September 2028. 40 models tracked.

Top models

#ModelScoreOverall rank
1Qwen 3 VL 235B A22B Instruct73#264
2Qwen 3 VL 30B A3B Instruct73#365
3GPT-567.2#91
4Gemini 2.5 Pro66.9#145
5GPT-4.166.5#240
6GLM-4.5V66.4#339
7GPT-4.1 Mini65.9#346
8GPT-5.165.9#131
9GPT-4o64.8#333
10Gemini 2.5 Flash63.6#237
11Qwen 2.5 VL 72B Instruct63.4#364
12Claude Sonnet 4.562.8#138
13MedGemma-4B62.1#731
14Qwen 2.5 VL 32B Instruct61.9#443
15Qwen 2 VL 7B Instruct61.7#816

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=medq-deg-clinical-understanding · How It Works · Data refreshed daily, snapshot 2026-10-11.