MedQ-Deg - Motion Degradation: leaderboard

Metric: Accuracy (%) on MedQ-Deg items whose images carry motion-interference degradations, pooled over the L1 (mild) and L2 (severe) severities, multiple-choice medical VQA items merged from OmniMedVQA, GMAI-MMBench and MedXpertQA, each image corrupted by modality-specific degradations whose severity three radiologists calibrated (24,894 QA pairs over 7 modalities and 18 degradation types), single-letter answer, temperature 1.0, mean of 3 runs; higher is better. Source: arxiv.org. Saturation forecast: Around June 2027. 40 models tracked.

Top models

#ModelScoreOverall rank
1Qwen 3 VL 235B A22B Instruct71.1#264
2Qwen 3 VL 30B A3B Instruct70.3#365
3Gemini 2.5 Flash68.7#237
4Gemini 2.5 Pro68.6#145
5GPT-568.4#91
6GPT-4.167.4#240
7GPT-5.167.2#131
8GPT-4.1 Mini66.8#346
9GPT-4o65#333
10GLM-4.5V63#339
11Gemma 3 27B59.1#596
12Qwen 2.5 VL 72B Instruct57.5#364
13Qwen 2.5 VL 32B Instruct56#443
14MedGemma-4B55.5#731
15Claude Sonnet 4.554.3#138

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=medq-deg-motion-degradation · How It Works · Data refreshed daily, snapshot 2026-10-11.