MedQ-Deg - Clean (L0): leaderboard

Metric: Accuracy (%) on the MedQ-Deg items with the original, undegraded images (L0), multiple-choice medical VQA items merged from OmniMedVQA, GMAI-MMBench and MedXpertQA, each image corrupted by modality-specific degradations whose severity three radiologists calibrated (24,894 QA pairs over 7 modalities and 18 degradation types), single-letter answer, temperature 1.0, mean of 3 runs; higher is better. Source: arxiv.org. Saturation forecast: Around June 2027. 40 models tracked.

Top models

#ModelScoreOverall rank
1Qwen 3 VL 235B A22B Instruct73.15#264
2Qwen 3 VL 30B A3B Instruct72.74#365
3Gemini 2.5 Pro71.39#145
4GPT-570.27#91
5GPT-5.168.56#131
6Gemini 2.5 Flash67.39#237
7GPT-4o66.98#333
8GPT-4.166#240
9GLM-4.5V65.77#339
10GPT-4.1 Mini65.21#346
11Gemma 3 27B61.03#596
12Qwen 2.5 VL 72B Instruct59.68#364
13Claude Sonnet 4.559.64#138
14MedGemma-4B58.48#731
15GPT-4o Mini57.32#588

No result here: #3 Claude Opus 5.5, #5 GPT-6 Astra, #8 Claude Fable 5.1.

Interactive version: theaggregate.ai/benchmark?slug=medq-deg-clean-l0 · How It Works · Data refreshed daily, snapshot 2026-10-11.