MedGemma-4B: benchmark results

Provider: Other.

Unified ELO 1462 ± 16, rank #884 of 1607 rated models, from 60 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
LoMeVQA - Progress Classification52.6Accuracy (%; three-option classification of disease progress93.3
EPAG - Top-k Diagnosis Accuracy82.31Disease Diagnosis Accuracy, Gold Anywhere in Ranked List (%)85
Large Language Models Lack Temporal Awareness60.1Accuracy (self-reported)84.6
Medmarks - MedHallu Hard50.37Score (%)82.9
LoMeVQA - Progress Description20.95F1-RadGraph clinical efficacy (%; free-text description of t80
PSF-Med - Open Medical - PadChest13.4Paraphrase Flip Rate (%)75
EarlyDx32Primary-track F1 (%): micro-averaged F1 of the predicted fre66.7
CardioLens51.53F1 (Random) (self-reported)65.2
PSF-Med - Open Medical12.33Paraphrase Flip Rate (%)60
LoMeVQA - Progress Report Generation10.73F1-RadGraph clinical efficacy (%; progress report written fr58.3
AgentRx (Patient Summary, Few-Shot) - In-Hospital Mortality0.69AUROC (0-1; in-hospital mortality; MIMIC-IV ICU test split; 50
ECGQuest59.5Accuracy (%; 1,050 held-out true/false questions on ECG know50

Interactive version: theaggregate.ai/model?slug=medgemma-4b · How It Works · Data refreshed daily, snapshot 2026-09-29.