MedGemma-4B: benchmark results
Provider: Other.
Unified ELO 1462 ± 16, rank #884 of 1607 rated models, from 60 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| LoMeVQA - Progress Classification | 52.6 | Accuracy (%; three-option classification of disease progress | 93.3 |
| EPAG - Top-k Diagnosis Accuracy | 82.31 | Disease Diagnosis Accuracy, Gold Anywhere in Ranked List (%) | 85 |
| Large Language Models Lack Temporal Awareness | 60.1 | Accuracy (self-reported) | 84.6 |
| Medmarks - MedHallu Hard | 50.37 | Score (%) | 82.9 |
| LoMeVQA - Progress Description | 20.95 | F1-RadGraph clinical efficacy (%; free-text description of t | 80 |
| PSF-Med - Open Medical - PadChest | 13.4 | Paraphrase Flip Rate (%) | 75 |
| EarlyDx | 32 | Primary-track F1 (%): micro-averaged F1 of the predicted fre | 66.7 |
| CardioLens | 51.53 | F1 (Random) (self-reported) | 65.2 |
| PSF-Med - Open Medical | 12.33 | Paraphrase Flip Rate (%) | 60 |
| LoMeVQA - Progress Report Generation | 10.73 | F1-RadGraph clinical efficacy (%; progress report written fr | 58.3 |
| AgentRx (Patient Summary, Few-Shot) - In-Hospital Mortality | 0.69 | AUROC (0-1; in-hospital mortality; MIMIC-IV ICU test split; | 50 |
| ECGQuest | 59.5 | Accuracy (%; 1,050 held-out true/false questions on ECG know | 50 |
Interactive version: theaggregate.ai/model?slug=medgemma-4b · How It Works · Data refreshed daily, snapshot 2026-09-29.