MedGemma-4B-IT: benchmark results
Provider: Other. Released 2025-05-20. Access: Open.
Unified ELO 1536 ± 22, rank #911 of 2656 rated models, from 33 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Medical MM Leaderboard - SLAKE | 76.4 | Accuracy (%) | 90.5 |
| Medical MM Leaderboard - VQA-RAD | 72.5 | Accuracy (%) | 90.5 |
| MIMIC-CDM - Cholecystitis | 61.6 | Accuracy (%) | 85.7 |
| HealthBench Hard | 45 | Overall score (self-reported) | 71.7 |
| Open Portuguese LLM - FaQuAD NLI | 73.37 | Macro F1 (%) | 68.6 |
| MedPRESS - Unsafe Agreement Rate | 46.2 | Unsafe-agreement answers (%; turns 1-5, LLM judge) | 68.4 |
| Open Portuguese LLM - ENEM | 67.25 | Accuracy (%) | 65.7 |
| FACTS Leaderboard | 37.71 | Combined Score (%) | 63.6 |
| MedPRESS - Conversation Failure Rate | 90.4 | Conversations with any unsafe agreement (%) | 63.2 |
| Medical MM Leaderboard - PathVQA | 48.8 | Accuracy (%) | 61.9 |
| Open Portuguese LLM - OAB Exams | 45.6 | Accuracy (%) | 61.4 |
| MEDIC Benchmark | 64.2 | MEDIC Public Table Average | 61.1 |
Interactive version: theaggregate.ai/model?slug=medgemma-4b-it · How It Works · Data refreshed daily, snapshot 2026-09-19.