MedGemma-4B-IT: benchmark results

Provider: Other. Released 2025-05-20. Access: Open.

Unified ELO 1536 ± 22, rank #911 of 2656 rated models, from 33 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Medical MM Leaderboard - SLAKE76.4Accuracy (%)90.5
Medical MM Leaderboard - VQA-RAD72.5Accuracy (%)90.5
MIMIC-CDM - Cholecystitis61.6Accuracy (%)85.7
HealthBench Hard45Overall score (self-reported)71.7
Open Portuguese LLM - FaQuAD NLI73.37Macro F1 (%)68.6
MedPRESS - Unsafe Agreement Rate46.2Unsafe-agreement answers (%; turns 1-5, LLM judge)68.4
Open Portuguese LLM - ENEM67.25Accuracy (%)65.7
FACTS Leaderboard37.71Combined Score (%)63.6
MedPRESS - Conversation Failure Rate90.4Conversations with any unsafe agreement (%)63.2
Medical MM Leaderboard - PathVQA48.8Accuracy (%)61.9
Open Portuguese LLM - OAB Exams45.6Accuracy (%)61.4
MEDIC Benchmark64.2MEDIC Public Table Average61.1

Interactive version: theaggregate.ai/model?slug=medgemma-4b-it · How It Works · Data refreshed daily, snapshot 2026-09-19.