MedGemma-27B-IT: benchmark results

Provider: Other. Released 2025-05-20. Access: Open.

Unified ELO 1655 ± 49, rank #441 of 2656 rated models, from 14 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Open Portuguese LLM - ASSIN2 RTE94.56Macro F1 (%)97
Open Portuguese LLM - FaQuAD NLI83.25Macro F1 (%)96.4
BRIDGE Medical Leaderboard - Few-Shot51.97Average Performance (%)92.6
Open Portuguese LLM - ENEM78.66Accuracy (%)92.2
Open Portuguese LLM - OAB Exams59.13Accuracy (%)91.7
Open Portuguese LLM - BLUEX69.26Accuracy (%)91.5
BRIDGE Medical Leaderboard43.66Average Performance (%)87
MedPRESS - Safe Stance Adherence Rate32.2Safe-stance answers (%; turns 1-5, LLM judge)84.2
BRIDGE Medical Leaderboard - Zero-Shot40.8Average Performance (%)83.3
BRIDGE Medical Leaderboard - CoT38.2Average Performance (%)80.6
MedPRESS - Conversation Failure Rate87.9Conversations with any unsafe agreement (%)68.4
MedPRESS - Turn of Flip2.16Pressure turns withstood before first unsafe agreement (0-5)68.4

Interactive version: theaggregate.ai/model?slug=medgemma-27b-it · How It Works · Data refreshed daily, snapshot 2026-09-19.