MedGemma-27B-IT: benchmark results
Provider: Other. Released 2025-05-20. Access: Open.
Unified ELO 1655 ± 49, rank #441 of 2656 rated models, from 14 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Open Portuguese LLM - ASSIN2 RTE | 94.56 | Macro F1 (%) | 97 |
| Open Portuguese LLM - FaQuAD NLI | 83.25 | Macro F1 (%) | 96.4 |
| BRIDGE Medical Leaderboard - Few-Shot | 51.97 | Average Performance (%) | 92.6 |
| Open Portuguese LLM - ENEM | 78.66 | Accuracy (%) | 92.2 |
| Open Portuguese LLM - OAB Exams | 59.13 | Accuracy (%) | 91.7 |
| Open Portuguese LLM - BLUEX | 69.26 | Accuracy (%) | 91.5 |
| BRIDGE Medical Leaderboard | 43.66 | Average Performance (%) | 87 |
| MedPRESS - Safe Stance Adherence Rate | 32.2 | Safe-stance answers (%; turns 1-5, LLM judge) | 84.2 |
| BRIDGE Medical Leaderboard - Zero-Shot | 40.8 | Average Performance (%) | 83.3 |
| BRIDGE Medical Leaderboard - CoT | 38.2 | Average Performance (%) | 80.6 |
| MedPRESS - Conversation Failure Rate | 87.9 | Conversations with any unsafe agreement (%) | 68.4 |
| MedPRESS - Turn of Flip | 2.16 | Pressure turns withstood before first unsafe agreement (0-5) | 68.4 |
Interactive version: theaggregate.ai/model?slug=medgemma-27b-it · How It Works · Data refreshed daily, snapshot 2026-09-19.