Gemma 3 4B (IT) — benchmark results
Google's instruction-tuned 4B Gemma 3 model (March 2025) taking image and text input, with 128K context and support for 140+ languages. Provider: Google. Released 2025-03-12. Access: Open.
Unified ELO 1443 ± 6, rank #1034 of 1776 rated models, from 1154 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| EVALITA - MAIA-GEN | 58.39 | CPS | 100 |
| MEGA-Bench Task - Face Swap | 71.4 | Task Score (%) | 100 |
| AGC-Bench - sdat | 0.67 | Dataset z-score | 97.5 |
| EuroEval Croatian NLU - MMS HR | 45.57 | Sentiment classification Score (%) | 95.3 |
| MEGA-Bench Task - Logical Reasoning 2D Folding | 42.9 | Task Score (%) | 95.3 |
| MEGA-Bench Task - OCR Table To CSV | 64.3 | Task Score (%) | 95.3 |
| EuroEval Dutch Summarization - Wiki Lingua NL | 35.47 | Score (%) | 94.7 |
| MEGA-Bench Task - Multiview Reasoning Camera Moving | 71.4 | Task Score (%) | 94.2 |
| MT-Bench PL - Roleplay | 9.45 | Judge Score (0-10) | 93.9 |
| MEGA-Bench Task - Poetry Shakespearean Sonnet | 20 | Task Score (%) | 93 |
| MEGA-Bench Task - AV Vehicle Multiview Counting | 26.7 | Task Score (%) | 91.9 |
| Ukrainian LLM - Long Flores UK CRH | 69.72 | Score (%) | 91.7 |
Interactive version: theaggregate.ai/model?slug=gemma-3-4b-it · How the rankings work · Data refreshed daily, snapshot 2026-07-22.