Gemma 3 4B (IT) — benchmark results

Google's instruction-tuned 4B Gemma 3 model (March 2025) taking image and text input, with 128K context and support for 140+ languages. Provider: Google. Released 2025-03-12. Access: Open.

Unified ELO 1443 ± 6, rank #1034 of 1776 rated models, from 1154 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
EVALITA - MAIA-GEN58.39CPS100
MEGA-Bench Task - Face Swap71.4Task Score (%)100
AGC-Bench - sdat0.67Dataset z-score97.5
EuroEval Croatian NLU - MMS HR45.57Sentiment classification Score (%)95.3
MEGA-Bench Task - Logical Reasoning 2D Folding42.9Task Score (%)95.3
MEGA-Bench Task - OCR Table To CSV64.3Task Score (%)95.3
EuroEval Dutch Summarization - Wiki Lingua NL35.47Score (%)94.7
MEGA-Bench Task - Multiview Reasoning Camera Moving71.4Task Score (%)94.2
MT-Bench PL - Roleplay9.45Judge Score (0-10)93.9
MEGA-Bench Task - Poetry Shakespearean Sonnet20Task Score (%)93
MEGA-Bench Task - AV Vehicle Multiview Counting26.7Task Score (%)91.9
Ukrainian LLM - Long Flores UK CRH69.72Score (%)91.7

Interactive version: theaggregate.ai/model?slug=gemma-3-4b-it · How the rankings work · Data refreshed daily, snapshot 2026-07-22.