gemma-4-E4B-it — benchmark results

Google's Gemma 4 instruction-tuned on-device model (~4B effective params via per-layer embeddings) with text, image and audio input and a 128K context (April 2026). Provider: Google. Released 2026-04-02. Access: Open.

Unified ELO 1537 ± 8, rank #656 of 1776 rated models, from 311 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
EuroEval Italian Summarization - Ilpost SUM38.85Score (%)97.3
EuroEval Portuguese NLU - SST-2 PT86.93Sentiment classification Score (%)95.9
EuroEval Lithuanian NLU - Atsiliepimai49.2Sentiment classification Score (%)95
EuroEval Portuguese NLU - ScaLA PT48.86Linguistic acceptability Score (%)93.6
EuroEval Portuguese NLU63.05NLU Average Score (%)93.4
EuroEval Hungarian NLU - MultiWikiQA HU67.2Reading comprehension Score (%)92.6
EuroEval Danish Summarization - Nordjylland News37.41Score (%)91.8
EuroEval Icelandic NLU - Hotter and Colder Sentiment55.28Sentiment classification Score (%)91.6
EuroEval Bulgarian NLU - Cinexio56.94Sentiment classification Score (%)91
EuroEval Italian NLU - ScaLA IT47.43Linguistic acceptability Score (%)91
EuroEval Albanian Summarization - LR SUM SQ33.75Score (%)88.9
EuroEval Swedish Summarization - Swedn37.64Score (%)86.8

Interactive version: theaggregate.ai/model?slug=gemma-4-e4b-it · How the rankings work · Data refreshed daily, snapshot 2026-07-22.