gemma-4-E4B-it — benchmark results
Google's Gemma 4 instruction-tuned on-device model (~4B effective params via per-layer embeddings) with text, image and audio input and a 128K context (April 2026). Provider: Google. Released 2026-04-02. Access: Open.
Unified ELO 1537 ± 8, rank #656 of 1776 rated models, from 311 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| EuroEval Italian Summarization - Ilpost SUM | 38.85 | Score (%) | 97.3 |
| EuroEval Portuguese NLU - SST-2 PT | 86.93 | Sentiment classification Score (%) | 95.9 |
| EuroEval Lithuanian NLU - Atsiliepimai | 49.2 | Sentiment classification Score (%) | 95 |
| EuroEval Portuguese NLU - ScaLA PT | 48.86 | Linguistic acceptability Score (%) | 93.6 |
| EuroEval Portuguese NLU | 63.05 | NLU Average Score (%) | 93.4 |
| EuroEval Hungarian NLU - MultiWikiQA HU | 67.2 | Reading comprehension Score (%) | 92.6 |
| EuroEval Danish Summarization - Nordjylland News | 37.41 | Score (%) | 91.8 |
| EuroEval Icelandic NLU - Hotter and Colder Sentiment | 55.28 | Sentiment classification Score (%) | 91.6 |
| EuroEval Bulgarian NLU - Cinexio | 56.94 | Sentiment classification Score (%) | 91 |
| EuroEval Italian NLU - ScaLA IT | 47.43 | Linguistic acceptability Score (%) | 91 |
| EuroEval Albanian Summarization - LR SUM SQ | 33.75 | Score (%) | 88.9 |
| EuroEval Swedish Summarization - Swedn | 37.64 | Score (%) | 86.8 |
Interactive version: theaggregate.ai/model?slug=gemma-4-e4b-it · How the rankings work · Data refreshed daily, snapshot 2026-07-22.