gemma-4-E4B-it: benchmark results
Google's Gemma 4 instruction-tuned on-device model (~4B effective params via per-layer embeddings) with text, image and audio input and a 128K context (April 2026). Provider: Google. Released 2026-04-02. Access: Open.
Unified ELO 1535 ± 1, rank #510 of 1392 rated models, from 392 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| JMeetEval | 69 | Prompt-level strict accuracy (%) | 100 |
| JMeetEval - Instruction Level | 83.3 | Instruction-level strict accuracy (%) | 100 |
| EuroEval Italian Summarization - Ilpost SUM | 38.85 | Score (%) | 97.3 |
| EuroEval Portuguese NLU - SST-2 PT | 86.93 | Sentiment classification Score (%) | 95.9 |
| EuroEval Lithuanian NLU - Atsiliepimai | 49.2 | Sentiment classification Score (%) | 95 |
| EuroEval Portuguese NLU - ScaLA PT | 48.86 | Linguistic acceptability Score (%) | 93.6 |
| EuroEval Portuguese NLU | 63.05 | NLU Average Score (%) | 93.4 |
| EuroEval Hungarian NLU - MultiWikiQA HU | 67.2 | Reading comprehension Score (%) | 92.6 |
| EuroEval Danish Summarization - Nordjylland News | 37.41 | Score (%) | 91.8 |
| EuroEval Icelandic NLU - Hotter and Colder Sentiment | 55.28 | Sentiment classification Score (%) | 91.6 |
| EuroEval Bulgarian NLU - Cinexio | 56.94 | Sentiment classification Score (%) | 91 |
| EuroEval Italian NLU - ScaLA IT | 47.43 | Linguistic acceptability Score (%) | 91 |
Interactive version: theaggregate.ai/model?slug=gemma-4-e4b-it · How It Works · Data refreshed daily, snapshot 2026-09-05.