Gemma 4 E2B (IT) (Thinking): benchmark results
Provider: Google. Released 2026-04-02. Access: Open.
Unified ELO 1511 ± 1, rank #1301 of 3078 rated models, from 73 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Horangi 4 - BFCL | 69.32 | Accuracy (%) | 91.2 |
| Swallow - Japanese MT-Bench - Writing | 69.2 | Judge Score (normalized, %) | 82.1 |
| Swallow - Japanese MT-Bench - Humanities | 69.4 | Judge Score (normalized, %) | 80.6 |
| Swallow - Post-trained Japanese - WMT20 En-Ja | 25.6 | BLEU | 77.6 |
| FrameBench - Frame Identification - Japanese | 72.3 | Accuracy (%; Japanese FrameNet candidate frames) | 70 |
| Swallow - English MT-Bench - Extraction | 79.9 | Judge Score (normalized, %) | 68.7 |
| Swallow - English MT-Bench - Roleplay | 75.2 | Judge Score (normalized, %) | 68.7 |
| Swallow - Post-trained Japanese - WMT20 Ja-En | 22.7 | BLEU | 68.7 |
| Horangi 4 - Ko-HalluLens (Nonexistent Entities) | 79 | Refusal rate (%) | 63 |
| Swallow - English MT-Bench - Humanities | 72.2 | Judge Score (normalized, %) | 61.2 |
| Swallow - Japanese MT-Bench - Average | 72.6 | Judge Score (normalized, %) | 61.2 |
| Swallow - English MT-Bench - Coding | 76.4 | Judge Score (normalized, %) | 60.4 |
Interactive version: theaggregate.ai/model?slug=gemma-4-e2b-it-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.