Gemma 4 E4B (IT) (Thinking): benchmark results
Provider: Google. Released 2026-04-02. Access: Open.
Unified ELO 1556 ± 1, rank #922 of 3078 rated models, from 73 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Horangi 4 - BFCL | 70.95 | Accuracy (%) | 98.8 |
| FrameBench - Frame Identification - Japanese | 76.8 | Accuracy (%; Japanese FrameNet candidate frames) | 92 |
| Swallow - Japanese MT-Bench - Roleplay | 71.9 | Judge Score (normalized, %) | 86.6 |
| Swallow - Japanese MT-Bench - Humanities | 70.4 | Judge Score (normalized, %) | 85.1 |
| Swallow - Japanese MT-Bench - Writing | 69.7 | Judge Score (normalized, %) | 85.1 |
| Swallow - English MT-Bench - Coding | 79.3 | Judge Score (normalized, %) | 83.6 |
| Swallow - English MT-Bench - Extraction | 82.1 | Judge Score (normalized, %) | 82.1 |
| Swallow - Japanese MT-Bench - Coding | 79.9 | Judge Score (normalized, %) | 82.1 |
| Swallow - English MT-Bench - Roleplay | 77.2 | Judge Score (normalized, %) | 80.6 |
| Swallow - Post-trained Japanese - WMT20 En-Ja | 26.1 | BLEU | 80.6 |
| Swallow - Japanese MT-Bench - Average | 75.8 | Judge Score (normalized, %) | 78.4 |
| Horangi 4 - Ko-MT-Bench | 88.25 | Judge rating (1-10, x10) | 77.4 |
Interactive version: theaggregate.ai/model?slug=gemma-4-e4b-it-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.