Gemma 4 E2B (IT) (Thinking): benchmark results

Provider: Google. Released 2026-04-02. Access: Open.

Unified ELO 1511 ± 1, rank #1301 of 3078 rated models, from 73 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Horangi 4 - BFCL69.32Accuracy (%)91.2
Swallow - Japanese MT-Bench - Writing69.2Judge Score (normalized, %)82.1
Swallow - Japanese MT-Bench - Humanities69.4Judge Score (normalized, %)80.6
Swallow - Post-trained Japanese - WMT20 En-Ja25.6BLEU77.6
FrameBench - Frame Identification - Japanese72.3Accuracy (%; Japanese FrameNet candidate frames)70
Swallow - English MT-Bench - Extraction79.9Judge Score (normalized, %)68.7
Swallow - English MT-Bench - Roleplay75.2Judge Score (normalized, %)68.7
Swallow - Post-trained Japanese - WMT20 Ja-En22.7BLEU68.7
Horangi 4 - Ko-HalluLens (Nonexistent Entities)79Refusal rate (%)63
Swallow - English MT-Bench - Humanities72.2Judge Score (normalized, %)61.2
Swallow - Japanese MT-Bench - Average72.6Judge Score (normalized, %)61.2
Swallow - English MT-Bench - Coding76.4Judge Score (normalized, %)60.4

Interactive version: theaggregate.ai/model?slug=gemma-4-e2b-it-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.