Gemma 4 E4B (IT) (Thinking): benchmark results

Provider: Google. Released 2026-04-02. Access: Open.

Unified ELO 1556 ± 1, rank #922 of 3078 rated models, from 73 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Horangi 4 - BFCL70.95Accuracy (%)98.8
FrameBench - Frame Identification - Japanese76.8Accuracy (%; Japanese FrameNet candidate frames)92
Swallow - Japanese MT-Bench - Roleplay71.9Judge Score (normalized, %)86.6
Swallow - Japanese MT-Bench - Humanities70.4Judge Score (normalized, %)85.1
Swallow - Japanese MT-Bench - Writing69.7Judge Score (normalized, %)85.1
Swallow - English MT-Bench - Coding79.3Judge Score (normalized, %)83.6
Swallow - English MT-Bench - Extraction82.1Judge Score (normalized, %)82.1
Swallow - Japanese MT-Bench - Coding79.9Judge Score (normalized, %)82.1
Swallow - English MT-Bench - Roleplay77.2Judge Score (normalized, %)80.6
Swallow - Post-trained Japanese - WMT20 En-Ja26.1BLEU80.6
Swallow - Japanese MT-Bench - Average75.8Judge Score (normalized, %)78.4
Horangi 4 - Ko-MT-Bench88.25Judge rating (1-10, x10)77.4

Interactive version: theaggregate.ai/model?slug=gemma-4-e4b-it-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.