gemma-4-26B-A4B-it (Thinking): benchmark results

Provider: Google. Released 2026-04-02. Access: Open.

Unified ELO 1631 ± 1, rank #441 of 3078 rated models, from 77 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Swallow - Post-trained English - IFBench72.7Instruction-Level Strict Accuracy (%)98.5
Swallow - Post-trained English - LiveCodeBench75.5Pass@1 (%)98.5
Swallow - Post-trained Japanese - JHumanEval95.9Pass@1 (%)98.5
Swallow - Post-trained Japanese - WMT20 Ja-En26.1BLEU98.5
Swallow - Post-trained Japanese - WMT20 En-Ja28BLEU97
Swallow - Japanese MT-Bench - Extraction77.2Judge Score (normalized, %)95.5
Swallow - Post-trained English - Average84.1Average Score (%)95.5
Swallow - Post-trained Japanese - M-IFEval-Ja86.7Instruction-Level Strict Accuracy (%)95.5
Swallow - Post-trained Japanese - PolyMath High and Top60Accuracy (%)95.5
Swallow - English MT-Bench - Roleplay81.5Judge Score (normalized, %)94
Swallow - Japanese MT-Bench - Coding84Judge Score (normalized, %)94
Swallow - Japanese MT-Bench - Roleplay75.3Judge Score (normalized, %)94

Interactive version: theaggregate.ai/model?slug=gemma-4-26b-a4b-it-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.