gemma-4-26B-A4B-it (Thinking): benchmark results
Provider: Google. Released 2026-04-02. Access: Open.
Unified ELO 1631 ± 1, rank #441 of 3078 rated models, from 77 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Swallow - Post-trained English - IFBench | 72.7 | Instruction-Level Strict Accuracy (%) | 98.5 |
| Swallow - Post-trained English - LiveCodeBench | 75.5 | Pass@1 (%) | 98.5 |
| Swallow - Post-trained Japanese - JHumanEval | 95.9 | Pass@1 (%) | 98.5 |
| Swallow - Post-trained Japanese - WMT20 Ja-En | 26.1 | BLEU | 98.5 |
| Swallow - Post-trained Japanese - WMT20 En-Ja | 28 | BLEU | 97 |
| Swallow - Japanese MT-Bench - Extraction | 77.2 | Judge Score (normalized, %) | 95.5 |
| Swallow - Post-trained English - Average | 84.1 | Average Score (%) | 95.5 |
| Swallow - Post-trained Japanese - M-IFEval-Ja | 86.7 | Instruction-Level Strict Accuracy (%) | 95.5 |
| Swallow - Post-trained Japanese - PolyMath High and Top | 60 | Accuracy (%) | 95.5 |
| Swallow - English MT-Bench - Roleplay | 81.5 | Judge Score (normalized, %) | 94 |
| Swallow - Japanese MT-Bench - Coding | 84 | Judge Score (normalized, %) | 94 |
| Swallow - Japanese MT-Bench - Roleplay | 75.3 | Judge Score (normalized, %) | 94 |
Interactive version: theaggregate.ai/model?slug=gemma-4-26b-a4b-it-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.