Gemma 4 31B (IT) (Thinking): benchmark results
Provider: Google. Released 2026-04-02. Access: Open.
Unified ELO 1658 ± 1, rank #294 of 3078 rated models, from 83 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| FrameBench - Japanese | 99.1 | Accuracy (%; mean of five prompt templates) | 100 |
| Swallow - Post-trained English - Average | 87.4 | Average Score (%) | 100 |
| Swallow - Post-trained English - IFBench | 76.5 | Instruction-Level Strict Accuracy (%) | 100 |
| Swallow - Post-trained English - LiveCodeBench | 81 | Pass@1 (%) | 100 |
| Swallow - Post-trained Japanese - WMT20 En-Ja | 28.5 | BLEU | 100 |
| Horangi 4 - KoBALT-700 (Syntax) | 88 | Accuracy (%) | 98.6 |
| Swallow - Post-trained Japanese - M-IFEval-Ja | 89.8 | Instruction-Level Strict Accuracy (%) | 98.5 |
| Swallow - Post-trained Japanese - PolyMath High and Top | 68.4 | Accuracy (%) | 98.5 |
| Swallow - Japanese MT-Bench - Coding | 87.2 | Judge Score (normalized, %) | 97.8 |
| Horangi 4 - KoBBQ | 98 | Accuracy (%) | 97.1 |
| Swallow - Japanese MT-Bench - Reasoning | 84.4 | Judge Score (normalized, %) | 97 |
| Swallow - Post-trained English - MATH-500 | 99 | Accuracy (%) | 97 |
Interactive version: theaggregate.ai/model?slug=gemma-4-31b-it-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.