Gemma 4 31B (IT) (Thinking): benchmark results

Provider: Google. Released 2026-04-02. Access: Open.

Unified ELO 1658 ± 1, rank #294 of 3078 rated models, from 83 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
FrameBench - Japanese99.1Accuracy (%; mean of five prompt templates)100
Swallow - Post-trained English - Average87.4Average Score (%)100
Swallow - Post-trained English - IFBench76.5Instruction-Level Strict Accuracy (%)100
Swallow - Post-trained English - LiveCodeBench81Pass@1 (%)100
Swallow - Post-trained Japanese - WMT20 En-Ja28.5BLEU100
Horangi 4 - KoBALT-700 (Syntax)88Accuracy (%)98.6
Swallow - Post-trained Japanese - M-IFEval-Ja89.8Instruction-Level Strict Accuracy (%)98.5
Swallow - Post-trained Japanese - PolyMath High and Top68.4Accuracy (%)98.5
Swallow - Japanese MT-Bench - Coding87.2Judge Score (normalized, %)97.8
Horangi 4 - KoBBQ98Accuracy (%)97.1
Swallow - Japanese MT-Bench - Reasoning84.4Judge Score (normalized, %)97
Swallow - Post-trained English - MATH-50099Accuracy (%)97

Interactive version: theaggregate.ai/model?slug=gemma-4-31b-it-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.