Gemma 4 31B: benchmark results

Google's flagship dense Gemma 4 31B open model. Provider: Google. Released 2026-04-02. Access: Open.

Unified ELO 1607 ± 1, rank #243 of 1392 rated models, from 441 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Codeforces ELO2150ELO (self-reported)100
SEA-HELM75.19Mean Score (%)100
TeleResilienceBench29.1Macro Average CFR (self-reported)100
Kaggle FACTS Grounding80.67Score (%)97.8
RAI-Bench - Refusal Rate (JD)100Rate (%)97.4
ReasonScape R12927.77ReasonScore95.6
P3B398LLM pt-PT (↑) (self-reported)94.7
RAI-Bench - Refusal Rate (Career)100Rate (%)93.5
Fibble4 Arena8.33Win Rate (%)92.9
MathVision85.6Overall Accuracy (%)92.5
ALEM (Multi-Agent Coordination)10.1Total Achievement % (medium)91.7
Wordle Arena91.67Win Rate (%)91.7

Interactive version: theaggregate.ai/model?slug=gemma-4-31b · How It Works · Data refreshed daily, snapshot 2026-09-05.