Gemma 4 31B — benchmark results
Google's flagship dense Gemma 4 31B open model. Provider: Google. Released 2026-04-02. Access: Open.
Unified ELO 1585 ± 15, rank #514 of 1776 rated models, from 402 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| BIG-Bench Extra Hard | 74.4 | Score (self-reported) | 100 |
| Codeforces ELO | 2150 | ELO (self-reported) | 100 |
| Momento | 78.26 | Pass@k (self-reported) | 100 |
| SEA-HELM | 75.16 | Mean Score (%) | 100 |
| TeleResilienceBench | 29.1 | Macro Average CFR (self-reported) | 100 |
| TuRTLe Spec-to-RTL (Icarus Verilog) | 81.51 | Aggregated Score (self-reported) | 100 |
| TuRTLe Spec-to-RTL (Verilator) | 79.83 | Aggregated Score (self-reported) | 100 |
| Kaggle FACTS Grounding | 80.67 | Score (%) | 97.4 |
| Agent Arena | 14.51 | Net Improvement (%) | 97.3 |
| Agent Arena - Bash Recovery | 33.53 | Bash Recovery (%) | 97.3 |
| ReasonScape R12 | 927.77 | ReasonScore | 95.5 |
| TuRTLe Code Completion (Icarus Verilog) | 82.57 | Aggregated Score (self-reported) | 95.3 |
Interactive version: theaggregate.ai/model?slug=gemma-4-31b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.