Gemma 4 31B (Reasoning) — benchmark results
Gemma 4 31B evaluated with reasoning enabled. Provider: Google. Released 2026-04-02. Access: Open.
Unified ELO 1699 ± 24, rank #269 of 1776 rated models, from 43 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA IFBench | 75.58 | Accuracy (%) | 93.4 |
| AA GPQA Diamond | 85.66 | Accuracy (%) | 85.1 |
| AA Terminal-Bench Hard | 36.36 | Accuracy (%) | 83.3 |
| AA Humanity's Last Exam | 22.66 | Accuracy (%) | 82.9 |
| AA SciCode | 43.4 | Accuracy (%) | 82.9 |
| AA Long Context Reasoning | 62 | Accuracy (%) | 75.3 |
| AA CritPt | 1.43 | Accuracy (%) | 75.1 |
| Artificial Analysis Intelligence Index | 29.35 | Intelligence Index | 74.4 |
| PACT (Lechmazur) | 1557 | CMS Points | 74 |
| AA MMMU-Pro | 73.41 | Accuracy (%) | 70.6 |
| AA Omniscience - Science, Engineering & Mathematics | 32 | Accuracy (%) | 69.6 |
| Tau3 Banking | 15.05 | Success Rate (%) | 66.5 |
Interactive version: theaggregate.ai/model?slug=gemma-4-31b-reasoning · How the rankings work · Data refreshed daily, snapshot 2026-07-22.