Gemma 4 31B (Reasoning): benchmark results
Gemma 4 31B evaluated with reasoning enabled. Provider: Google. Released 2026-04-02. Access: Open.
Unified ELO 1603 ± 1, rank #426 of 1761 rated models, from 46 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA IFBench | 75.58 | Accuracy (%) | 93.4 |
| AA Terminal-Bench Hard | 36.36 | Accuracy (%) | 83.3 |
| AA GPQA Diamond | 85.66 | Accuracy (%) | 79.4 |
| AA Humanity's Last Exam | 23.63 | Accuracy (%) | 75.7 |
| PACT (Lechmazur) | 1557 | PACT Bilateral Rating | 74 |
| AA Omniscience - Software Engineering (SWE) - PHP | 30 | Accuracy (%) | 72.5 |
| AA Omniscience - Software Engineering (SWE) - Dart | 24 | Accuracy (%) | 72.2 |
| AA Omniscience - Software Engineering (SWE) - JavaScript | 33.64 | Accuracy (%) | 69.1 |
| Artificial Analysis Intelligence Index | 22.23 | Intelligence Index | 68.5 |
| AA Omniscience - Software Engineering (SWE) - Go | 22 | Accuracy (%) | 68.4 |
| AA CritPt | 1.43 | Accuracy (%) | 68.1 |
| AA Omniscience - Software Engineering (SWE) - C | 41 | Accuracy (%) | 66.4 |
Interactive version: theaggregate.ai/model?slug=gemma-4-31b-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-05.