Gemma 4 31B (IT) (Non-reasoning): benchmark results
Provider: Google. Released 2026-04-02. Access: Open.
Unified ELO 1572 ± 1, rank #821 of 3078 rated models, from 30 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| CPTU Bench | 4.32 | Average Score (1-5) | 96 |
| FrameBench - Japanese | 97.6 | Accuracy (%; mean of five prompt templates) | 92 |
| MERA v2 - Enantiosemy | 55.8 | Score (%) | 81.6 |
| MERA v2 - Humor | 36.4 | Score (%) | 81.6 |
| MERA v2 - Reasoning | 62.4 | Score (%) | 81.6 |
| MERA v2 - NewReasoning | 68.9 | Score (%) | 78.9 |
| MERA v2 - SAGE | 68 | Score (%) | 78.9 |
| FrameBench - Frame Identification - English | 81 | Accuracy (%; FrameNet candidate frames) | 76.3 |
| MERA v2 - Culture Specific | 35.9 | Score (%) | 76.3 |
| FrameBench - English | 97.5 | Accuracy (%; mean of five prompt templates) | 73.7 |
| MERA v2 - Characters | 57.5 | Score (%) | 73.7 |
| MERA v2 - RUBIN | 63.7 | Score (%) | 73.7 |
Interactive version: theaggregate.ai/model?slug=gemma-4-31b-it-non-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-19.