Gemma 4 31B (Reasoning) — benchmark results

Gemma 4 31B evaluated with reasoning enabled. Provider: Google. Released 2026-04-02. Access: Open.

Unified ELO 1699 ± 24, rank #269 of 1776 rated models, from 43 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AA IFBench75.58Accuracy (%)93.4
AA GPQA Diamond85.66Accuracy (%)85.1
AA Terminal-Bench Hard36.36Accuracy (%)83.3
AA Humanity's Last Exam22.66Accuracy (%)82.9
AA SciCode43.4Accuracy (%)82.9
AA Long Context Reasoning62Accuracy (%)75.3
AA CritPt1.43Accuracy (%)75.1
Artificial Analysis Intelligence Index29.35Intelligence Index74.4
PACT (Lechmazur)1557CMS Points74
AA MMMU-Pro73.41Accuracy (%)70.6
AA Omniscience - Science, Engineering & Mathematics32Accuracy (%)69.6
Tau3 Banking15.05Success Rate (%)66.5

Interactive version: theaggregate.ai/model?slug=gemma-4-31b-reasoning · How the rankings work · Data refreshed daily, snapshot 2026-07-22.