Gemma 4 12B (Reasoning): benchmark results

Provider: Google. Released 2026-06-03. Access: Open.

Unified ELO 1507 ± 1, rank #945 of 2032 rated models, from 71 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AA IFBench73.54Accuracy (%)91.3
AA Omniscience - Software Engineering (SWE) - HTML34.69Accuracy (%)67.7
AA Humanity's Last Exam15.66Accuracy (%)63.9
IHBench - Recovery Quality - Filler25Recovery pass rate (%; 60 backchannel fillers, where the tur61.5
IHBench - Task Fulfillment51.1Win rate against GPT-4o Audio (%; share of 428 interruption 60
AA Omniscience - Software Engineering (SWE) - Rust50Accuracy (%)59
AA GPQA Diamond75.25Accuracy (%)57.8
AA Terminal-Bench Hard18.18Accuracy (%)57.3
Artificial Analysis Intelligence Index14.18Intelligence Index57.3
AA Long Context Reasoning63.67Accuracy (%)55.8
AA Omniscience - Software Engineering (SWE) - Swift32Accuracy (%)51.2
AA-Omniscience Hallucination Rate80.98Hallucination Rate (%)51.2

Interactive version: theaggregate.ai/model?slug=gemma-4-12b-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-26.