Gemma 4 12B (Non-reasoning): benchmark results

Provider: Google. Released 2026-06-03. Access: Open.

Unified ELO 1504 ± 1, rank #972 of 2032 rated models, from 67 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AA-Omniscience Hallucination Rate73.41Hallucination Rate (%)61.9
IHBench - Task Fulfillment50.5Win rate against GPT-4o Audio (%; share of 428 interruption 56
AA IFBench45.17Accuracy (%)51.7
AA Omniscience - Software Engineering (SWE) - HTML30Accuracy (%)51.7
AA Omniscience - Software Engineering (SWE) - Swift32Accuracy (%)51.2
MSI-Bench - English22.6All-pass rate (%; all 576 English test cases over six multi-50
MSI-Bench - English Speaker Authority42.7All-pass rate (%; the 96 speaker-authority cases, where an u50
MSI-Bench - Mandarin Selective Disclosure27.1All-pass rate (%; the 96 selective-disclosure cases, where e50
IHBench - Recovery Quality - Filler22Recovery pass rate (%; 60 backchannel fillers, where the tur48.1
AA Terminal-Bench Hard11.36Accuracy (%)46.8
AA Omniscience - Software Engineering (SWE) - Dart16Accuracy (%)43.7
MSI-Bench - English Constraint Prioritization26All-pass rate (%; the 96 constraint-prioritization cases, wh42.9

Interactive version: theaggregate.ai/model?slug=gemma-4-12b-non-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-26.