Gemma 4 12B (Non-reasoning): benchmark results
Provider: Google. Released 2026-06-03. Access: Open.
Unified ELO 1504 ± 1, rank #972 of 2032 rated models, from 67 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA-Omniscience Hallucination Rate | 73.41 | Hallucination Rate (%) | 61.9 |
| IHBench - Task Fulfillment | 50.5 | Win rate against GPT-4o Audio (%; share of 428 interruption | 56 |
| AA IFBench | 45.17 | Accuracy (%) | 51.7 |
| AA Omniscience - Software Engineering (SWE) - HTML | 30 | Accuracy (%) | 51.7 |
| AA Omniscience - Software Engineering (SWE) - Swift | 32 | Accuracy (%) | 51.2 |
| MSI-Bench - English | 22.6 | All-pass rate (%; all 576 English test cases over six multi- | 50 |
| MSI-Bench - English Speaker Authority | 42.7 | All-pass rate (%; the 96 speaker-authority cases, where an u | 50 |
| MSI-Bench - Mandarin Selective Disclosure | 27.1 | All-pass rate (%; the 96 selective-disclosure cases, where e | 50 |
| IHBench - Recovery Quality - Filler | 22 | Recovery pass rate (%; 60 backchannel fillers, where the tur | 48.1 |
| AA Terminal-Bench Hard | 11.36 | Accuracy (%) | 46.8 |
| AA Omniscience - Software Engineering (SWE) - Dart | 16 | Accuracy (%) | 43.7 |
| MSI-Bench - English Constraint Prioritization | 26 | All-pass rate (%; the 96 constraint-prioritization cases, wh | 42.9 |
Interactive version: theaggregate.ai/model?slug=gemma-4-12b-non-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-26.