Gemma 4 12B (Reasoning): benchmark results
Provider: Google. Released 2026-06-03. Access: Open.
Unified ELO 1507 ± 1, rank #945 of 2032 rated models, from 71 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA IFBench | 73.54 | Accuracy (%) | 91.3 |
| AA Omniscience - Software Engineering (SWE) - HTML | 34.69 | Accuracy (%) | 67.7 |
| AA Humanity's Last Exam | 15.66 | Accuracy (%) | 63.9 |
| IHBench - Recovery Quality - Filler | 25 | Recovery pass rate (%; 60 backchannel fillers, where the tur | 61.5 |
| IHBench - Task Fulfillment | 51.1 | Win rate against GPT-4o Audio (%; share of 428 interruption | 60 |
| AA Omniscience - Software Engineering (SWE) - Rust | 50 | Accuracy (%) | 59 |
| AA GPQA Diamond | 75.25 | Accuracy (%) | 57.8 |
| AA Terminal-Bench Hard | 18.18 | Accuracy (%) | 57.3 |
| Artificial Analysis Intelligence Index | 14.18 | Intelligence Index | 57.3 |
| AA Long Context Reasoning | 63.67 | Accuracy (%) | 55.8 |
| AA Omniscience - Software Engineering (SWE) - Swift | 32 | Accuracy (%) | 51.2 |
| AA-Omniscience Hallucination Rate | 80.98 | Hallucination Rate (%) | 51.2 |
Interactive version: theaggregate.ai/model?slug=gemma-4-12b-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-26.