Gemma 4 12B — benchmark results

Provider: Google. Released 2026-04-02. Access: API.

Unified ELO 1438 ± 74, rank #1053 of 1776 rated models, from 18 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
BIG-Bench Extra Hard53Score (self-reported)86.7
SEA-HELM65.23Mean Score (%)81.4
LLM Stats (MathVision)79.7Score (%)61.3
ZeroEval GPQA Diamond78.8GPQA Diamond Score60.2
LLM Stats (MRCR v2 (8-needle))43.4Score (%)60
MedXpertQA48.7Score (self-reported)55.6
Codeforces ELO1659ELO (self-reported)50
LLM Stats (FLEURS)93.1Score (%)40
LLM Stats (MedXpertQA)48.7Score (%)36.4
BenchLM47.3Overall Score32
LLM Stats (MMMLU)83.4Score (%)30.4
CritPt0Accuracy (self-reported)30.3

Interactive version: theaggregate.ai/model?slug=gemma-4-12b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.