DiffusionGemma 26B-A4B — benchmark results

Google DeepMind's open-weight text-diffusion model (June 2026) built on the Gemma 4 26B-A4B MoE backbone: 25.2B total / 3.8B active parameters, denoising 256-token canvases in parallel for ~4x faster generation. Provider: Google. Released 2026-06-10. Access: Open.

Unified ELO 1557 ± 21, rank #670 of 1839 rated models, from 41 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AA IFBench59.46Accuracy (%)71
AA Humanity's Last Exam10.19Accuracy (%)61.6
AA CritPt0.29Accuracy (%)57.9
AA Omniscience - Software Engineering (SWE) - Rust52Accuracy (%)54.1
MedXpertQA49Score (self-reported)52
LLM Stats (MathVision)70.5Score (%)51.6
Epoch AI - Scicode34.26Score51.4
ZeroEval GPQA Diamond73.2GPQA Diamond Score49.3
LLM Stats Score20.41LLM Stats Score (conservative rating)48.7
AA MMMU-Pro66.53Accuracy (%)48.5
AA Omniscience - Science, Engineering & Mathematics25.8Accuracy (%)48.4
AA Omniscience - Software Engineering (SWE) - Julia12Accuracy (%)47.3

Interactive version: theaggregate.ai/model?slug=diffusiongemma-26b-a4b · How It Works · Data refreshed daily, snapshot 2026-08-05.