DiffusionGemma 26B-A4B — benchmark results
Google DeepMind's open-weight text-diffusion model (June 2026) built on the Gemma 4 26B-A4B MoE backbone: 25.2B total / 3.8B active parameters, denoising 256-token canvases in parallel for ~4x faster generation. Provider: Google. Released 2026-06-10. Access: Open.
Unified ELO 1557 ± 21, rank #670 of 1839 rated models, from 41 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA IFBench | 59.46 | Accuracy (%) | 71 |
| AA Humanity's Last Exam | 10.19 | Accuracy (%) | 61.6 |
| AA CritPt | 0.29 | Accuracy (%) | 57.9 |
| AA Omniscience - Software Engineering (SWE) - Rust | 52 | Accuracy (%) | 54.1 |
| MedXpertQA | 49 | Score (self-reported) | 52 |
| LLM Stats (MathVision) | 70.5 | Score (%) | 51.6 |
| Epoch AI - Scicode | 34.26 | Score | 51.4 |
| ZeroEval GPQA Diamond | 73.2 | GPQA Diamond Score | 49.3 |
| LLM Stats Score | 20.41 | LLM Stats Score (conservative rating) | 48.7 |
| AA MMMU-Pro | 66.53 | Accuracy (%) | 48.5 |
| AA Omniscience - Science, Engineering & Mathematics | 25.8 | Accuracy (%) | 48.4 |
| AA Omniscience - Software Engineering (SWE) - Julia | 12 | Accuracy (%) | 47.3 |
Interactive version: theaggregate.ai/model?slug=diffusiongemma-26b-a4b · How It Works · Data refreshed daily, snapshot 2026-08-05.