DeepSeek V4 Flash (Reasoning, Max Effort) — benchmark results
DeepSeek V4 Flash reasoning max-effort evaluation mode. Provider: DeepSeek. Released 2026-04-23. Access: Open.
Unified ELO 1808 ± 21, rank #158 of 1806 rated models, from 37 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA IFBench | 79.18 | Accuracy (%) | 97.5 |
| AA TAU-2 Bench | 95.03 | Accuracy (%) | 93.8 |
| AA GPQA Diamond | 89.39 | Accuracy (%) | 90.9 |
| AA Humanity's Last Exam | 34.85 | Accuracy (%) | 90.6 |
| Artificial Analysis Intelligence Index | 42.12 | Intelligence Index | 90.1 |
| CritPt | 7.1 | Accuracy (self-reported) | 88.9 |
| AA Omniscience - Software Engineering (SWE) - Rust | 77.08 | Accuracy (%) | 88.6 |
| AA Omniscience - Health | 37.3 | Accuracy (%) | 87.1 |
| AA Omniscience - Science, Engineering & Mathematics | 42.8 | Accuracy (%) | 86.7 |
| AA CritPt | 7.14 | Accuracy (%) | 86.3 |
| AA Omniscience - Software Engineering (SWE) - C | 70 | Accuracy (%) | 85.6 |
| AA Omniscience - Software Engineering (SWE) - Go | 50 | Accuracy (%) | 84.8 |
Interactive version: theaggregate.ai/model?slug=deepseek-v4-flash-reasoning-max-effort · How It Works · Data refreshed daily, snapshot 2026-08-07.