DeepSeek V4 Pro (Max): benchmark results
DeepSeek V4 Pro evaluated at the max reasoning-effort setting. Provider: DeepSeek. Released 2026-04-23. Access: Open.
Unified ELO 1675 ± 1, rank #162 of 1761 rated models, from 81 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| LLM Stats (CSimpleQA) | 84.4 | Score (%) | 100 |
| LLM Stats (MathArena Apex) | 90.2 | Score (%) | 100 |
| LLM Stats (CodeForces) | 100 | Score | 96.9 |
| MathArena - ArXiv Math Jan 2026 | 73.91 | Accuracy (%) | 94.2 |
| Vals AI Finance Agent | 60.39 | Accuracy (%) | 93.3 |
| OTIS Mock AIME 2024-25 | 96.67 | Accuracy (%) | 91.7 |
| HANDBOOK.md Agents | 26.9 | Score (%) | 91.2 |
| Vals AI LiveCodeBench | 87.48 | Accuracy (%) | 90.8 |
| LLM2014 Logic 2026-04 | 70.49 | Median Score | 90 |
| LLM Stats Score | 43.51 | LLM Stats Score (conservative rating) | 89.7 |
| ZeroEval GPQA Diamond | 90.1 | GPQA Diamond Score | 88.8 |
| MathArena - APEX Shortlist 2025 | 87.77 | Accuracy (%) | 86.5 |
Interactive version: theaggregate.ai/model?slug=deepseek-v4-pro-max · How It Works · Data refreshed daily, snapshot 2026-09-05.