DeepSeek V4 Pro (Max) — benchmark results
DeepSeek V4 Pro evaluated at the max reasoning-effort setting. Provider: DeepSeek. Released 2026-04-23. Access: Open.
Unified ELO 1873 ± 21, rank #84 of 1776 rated models, from 49 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| LLM Stats (CSimpleQA) | 84.4 | Score (%) | 100 |
| LLM Stats (MathArena Apex) | 90.2 | Score (%) | 100 |
| LLM Stats (CodeForces) | 100 | Score | 96.7 |
| CritPt | 12.9 | Accuracy (self-reported) | 96 |
| MathArena - ArXiv Math Jan 2026 | 73.91 | Accuracy (%) | 94.2 |
| OTIS Mock AIME 2024-25 | 96.67 | Accuracy (%) | 92.9 |
| Finance Agent v1.1 | 60.39 | Score (self-reported) | 90.9 |
| ZeroEval GPQA Diamond | 90.1 | GPQA Diamond Score | 90.3 |
| LLM Stats (HMMT Feb 26) | 95.2 | Score (%) | 90 |
| LLM2014 Logic 2026-04 | 70.49 | Median Score | 90 |
| AA-LCR | 66.3 | Score (self-reported) | 88.3 |
| MathArena - APEX Shortlist 2025 | 87.77 | Accuracy (%) | 86.1 |
Interactive version: theaggregate.ai/model?slug=deepseek-v4-pro-max · How the rankings work · Data refreshed daily, snapshot 2026-07-22.