DeepSeek V4 Pro (High): benchmark results
DeepSeek V4 Pro evaluated at the high reasoning-effort setting. Provider: DeepSeek. Released 2026-04-23. Access: Open.
Unified ELO 1667 ± 1, rank #195 of 1761 rated models, from 14 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| OTIS Mock AIME 2024-25 | 95.56 | Accuracy (%) | 88.7 |
| LisanBench | 0.21 | Mean Path Length / Current Maximum | 85 |
| ALE-Bench | 1006.08 | Performance (Self-Refine x1) (self-reported) | 72.7 |
| OckBench | 84 | Accuracy (%) | 71.8 |
| FutureEval | 8.73 | Unified Forecasting Score | 69.6 |
| Creative Writing (Lechmazur) | 0.8 | Mean Score | 68.1 |
| Epoch AI - Critpt | 10 | Score | 66.3 |
| WebDev Arena | 1463.8 | Arena Score | 57.3 |
| WeirdML | 46.53 | Average Score | 54.1 |
| Chess Puzzles (Epoch AI) | 13 | Accuracy (%) | 47.5 |
| Surface Evolver Bench Pass Rate | 25 | Pass Rate (%) | 46 |
| SWE-rebench | 41.43 | Resolved (%) | 45.8 |
Interactive version: theaggregate.ai/model?slug=deepseek-v4-pro-high · How It Works · Data refreshed daily, snapshot 2026-09-05.