DeepSeek V4 Pro Preview: benchmark results
Provider: DeepSeek. Released 2026-04-23. Access: Open.
Unified ELO 1683 ± 1, rank #81 of 1392 rated models, from 10 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| BenchmarkList ECI | 142.96 | Capability Index (ECI) | 91.3 |
| MineBench | 1551 | Elo Rating | 51.6 |
| NYT Connections Extended | 59.9 | Score (%) | 49 |
| Creative Writing (Lechmazur) | -0.2 | Mean Score | 46.8 |
| Multi-turn Debate (Lechmazur) | 1471.8 | Bradley-Terry Rating | 40 |
| Toolathlon | 55.9 | Score (self-reported) | 37 |
| NL2Repo | 38.5 | Score (self-reported) | 30.8 |
| ComplexConstraints | 28 | Score (%) | 28.2 |
| HANDBOOK.md Agents | 6.9 | Score (%) | 20.2 |
| GDPevo | 43.58 | base accuracy (%, no-evolution mode) | 0 |
Interactive version: theaggregate.ai/model?slug=deepseek-v4-pro-preview · How It Works · Data refreshed daily, snapshot 2026-09-05.