DeepSeek V4 Pro (Reasoning, High Effort) — benchmark results
DeepSeek V4 Pro reasoning high-effort evaluation mode. Provider: DeepSeek. Released 2026-04-23. Access: Open.
Unified ELO 1853 ± 26, rank #100 of 1841 rated models, from 35 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA GPQA Diamond | 90.51 | Accuracy (%) | 94.2 |
| AA Omniscience - Science, Engineering & Mathematics | 45.3 | Accuracy (%) | 92.9 |
| AA Omniscience - Humanities & Social Sciences | 43.6 | Accuracy (%) | 92.7 |
| Artificial Analysis Intelligence Index | 43.11 | Intelligence Index | 92.3 |
| AA Omniscience - Business | 37.5 | Accuracy (%) | 92.1 |
| AA TAU-2 Bench | 94.15 | Accuracy (%) | 92 |
| AA Humanity's Last Exam | 33.5 | Accuracy (%) | 91.8 |
| AA CritPt | 10 | Accuracy (%) | 90.4 |
| AA Omniscience - Law | 36.6 | Accuracy (%) | 90.3 |
| AA-Omniscience Accuracy | 41.83 | Accuracy (%) | 89.5 |
| AA Omniscience - Health | 36.8 | Accuracy (%) | 88.4 |
| AA Terminal-Bench Hard | 41.67 | Accuracy (%) | 88.4 |
Interactive version: theaggregate.ai/model?slug=deepseek-v4-pro-reasoning-high-effort · How It Works · Data refreshed daily, snapshot 2026-07-25.