DeepSeek V4 Pro (Reasoning, Max Effort) — benchmark results
DeepSeek V4 Pro reasoning max-effort evaluation mode. Provider: DeepSeek. Released 2026-04-23. Access: Open.
Unified ELO 1832 ± 13, rank #112 of 1776 rated models, from 77 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Qwen3.7 Launch - CritPT | 12.9 | Score (%) | 100 |
| Qwen3.7 Launch - LiveCodeBench | 93.5 | Score (%) | 100 |
| Qwen3.7 Launch - Vitabench | 51.9 | Score (%) | 100 |
| AA TAU-2 Bench | 96.2 | Accuracy (%) | 96.8 |
| AA IFBench | 76.46 | Accuracy (%) | 95.5 |
| AA Omniscience - Humanities & Social Sciences | 47.7 | Accuracy (%) | 95.5 |
| AA Omniscience - Science, Engineering & Mathematics | 45.5 | Accuracy (%) | 94.2 |
| Artificial Analysis Intelligence Index | 44.27 | Intelligence Index | 94.2 |
| AA Omniscience - Health | 40.6 | Accuracy (%) | 94 |
| AA Humanity's Last Exam | 35.87 | Accuracy (%) | 93.7 |
| AA CritPt | 12.86 | Accuracy (%) | 93.4 |
| AA Terminal-Bench Hard | 46.21 | Accuracy (%) | 92.9 |
Interactive version: theaggregate.ai/model?slug=deepseek-v4-pro-reasoning-max-effort · How the rankings work · Data refreshed daily, snapshot 2026-07-22.