DeepSeek V4 Pro: benchmark results
DeepSeek's open 1600B V4 Pro model for high-end reasoning and general tasks. Provider: DeepSeek. Released 2026-04-23. Access: Open.
Unified ELO 1663 ± 1, rank #107 of 1392 rated models, from 435 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Beyond the All-in-One Agent | 62 | avg. (self-reported) | 100 |
| CAM-Bench | 19.67 | Pass@32 (self-reported) | 100 |
| Can Agents Price a Reaction? Evaluating LLMs o | 37 | C.10 Clean (self-reported) | 100 |
| GroupTravelBench | 10.47 | GU (Group Utility) (self-reported) | 100 |
| Vellum - LiveCodeBench | 93.5 | Pass@1 (%) | 100 |
| WorldCupBench - Brier Skill | 85.35 | 100 - Brier Total | 100 |
| MERA Code - ruCodeEval | 75.79 | pass@1 (%) | 98.5 |
| MERA Code - stRuCom | 37.45 | chrF (%) | 98.5 |
| ClawProBench | 64.38 | Final Score (self-reported) | 98.2 |
| RAI-Bench - RAG Robustness (HY Abstention) | 91 | Rate (%) | 97.1 |
| MERA Code | 0.51 | Total score | 96.9 |
| RewardBench 2 Focus | 93.13 | Accuracy (%) | 96.1 |
Interactive version: theaggregate.ai/model?slug=deepseek-v4-pro · How It Works · Data refreshed daily, snapshot 2026-09-05.