DeepSeek V4 Pro — benchmark results
DeepSeek's open 1600B V4 Pro model for high-end reasoning and general tasks. Provider: DeepSeek. Released 2026-04-23. Access: Open.
Unified ELO 1756 ± 15, rank #185 of 1776 rated models, from 298 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Beyond the All-in-One Agent | 62 | avg. (self-reported) | 100 |
| CAM-Bench | 19.67 | Pass@32 (self-reported) | 100 |
| Can Agents Price a Reaction? Evaluating LLMs o | 37 | C.10 Clean (self-reported) | 100 |
| GroupTravelBench | 10.47 | GU (Group Utility) (self-reported) | 100 |
| OpenCompass Agent - Tool Use | 55.3 | Score (%) | 100 |
| OpenCompass LLM - Agent | 55.3 | Score (%) | 100 |
| Vellum - LiveCodeBench | 93.5 | Pass@1 (%) | 100 |
| WDCD R3 Pressure Integrity | 200 | Score (%) | 100 |
| WorldCupBench - Brier Skill | 85.26 | 100 - Brier Total | 100 |
| ClawProBench | 64.38 | Final Score (self-reported) | 98.2 |
| IOI | 580.1 | Score (self-reported) | 98.1 |
| RewardBench 2 Focus | 93.13 | Accuracy (%) | 96.1 |
Interactive version: theaggregate.ai/model?slug=deepseek-v4-pro · How the rankings work · Data refreshed daily, snapshot 2026-07-22.