DeepSeek V4 Pro (Thinking): benchmark results
Provider: DeepSeek. Released 2026-04-24. Access: Open.
Unified ELO 1634, rank #231 of 2032 rated models, from 22 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| PortBench-QA - Max-Sharpe Allocation | 99.2 | Item score (%; long-only maximum-Sharpe weights for three or | 100 |
| VitaBench 2.0 | 47.2 | Avg@4 Full Context (self-reported) | 87.5 |
| PortBench-QA | 81.8 | Mean item score (%; mean of seven QA templates, 50 test ques | 77.8 |
| PortBench-QA - Position Sizing | 96.3 | Item score (%; fixed-fractional sizing from a drawdown limit | 72.2 |
| NetConfArena - Task Pass Rate | 73.8 | Instances with every test case satisfied (%) | 71.4 |
| ComboShoppingBench - Overall Success | 34.4 | Pass rate (%; all judged and rule-based checks) | 66.7 |
| ComboShoppingBench - Response Quality | 89.3 | Pass rate (%; LLM-judged) | 66.7 |
| ComboShoppingBench - Budget Compliance | 78.7 | Pass rate (%) | 61.9 |
| ComboShoppingBench - Claim Faithfulness | 80.1 | Pass rate (%; LLM-judged) | 61.9 |
| WebGameBench | 62.2 | Usable Rate (%; validly scored artifacts labelled Excellent | 61.5 |
| ComboShoppingBench - Coupon Optimality | 77.3 | Pass rate (%) | 57.1 |
| ComboShoppingBench - Semantic Requirements | 69.1 | Pass rate (%; LLM-judged) | 57.1 |
Interactive version: theaggregate.ai/model?slug=deepseek-v4-pro-thinking · How It Works · Data refreshed daily, snapshot 2026-09-26.