DeepSeek V4 Pro (Thinking): benchmark results

Provider: DeepSeek. Released 2026-04-24. Access: Open.

Unified ELO 1634, rank #231 of 2032 rated models, from 22 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
PortBench-QA - Max-Sharpe Allocation99.2Item score (%; long-only maximum-Sharpe weights for three or100
VitaBench 2.047.2Avg@4 Full Context (self-reported)87.5
PortBench-QA81.8Mean item score (%; mean of seven QA templates, 50 test ques77.8
PortBench-QA - Position Sizing96.3Item score (%; fixed-fractional sizing from a drawdown limit72.2
NetConfArena - Task Pass Rate73.8Instances with every test case satisfied (%)71.4
ComboShoppingBench - Overall Success34.4Pass rate (%; all judged and rule-based checks)66.7
ComboShoppingBench - Response Quality89.3Pass rate (%; LLM-judged)66.7
ComboShoppingBench - Budget Compliance78.7Pass rate (%)61.9
ComboShoppingBench - Claim Faithfulness80.1Pass rate (%; LLM-judged)61.9
WebGameBench62.2Usable Rate (%; validly scored artifacts labelled Excellent 61.5
ComboShoppingBench - Coupon Optimality77.3Pass rate (%)57.1
ComboShoppingBench - Semantic Requirements69.1Pass rate (%; LLM-judged)57.1

Interactive version: theaggregate.ai/model?slug=deepseek-v4-pro-thinking · How It Works · Data refreshed daily, snapshot 2026-09-26.