DeepSeek V4 Pro: benchmark results

DeepSeek's open 1600B V4 Pro model for high-end reasoning and general tasks. Provider: DeepSeek. Released 2026-04-23. Access: Open.

Unified ELO 1663 ± 1, rank #107 of 1392 rated models, from 435 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Beyond the All-in-One Agent62avg. (self-reported)100
CAM-Bench19.67Pass@32 (self-reported)100
Can Agents Price a Reaction? Evaluating LLMs o37C.10 Clean (self-reported)100
GroupTravelBench10.47GU (Group Utility) (self-reported)100
Vellum - LiveCodeBench93.5Pass@1 (%)100
WorldCupBench - Brier Skill85.35100 - Brier Total100
MERA Code - ruCodeEval75.79pass@1 (%)98.5
MERA Code - stRuCom37.45chrF (%)98.5
ClawProBench64.38Final Score (self-reported)98.2
RAI-Bench - RAG Robustness (HY Abstention)91Rate (%)97.1
MERA Code0.51Total score96.9
RewardBench 2 Focus93.13Accuracy (%)96.1

Interactive version: theaggregate.ai/model?slug=deepseek-v4-pro · How It Works · Data refreshed daily, snapshot 2026-09-05.