DeepSeek V4 Pro — benchmark results

DeepSeek's open 1600B V4 Pro model for high-end reasoning and general tasks. Provider: DeepSeek. Released 2026-04-23. Access: Open.

Unified ELO 1756 ± 15, rank #185 of 1776 rated models, from 298 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Beyond the All-in-One Agent62avg. (self-reported)100
CAM-Bench19.67Pass@32 (self-reported)100
Can Agents Price a Reaction? Evaluating LLMs o37C.10 Clean (self-reported)100
GroupTravelBench10.47GU (Group Utility) (self-reported)100
OpenCompass Agent - Tool Use55.3Score (%)100
OpenCompass LLM - Agent55.3Score (%)100
Vellum - LiveCodeBench93.5Pass@1 (%)100
WDCD R3 Pressure Integrity200Score (%)100
WorldCupBench - Brier Skill85.26100 - Brier Total100
ClawProBench64.38Final Score (self-reported)98.2
IOI580.1Score (self-reported)98.1
RewardBench 2 Focus93.13Accuracy (%)96.1

Interactive version: theaggregate.ai/model?slug=deepseek-v4-pro · How the rankings work · Data refreshed daily, snapshot 2026-07-22.