DeepSeek V4 Pro (Max) — benchmark results

DeepSeek V4 Pro evaluated at the max reasoning-effort setting. Provider: DeepSeek. Released 2026-04-23. Access: Open.

Unified ELO 1873 ± 21, rank #84 of 1776 rated models, from 49 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
LLM Stats (CSimpleQA)84.4Score (%)100
LLM Stats (MathArena Apex)90.2Score (%)100
LLM Stats (CodeForces)100Score96.7
CritPt12.9Accuracy (self-reported)96
MathArena - ArXiv Math Jan 202673.91Accuracy (%)94.2
OTIS Mock AIME 2024-2596.67Accuracy (%)92.9
Finance Agent v1.160.39Score (self-reported)90.9
ZeroEval GPQA Diamond90.1GPQA Diamond Score90.3
LLM Stats (HMMT Feb 26)95.2Score (%)90
LLM2014 Logic 2026-0470.49Median Score90
AA-LCR66.3Score (self-reported)88.3
MathArena - APEX Shortlist 202587.77Accuracy (%)86.1

Interactive version: theaggregate.ai/model?slug=deepseek-v4-pro-max · How the rankings work · Data refreshed daily, snapshot 2026-07-22.