DeepSeek V4 Pro (Max): benchmark results

DeepSeek V4 Pro evaluated at the max reasoning-effort setting. Provider: DeepSeek. Released 2026-04-23. Access: Open.

Unified ELO 1675 ± 1, rank #162 of 1761 rated models, from 81 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
LLM Stats (CSimpleQA)84.4Score (%)100
LLM Stats (MathArena Apex)90.2Score (%)100
LLM Stats (CodeForces)100Score96.9
MathArena - ArXiv Math Jan 202673.91Accuracy (%)94.2
Vals AI Finance Agent60.39Accuracy (%)93.3
OTIS Mock AIME 2024-2596.67Accuracy (%)91.7
HANDBOOK.md Agents26.9Score (%)91.2
Vals AI LiveCodeBench87.48Accuracy (%)90.8
LLM2014 Logic 2026-0470.49Median Score90
LLM Stats Score43.51LLM Stats Score (conservative rating)89.7
ZeroEval GPQA Diamond90.1GPQA Diamond Score88.8
MathArena - APEX Shortlist 202587.77Accuracy (%)86.5

Interactive version: theaggregate.ai/model?slug=deepseek-v4-pro-max · How It Works · Data refreshed daily, snapshot 2026-09-05.