DeepSeek V4 Pro (Reasoning, High Effort) — benchmark results

DeepSeek V4 Pro reasoning high-effort evaluation mode. Provider: DeepSeek. Released 2026-04-23. Access: Open.

Unified ELO 1853 ± 26, rank #100 of 1841 rated models, from 35 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AA GPQA Diamond90.51Accuracy (%)94.2
AA Omniscience - Science, Engineering & Mathematics45.3Accuracy (%)92.9
AA Omniscience - Humanities & Social Sciences43.6Accuracy (%)92.7
Artificial Analysis Intelligence Index43.11Intelligence Index92.3
AA Omniscience - Business37.5Accuracy (%)92.1
AA TAU-2 Bench94.15Accuracy (%)92
AA Humanity's Last Exam33.5Accuracy (%)91.8
AA CritPt10Accuracy (%)90.4
AA Omniscience - Law36.6Accuracy (%)90.3
AA-Omniscience Accuracy41.83Accuracy (%)89.5
AA Omniscience - Health36.8Accuracy (%)88.4
AA Terminal-Bench Hard41.67Accuracy (%)88.4

Interactive version: theaggregate.ai/model?slug=deepseek-v4-pro-reasoning-high-effort · How It Works · Data refreshed daily, snapshot 2026-07-25.