DeepSeek V4: benchmark results

Provider: DeepSeek. Access: Open.

Unified ELO 1780 ± 28, rank #93 of 1607 rated models, from 15 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
BioMedHop (Direct Prompting) - Path-based Counting14.5Accuracy (%; exact count of the distinct diseases that satis100
BioMedHop (Direct Prompting) - Path-based Reasoning28.2Accuracy (%; mean of multiple-choice accuracy and alias-norm100
EnterpriseArena60Full Survival % (self-reported)100
WorldMemArena69.13QA Correct (%; share of answers a GPT-5.4-mini judge marks c100
BioMedHop (Direct Prompting)18.5Overall accuracy (%; mean of the four task-family scores, wh87.5
FrontierMath - Tiers 1-3 (v2)45.26Accuracy (%, 285 private v2 problems)64.1
BioMedHop (Direct Prompting) - Intersection Reasoning20.5Accuracy (%; mean of multiple-choice accuracy and alias-norm62.5
BioMedHop (Direct Prompting) - Entity Pair Matching10.8Accuracy (%; mean of multiple-choice accuracy and alias-norm56.2
OpenSkillEval4.3Overall avg. (self-reported)55.6
Are Agents Ready to Teach? A Multi-Stage Bench55.7Eq. pass (self-reported)50
SCOUT-450 (LLM Judge)87.1Accuracy (%; binary attack-or-benign verdicts on 450 held-ou50
PicoRV32 RTL-to-GDS (Claude Code, 350 MHz)13.33End-to-end design score (0-100) for the PicoRV32 RTL-to-GDS 33.3

Interactive version: theaggregate.ai/model?slug=deepseek-v4 · How It Works · Data refreshed daily, snapshot 2026-09-29.