DeepSeek V4: benchmark results
Provider: DeepSeek. Access: Open.
Unified ELO 1780 ± 28, rank #93 of 1607 rated models, from 15 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| BioMedHop (Direct Prompting) - Path-based Counting | 14.5 | Accuracy (%; exact count of the distinct diseases that satis | 100 |
| BioMedHop (Direct Prompting) - Path-based Reasoning | 28.2 | Accuracy (%; mean of multiple-choice accuracy and alias-norm | 100 |
| EnterpriseArena | 60 | Full Survival % (self-reported) | 100 |
| WorldMemArena | 69.13 | QA Correct (%; share of answers a GPT-5.4-mini judge marks c | 100 |
| BioMedHop (Direct Prompting) | 18.5 | Overall accuracy (%; mean of the four task-family scores, wh | 87.5 |
| FrontierMath - Tiers 1-3 (v2) | 45.26 | Accuracy (%, 285 private v2 problems) | 64.1 |
| BioMedHop (Direct Prompting) - Intersection Reasoning | 20.5 | Accuracy (%; mean of multiple-choice accuracy and alias-norm | 62.5 |
| BioMedHop (Direct Prompting) - Entity Pair Matching | 10.8 | Accuracy (%; mean of multiple-choice accuracy and alias-norm | 56.2 |
| OpenSkillEval | 4.3 | Overall avg. (self-reported) | 55.6 |
| Are Agents Ready to Teach? A Multi-Stage Bench | 55.7 | Eq. pass (self-reported) | 50 |
| SCOUT-450 (LLM Judge) | 87.1 | Accuracy (%; binary attack-or-benign verdicts on 450 held-ou | 50 |
| PicoRV32 RTL-to-GDS (Claude Code, 350 MHz) | 13.33 | End-to-end design score (0-100) for the PicoRV32 RTL-to-GDS | 33.3 |
Interactive version: theaggregate.ai/model?slug=deepseek-v4 · How It Works · Data refreshed daily, snapshot 2026-09-29.