DeepSeek V2.5: benchmark results
DeepSeek's open 236B MoE (21B active, 128K context) merging DeepSeek-V2-Chat and Coder-V2 into one general+coding model (September 2024). Provider: DeepSeek. Released 2024-09-05. Access: Open.
Unified ELO 1497 ± 1, rank #697 of 1392 rated models, from 35 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| VNTL Leaderboard | 71.14 | Accuracy (%) | 88.4 |
| FullStackBench en | 58.65 | Score (self-reported) | 76.7 |
| WebApp1K | 83.38 | Pass@1 (%) | 69.7 |
| Halluverse-M3 | 73.76 | Macro Accuracy (self-reported) | 69.2 |
| BenchTable | 50 | Total Score (%) | 52.9 |
| WebApp1K-React | 0.83 | pass@1 | 42.9 |
| Deception Effectiveness (Lechmazur) | 0.61 | Deception Score | 41.2 |
| Chatbot Arena (Text - Coding) | 1368 | Arena Score | 40.9 |
| RepairBench | 25.1 | Plausible@1 (%) | 39.5 |
| Chatbot Arena (Text - Chinese) | 1331 | Arena Score | 37.7 |
| Chatbot Arena (Text - Hard Prompts) | 1321 | Arena Score | 37.2 |
| Chatbot Arena (Text - Math) | 1287 | Arena Score | 35.5 |
Interactive version: theaggregate.ai/model?slug=deepseek-v2-5 · How It Works · Data refreshed daily, snapshot 2026-09-05.