DeepSeek Coder V2: benchmark results
DeepSeek's open MoE code model (236B total/21B active, 128K context, 338 languages) that rivaled GPT-4 Turbo on coding (June 2024). Provider: DeepSeek. Released 2024-06-17. Access: Open.
Unified ELO 1509 ± 1, rank #655 of 1392 rated models, from 25 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| OlympicArena | 29.31 | Overall Accuracy (%) | 88.9 |
| BenchBench | 71.31 | Aggregate Score (%) | 74.3 |
| WildBench | 47.4 | WB Score Task-Macro | 69.4 |
| Web-Bench | 16.7 | Pass@1 (%) | 66 |
| Omni-MATH | 25.78 | Overall Accuracy (%) | 35.7 |
| BenchTable | 36 | Total Score (%) | 33.7 |
| SciCode | 21.2 | Subproblem Resolve Rate (%) | 33.3 |
| Chatbot Arena (Text - Coding) | 1342 | Arena Score | 33.1 |
| Chatbot Arena (Text - Math) | 1272 | Arena Score | 31.4 |
| Chatbot Arena (Text - Hard Prompts) | 1288 | Arena Score | 29.8 |
| OJBench | 8.24 | Score (self-reported) | 29.4 |
| Chatbot Arena (Text - Chinese) | 1291 | Arena Score | 27.2 |
Interactive version: theaggregate.ai/model?slug=deepseek-coder-v2 · How It Works · Data refreshed daily, snapshot 2026-09-05.