DeepSeek V2 — benchmark results
DeepSeek's 236B open MoE (21B active) with Multi-head Latent Attention and a 128K context. Provider: DeepSeek. Released 2024-05-06. Access: Open.
Unified ELO 1442 ± 22, rank #1037 of 1776 rated models, from 11 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| ARC Challenge (AI2) | 92.2 | Accuracy (%) | 94.9 |
| WinoGrande | 86.3 | Accuracy (%) | 92.5 |
| HellaSwag | 87.1 | Accuracy (%) | 92.1 |
| Big-Bench Hard | 78.8 | Average (%) | 87.5 |
| PIQA | 83.9 | Accuracy (%) | 86.7 |
| MMLU | 78.4 | Accuracy (%) | 77 |
| MixEval | 51.7 | Score | 68.6 |
| TriviaQA | 80 | Accuracy (%) | 64.1 |
| LLM2014 Logic 2024-05 | 39.45 | Score (%) | 57.1 |
| LLM2014 Logic 2024-06 | 39.49 | Score (%) | 41.4 |
| Epoch AI - ECI | 124.57 | ECI Score | 21.4 |
Interactive version: theaggregate.ai/model?slug=deepseek-v2 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.