DeepSeek V2.5 — benchmark results
DeepSeek's open 236B MoE (21B active, 128K context) merging DeepSeek-V2-Chat and Coder-V2 into one general+coding model (September 2024). Provider: DeepSeek. Released 2024-09-05. Access: Open.
Unified ELO 1533 ± 30, rank #674 of 1776 rated models, from 36 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| LiveBench Math Comp | 53.12 | Score | 98.6 |
| LiveBench Coding Completion | 50 | Score | 95.8 |
| LiveBench Olympiad | 64.29 | Score | 91.7 |
| LiveBench LCB Generation | 43 | Score | 90.3 |
| VNTL Leaderboard | 71.14 | Accuracy (%) | 88.4 |
| LiveBench AMPS Hard | 40 | Score | 87.5 |
| LiveBench Story Generation | 77.17 | Score | 87.5 |
| LiveBench Web Of Lies V2 | 58 | Score | 84.7 |
| LiveBench Table Reformat | 64 | Score | 83.3 |
| FullStackBench en | 58.65 | Score (self-reported) | 81.5 |
| LiveBench Paraphrase | 68.72 | Score | 77.8 |
| LiveBench Plot Unscrambling | 35.38 | Score | 77.8 |
Interactive version: theaggregate.ai/model?slug=deepseek-v2-5 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.