DeepSeek V2.5: benchmark results

DeepSeek's open 236B MoE (21B active, 128K context) merging DeepSeek-V2-Chat and Coder-V2 into one general+coding model (September 2024). Provider: DeepSeek. Released 2024-09-05. Access: Open.

Unified ELO 1497 ± 1, rank #697 of 1392 rated models, from 35 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
VNTL Leaderboard71.14Accuracy (%)88.4
FullStackBench en58.65Score (self-reported)76.7
WebApp1K83.38Pass@1 (%)69.7
Halluverse-M373.76Macro Accuracy (self-reported)69.2
BenchTable50Total Score (%)52.9
WebApp1K-React0.83pass@142.9
Deception Effectiveness (Lechmazur)0.61Deception Score41.2
Chatbot Arena (Text - Coding)1368Arena Score40.9
RepairBench25.1Plausible@1 (%)39.5
Chatbot Arena (Text - Chinese)1331Arena Score37.7
Chatbot Arena (Text - Hard Prompts)1321Arena Score37.2
Chatbot Arena (Text - Math)1287Arena Score35.5

Interactive version: theaggregate.ai/model?slug=deepseek-v2-5 · How It Works · Data refreshed daily, snapshot 2026-09-05.