DeepSeek V2.5 — benchmark results

DeepSeek's open 236B MoE (21B active, 128K context) merging DeepSeek-V2-Chat and Coder-V2 into one general+coding model (September 2024). Provider: DeepSeek. Released 2024-09-05. Access: Open.

Unified ELO 1533 ± 30, rank #674 of 1776 rated models, from 36 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
LiveBench Math Comp53.12Score98.6
LiveBench Coding Completion50Score95.8
LiveBench Olympiad64.29Score91.7
LiveBench LCB Generation43Score90.3
VNTL Leaderboard71.14Accuracy (%)88.4
LiveBench AMPS Hard40Score87.5
LiveBench Story Generation77.17Score87.5
LiveBench Web Of Lies V258Score84.7
LiveBench Table Reformat64Score83.3
FullStackBench en58.65Score (self-reported)81.5
LiveBench Paraphrase68.72Score77.8
LiveBench Plot Unscrambling35.38Score77.8

Interactive version: theaggregate.ai/model?slug=deepseek-v2-5 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.