DeepSeek V3 — benchmark results

DeepSeek V3 open MoE model (671B total, 37B active). Provider: DeepSeek. Released 2024-12-26. Access: Open.

Unified ELO 1594 ± 12, rank #488 of 1776 rated models, from 300 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AIM-Bench (inventory manager)14.72BG Avg. C (self-reported)100
BBH87.5Score (self-reported)100
DITING5.16Score100
DITING - Lexical Ambiguity5.52Score100
DITING - Tense Consistency5.46Score100
DITING - Terminology Localization5.12Score100
Diagnosing LLM Arbitration Behavior over Pre-e23.5Margin (self-reported)100
FinBen - FNS37.72Normalized Score100
LLM Stats (DROP)91.6Score (%)100
LongEmotion63.42Overall Score (%)100
MIRAGE26Violent completion rate, $C_1$ (%) (self-reported)100
Open FinLLM Reasoning61.3Accuracy (%)100

Interactive version: theaggregate.ai/model?slug=deepseek-v3 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.