DeepSeek V3 — benchmark results
DeepSeek V3 open MoE model (671B total, 37B active). Provider: DeepSeek. Released 2024-12-26. Access: Open.
Unified ELO 1594 ± 12, rank #488 of 1776 rated models, from 300 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AIM-Bench (inventory manager) | 14.72 | BG Avg. C (self-reported) | 100 |
| BBH | 87.5 | Score (self-reported) | 100 |
| DITING | 5.16 | Score | 100 |
| DITING - Lexical Ambiguity | 5.52 | Score | 100 |
| DITING - Tense Consistency | 5.46 | Score | 100 |
| DITING - Terminology Localization | 5.12 | Score | 100 |
| Diagnosing LLM Arbitration Behavior over Pre-e | 23.5 | Margin (self-reported) | 100 |
| FinBen - FNS | 37.72 | Normalized Score | 100 |
| LLM Stats (DROP) | 91.6 | Score (%) | 100 |
| LongEmotion | 63.42 | Overall Score (%) | 100 |
| MIRAGE | 26 | Violent completion rate, $C_1$ (%) (self-reported) | 100 |
| Open FinLLM Reasoning | 61.3 | Accuracy (%) | 100 |
Interactive version: theaggregate.ai/model?slug=deepseek-v3 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.