DeepSeek V3: benchmark results
DeepSeek V3 open MoE model (671B total, 37B active). Provider: DeepSeek. Released 2024-12-26. Access: Open.
Unified ELO 1576 ± 1, rank #331 of 1392 rated models, from 342 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AIM-Bench (inventory manager) | 14.72 | BG Avg. C (self-reported) | 100 |
| BBH | 87.5 | Score (self-reported) | 100 |
| DITING | 5.16 | Score | 100 |
| DITING - Lexical Ambiguity | 5.52 | Score | 100 |
| DITING - Tense Consistency | 5.46 | Score | 100 |
| DITING - Terminology Localization | 5.12 | Score | 100 |
| Diagnosing LLM Arbitration Behavior over Pre-e | 23.5 | Margin (self-reported) | 100 |
| FinBen - FNS | 37.72 | Normalized Score | 100 |
| LLM Stats (DROP) | 91.6 | Score (%) | 100 |
| LongEmotion | 63.42 | Overall Score (%) | 100 |
| Open FinLLM Reasoning | 61.3 | Accuracy (%) | 100 |
| Open FinLLM Reasoning - DM-Complong | 42.33 | Accuracy (%) | 100 |
Interactive version: theaggregate.ai/model?slug=deepseek-v3 · How It Works · Data refreshed daily, snapshot 2026-09-05.