DeepSeek V2 Chat: benchmark results
Provider: DeepSeek. Released 2024-05-06. Access: Open.
Unified ELO 1526 ± 1, rank #570 of 1392 rated models, from 11 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Open Portuguese LLM - ASSIN2 STS | 85.33 | Pearson Correlation (×100) | 97.6 |
| RepoQA | 83.4 | Score (self-reported) | 84.4 |
| WildBench | 48.21 | WB Score Task-Macro | 74.2 |
| BigCodeBench | 40.4 | Pass@1 (%) | 59.9 |
| Open Portuguese LLM - HateBR Offensive | 88.43 | F1 (%) | 51.2 |
| BenchTable | 49.2 | Total Score (%) | 51 |
| Open Portuguese LLM - Hate Speech | 72.72 | F1 (%) | 39 |
| SEAL - Chinese | 996 | Score | 35 |
| Open Portuguese LLM Leaderboard | 77.06 | Average Score (%) | 17.1 |
| Artificial Analysis Intelligence Index | 1 | Intelligence Index | 9.4 |
| Open Portuguese LLM - TweetSentBR | 68.35 | F1 (%) | 7.3 |
Interactive version: theaggregate.ai/model?slug=deepseek-v2-chat · How It Works · Data refreshed daily, snapshot 2026-09-05.