DeepSeek LLM Chat 67B — benchmark results
Provider: DeepSeek. Released 2023-11-29. Access: Open.
Unified ELO 1441 ± 47, rank #1054 of 1839 rated models, from 13 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM Safety Anthropic Red Team | 99.4 | LM Evaluated Safety score (%) | 71.5 |
| HELM NaturalQuestions (Closed) | 41.21 | F1 (%) | 67.8 |
| HELM NaturalQuestions (Open) | 73.26 | F1 (%) | 63.3 |
| HELM Lite | 53.67 | Mean win rate (self-reported) | 57.1 |
| HELM (Stanford) | 48.84 | Mean Win Rate (%) | 50 |
| HELM WMT 2014 | 18.63 | BLEU-4 (%) | 43.3 |
| HELM Safety HarmBench | 64.9 | LM Evaluated Safety score (%) | 33.7 |
| HELM Safety | 87.3 | Mean score (self-reported) | 28.4 |
| HELM Safety BBQ | 86.2 | BBQ accuracy (%) | 24.4 |
| HELM Safety SimpleSafetyTests | 96.8 | LM Evaluated Safety score (%) | 23.8 |
| HELM AIR-Bench | 50.5 | Refusal Rate (%) | 18.6 |
| HELM Safety XSTest | 88.9 | LM Evaluated Safety score (%) | 11.6 |
Interactive version: theaggregate.ai/model?slug=deepseek-llm-chat-67b · How It Works · Data refreshed daily, snapshot 2026-08-05.