deepseek-llm-67B-chat — benchmark results
DeepSeek's first-gen 67B chat model, instruction-tuned from a base trained on 2T English/Chinese tokens; open weights allowing commercial use. Provider: DeepSeek. Released 2023-11-29. Access: Open.
Unified ELO 1448 ± 45, rank #1008 of 1776 rated models, from 12 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Open LLM Leaderboard - MuSR | 23.93 | Score | 98.7 |
| ChineseSafe Benchmark | 76.76 | Accuracy (%) | 98.1 |
| InfiBench | 57.41 | Score (%) | 89.5 |
| Open LLM Leaderboard - GPQA | 8.84 | Score | 72.5 |
| Open LLM Leaderboard - MMLU-Pro | 32.71 | Score | 72.3 |
| Open LLM Leaderboard - IFEval | 55.87 | Score | 67.7 |
| Open LLM Leaderboard - BBH | 33.23 | Score | 64.9 |
| BenchBench | 57.35 | Aggregate Score (%) | 60.3 |
| AlpacaEval 2.0 | 17.84 | LC Win Rate (%) | 45.9 |
| Open LLM Leaderboard - MATH Level 5 | 9.29 | Score | 45.1 |
| Chatbot Arena (Text) | 1184 | Elo | 16.6 |
| MATH Level 5 | 6.39 | Accuracy (%) | 4.6 |
Interactive version: theaggregate.ai/model?slug=deepseek-llm-67b-chat · How the rankings work · Data refreshed daily, snapshot 2026-07-22.