DeepSeek V3.1 (Thinking) — benchmark results
DeepSeek V3.1 evaluated with thinking enabled. Provider: DeepSeek. Released 2025-08-21. Access: Open.
Unified ELO 1714 ± 15, rank #252 of 1776 rated models, from 65 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| LLM2014 Code 2025-09 - C# | 8.7 | Score | 100 |
| LLM2014 Code 2025-09 - TypeScript | 9.38 | Score | 100 |
| LLM2014 Code 2025-09 - C++ | 6.19 | Score | 94.7 |
| AA MMLU-Pro | 85.06 | Accuracy (%) | 91.3 |
| LLM2014 Logic 2025-09 | 54.13 | Median Score | 91.1 |
| UGI Leaderboard | 50.59 | UGI Score | 90.5 |
| AA LiveCodeBench | 78.41 | Pass@1 (%) | 90.2 |
| UGI - Writing | 51.85 | Writing Score | 89.9 |
| AA AIME 2025 | 89.67 | Accuracy (%) | 89.8 |
| UGI - Natural Intelligence | 46.89 | NatInt Score | 89.6 |
| LLM2014 Code 2025-09 - Python | 7.36 | Score | 84.2 |
| LLM2014 Logic 2025-08 | 56.82 | Median Score | 84.1 |
Interactive version: theaggregate.ai/model?slug=deepseek-v3-1-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.