DeepSeek V3.2 (Thinking) — benchmark results
DeepSeek V3.2 evaluated with thinking enabled. Provider: DeepSeek. Released 2025-12-01. Access: Open.
Unified ELO 1721 ± 14, rank #239 of 1776 rated models, from 126 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| UGI Leaderboard | 57.78 | UGI Score | 98 |
| AA LiveCodeBench | 86.24 | Pass@1 (%) | 97.1 |
| AA MMLU-Pro | 86.21 | Accuracy (%) | 95.1 |
| AA AIME 2025 | 92 | Accuracy (%) | 93.7 |
| UGI - Writing | 57.28 | Writing Score | 92.3 |
| UGI - Natural Intelligence | 48.11 | NatInt Score | 90.4 |
| AA Omniscience - Software Engineering (SWE) - Swift | 68 | Accuracy (%) | 88.4 |
| LLM2014 Logic 2025-10 | 54.16 | Median Score | 87.8 |
| BenchTable | 70.7 | Total Score (%) | 87.5 |
| YapBench | 462.3 | YapIndex (lower is better) | 86.4 |
| AA TAU-2 Bench | 90.64 | Accuracy (%) | 85.4 |
| LLM2014 Code 2025-11 - TypeScript | 8.17 | Score | 84 |
Interactive version: theaggregate.ai/model?slug=deepseek-v3-2-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.