DeepSeek V3.2 (Thinking) — benchmark results

DeepSeek V3.2 evaluated with thinking enabled. Provider: DeepSeek. Released 2025-12-01. Access: Open.

Unified ELO 1721 ± 14, rank #239 of 1776 rated models, from 126 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
UGI Leaderboard57.78UGI Score98
AA LiveCodeBench86.24Pass@1 (%)97.1
AA MMLU-Pro86.21Accuracy (%)95.1
AA AIME 202592Accuracy (%)93.7
UGI - Writing57.28Writing Score92.3
UGI - Natural Intelligence48.11NatInt Score90.4
AA Omniscience - Software Engineering (SWE) - Swift68Accuracy (%)88.4
LLM2014 Logic 2025-1054.16Median Score87.8
BenchTable70.7Total Score (%)87.5
YapBench462.3YapIndex (lower is better)86.4
AA TAU-2 Bench90.64Accuracy (%)85.4
LLM2014 Code 2025-11 - TypeScript8.17Score84

Interactive version: theaggregate.ai/model?slug=deepseek-v3-2-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.