DeepSeek V3.1 (Thinking) — benchmark results

DeepSeek V3.1 evaluated with thinking enabled. Provider: DeepSeek. Released 2025-08-21. Access: Open.

Unified ELO 1714 ± 15, rank #252 of 1776 rated models, from 65 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
LLM2014 Code 2025-09 - C#8.7Score100
LLM2014 Code 2025-09 - TypeScript9.38Score100
LLM2014 Code 2025-09 - C++6.19Score94.7
AA MMLU-Pro85.06Accuracy (%)91.3
LLM2014 Logic 2025-0954.13Median Score91.1
UGI Leaderboard50.59UGI Score90.5
AA LiveCodeBench78.41Pass@1 (%)90.2
UGI - Writing51.85Writing Score89.9
AA AIME 202589.67Accuracy (%)89.8
UGI - Natural Intelligence46.89NatInt Score89.6
LLM2014 Code 2025-09 - Python7.36Score84.2
LLM2014 Logic 2025-0856.82Median Score84.1

Interactive version: theaggregate.ai/model?slug=deepseek-v3-1-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.