DeepSeek V3.1 (Non-reasoning) — benchmark results

DeepSeek V3.1 evaluated with reasoning disabled. Provider: DeepSeek. Released 2025-08-21. Access: Open.

Unified ELO 1627 ± 22, rank #419 of 1776 rated models, from 39 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
UGI Leaderboard51.4UGI Score92.3
UGI - Writing51.06Writing Score89.4
UGI - Natural Intelligence45.29NatInt Score89
AA MMLU-Pro83.28Accuracy (%)84
AA Omniscience - Software Engineering (SWE) - Kotlin36Accuracy (%)80.5
AA Omniscience - Software Engineering (SWE) - Go32Accuracy (%)76.8
AA Omniscience - Software Engineering (SWE) - Java24Accuracy (%)71.2
AA Omniscience - Humanities & Social Sciences26.2Accuracy (%)71
AA Omniscience - Software Engineering (SWE) - Julia24Accuracy (%)70.4
AA Omniscience - Software Engineering (SWE) - Python30.5Accuracy (%)70.4
AA Omniscience - Software Engineering (SWE) - R22Accuracy (%)69.4
AA Omniscience - Software Engineering (SWE) - C49Accuracy (%)68.8

Interactive version: theaggregate.ai/model?slug=deepseek-v3-1-non-reasoning · How the rankings work · Data refreshed daily, snapshot 2026-07-22.