DeepSeek V3.1 Terminus (Reasoning) — benchmark results

DeepSeek V3.1 Terminus evaluated with reasoning enabled. Provider: DeepSeek. Released 2025-09-22. Access: Open.

Unified ELO 1722 ± 21, rank #237 of 1776 rated models, from 58 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
UGI Leaderboard51.26UGI Score91.9
AA LiveCodeBench79.79Pass@1 (%)91.8
AA MMLU-Pro85.11Accuracy (%)91.6
UGI - Writing54Writing Score91
UGI - Natural Intelligence49.47NatInt Score90.7
AA AIME 202589.67Accuracy (%)89.8
AA Long Context Reasoning65Accuracy (%)81.5
AA Omniscience - Humanities & Social Sciences29.6Accuracy (%)78.7
AA Omniscience - Health28.7Accuracy (%)77.9
AA CritPt1.71Accuracy (%)76.9
AA Omniscience - Science, Engineering & Mathematics34.9Accuracy (%)76
AA Omniscience - Business24.2Accuracy (%)75.9

Interactive version: theaggregate.ai/model?slug=deepseek-v3-1-terminus-reasoning · How the rankings work · Data refreshed daily, snapshot 2026-07-22.