DeepSeek V3.1 Terminus — benchmark results

DeepSeek's V3.1 Terminus update with improved agentic tool use and language consistency. Provider: DeepSeek. Released 2025-09-22. Access: Open.

Unified ELO 1662 ± 12, rank #340 of 1776 rated models, from 136 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
UGI Leaderboard57.58UGI Score97.9
AGC-Bench - ss_gen1.34Dataset z-score97.5
TuRTLe - Verilator Syntax95.25Average Score (%)95.3
AGC-Bench - conceptual_design1.3Dataset z-score93.9
TuRTLe - Icarus Syntax93.7Average Score (%)93
AGC-Bench - tinystories1.23Dataset z-score92.7
WebCoderBench - Performance98.19Score (%)92.3
AGC-Bench - puntuguese0.94Dataset z-score91.5
AGC-Bench - tinyfabulist1.1Dataset z-score91.5
AGC-Bench - outline_to_story1.14Dataset z-score91.4
AGC-Bench - showerthoughts1Dataset z-score91.2
TuRTLe - Verilator Power71.45Average Score (%)90.7

Interactive version: theaggregate.ai/model?slug=deepseek-v3-1-terminus · How the rankings work · Data refreshed daily, snapshot 2026-07-22.