DeepSeek V3.2 (High): benchmark results

DeepSeek V3.2 evaluated at the high reasoning-effort setting. Provider: DeepSeek. Released 2025-12-01. Access: Open.

Unified ELO 1650 ± 1, rank #346 of 3078 rated models, from 37 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Horangi 4 - Korean Hate Speech70Accuracy (%)94.7
Horangi 4 - HRM8K95.79Accuracy (%)88.5
Horangi 4 - GLP - Mathematical Reasoning95.89Score (%)87.5
Horangi 4 - Ko-AIME 202596Accuracy (%)81.7
Horangi 4 - ALT Average77.53Score (%)79.3
Horangi 4 - GLP - Specialized Knowledge53.28Score (%)76.9
Horangi 4 - Ko-HLE26.76Accuracy (%)76
Horangi 4 - BFCL68.22Accuracy (%)74.4
Horangi 4 - KMMLU-Pro79.8Accuracy (%)74
Horangi 4 - KoBALT-700 (Syntax)73Accuracy (%)73.6
Horangi 4 - Ko-ARC-AGI57.69Accuracy (%)73.1
SWE-bench Verified70Resolved (%)71.7

Interactive version: theaggregate.ai/model?slug=deepseek-v3-2-high · How It Works · Data refreshed daily, snapshot 2026-09-19.