DeepSeek-V3.2-Thinking-Speciale: benchmark results
Provider: DeepSeek. Access: Open.
Unified ELO 1866 ± 33, rank #90 of 2928 rated models, from 22 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| PM-LLM-Benchmark | 37.5 | Score | 94.6 |
| LisanBench | 0.33 | Mean Path Length / Current Maximum | 91.5 |
| AGI-Eval Community - Algorithmic Reasoning | 65.72 | Accuracy (%) | 87.1 |
| AGI-Eval Community - Learning (English) | 94.55 | Accuracy (%) | 84.1 |
| AGI-Eval Community - Subject Reasoning (Chinese) | 90.27 | Accuracy (%) | 83.7 |
| AGI-Eval Community - Mathematical Reasoning | 84.46 | Accuracy (%) | 83.6 |
| AGI-Eval Community - Objective Accuracy (Chinese) | 73.73 | Accuracy (%) | 83.3 |
| AGI-Eval Community - Subject Reasoning | 86.77 | Accuracy (%) | 82.9 |
| AGI-Eval Community - Objective Accuracy | 76.78 | Accuracy (%) | 81.4 |
| AGI-Eval Community - Subject Reasoning (English) | 85.21 | Accuracy (%) | 81.2 |
| AGI-Eval Community - Interaction (Chinese) | 82.51 | Accuracy (%) | 77.5 |
| AGI-Eval Community - Subject Knowledge | 86.91 | Accuracy (%) | 77.1 |
Interactive version: theaggregate.ai/model?slug=deepseek-v3-2-thinking-speciale · How It Works · Data refreshed daily, snapshot 2026-09-23.