DeepSeek Reasoner — benchmark results
DeepSeek's deepseek-reasoner API endpoint row, the thinking mode of the current DeepSeek flagship. Provider: DeepSeek. Released 2025-01-20. Access: API.
Unified ELO 1643 ± 34, rank #387 of 1776 rated models, from 17 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| BIRD-CRITIC | 33.67 | Score | 80 |
| OTIS Mock AIME 2024-25 | 87.82 | Accuracy (%) | 76.9 |
| Mizan LLM Leaderboard | 65.5 | Average Score (0-100) | 70.5 |
| Design Arena (Data Viz) | 1215 | Elo | 61.6 |
| Context-Bench Skills | 70.48 | Task Completion (%) | 52.4 |
| Design Arena (3D) | 1169 | Elo | 48.5 |
| Design Arena (Website) | 1177 | Elo | 45.5 |
| Design Arena (Game Dev) | 1157 | Elo | 40 |
| Design Arena (UI Components) | 1146 | Elo | 37.5 |
| HalluHard | 0.71 | Turn-1 Hallucination Rate | 29 |
| SimpleQA Verified | 27.5 | Accuracy (%) | 25 |
| Design Arena (SVG) | 1085 | Elo | 23.7 |
Interactive version: theaggregate.ai/model?slug=deepseek-reasoner · How the rankings work · Data refreshed daily, snapshot 2026-07-22.