DeepSeek Reasoner — benchmark results

DeepSeek's deepseek-reasoner API endpoint row, the thinking mode of the current DeepSeek flagship. Provider: DeepSeek. Released 2025-01-20. Access: API.

Unified ELO 1643 ± 34, rank #387 of 1776 rated models, from 17 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
BIRD-CRITIC33.67Score80
OTIS Mock AIME 2024-2587.82Accuracy (%)76.9
Mizan LLM Leaderboard65.5Average Score (0-100)70.5
Design Arena (Data Viz)1215Elo61.6
Context-Bench Skills70.48Task Completion (%)52.4
Design Arena (3D)1169Elo48.5
Design Arena (Website)1177Elo45.5
Design Arena (Game Dev)1157Elo40
Design Arena (UI Components)1146Elo37.5
HalluHard0.71Turn-1 Hallucination Rate29
SimpleQA Verified27.5Accuracy (%)25
Design Arena (SVG)1085Elo23.7

Interactive version: theaggregate.ai/model?slug=deepseek-reasoner · How the rankings work · Data refreshed daily, snapshot 2026-07-22.