DeepSeek R1 — benchmark results

DeepSeek's open R1 reasoning model trained with large-scale RL (January 2025). Provider: DeepSeek. Released 2025-01-20. Access: Open.

Unified ELO 1620 ± 11, rank #433 of 1776 rated models, from 337 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
CRMArena - KQA61.2KQA Score (%)100
Fibble3 Arena33.33Win Rate (%)100
HELM MedHELM - Patient Communication71.88Mean win rate100
HELM MedHELM - Race-Based Medicine91.62EM100
SEAL - Fortress74.39Score100
TACTL95.9Accuracy (%)100
VinCLAT (Catalan)28.25Average Accuracy (%)100
TuRTLe - Icarus Syntax95.22Average Score (%)97.7
TuRTLe - Verilator Syntax95.9Average Score (%)97.7
ReliableMath - Precision64.2Score (%)97.4
BRIDGE Medical Leaderboard - Zero-Shot44.25Average Performance (%)97.2
LLM2014 Logic 2025-0184.16Score (%)96.4

Interactive version: theaggregate.ai/model?slug=deepseek-r1 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.