DeepSeek R1 — benchmark results
DeepSeek's open R1 reasoning model trained with large-scale RL (January 2025). Provider: DeepSeek. Released 2025-01-20. Access: Open.
Unified ELO 1620 ± 11, rank #433 of 1776 rated models, from 337 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| CRMArena - KQA | 61.2 | KQA Score (%) | 100 |
| Fibble3 Arena | 33.33 | Win Rate (%) | 100 |
| HELM MedHELM - Patient Communication | 71.88 | Mean win rate | 100 |
| HELM MedHELM - Race-Based Medicine | 91.62 | EM | 100 |
| SEAL - Fortress | 74.39 | Score | 100 |
| TACTL | 95.9 | Accuracy (%) | 100 |
| VinCLAT (Catalan) | 28.25 | Average Accuracy (%) | 100 |
| TuRTLe - Icarus Syntax | 95.22 | Average Score (%) | 97.7 |
| TuRTLe - Verilator Syntax | 95.9 | Average Score (%) | 97.7 |
| ReliableMath - Precision | 64.2 | Score (%) | 97.4 |
| BRIDGE Medical Leaderboard - Zero-Shot | 44.25 | Average Performance (%) | 97.2 |
| LLM2014 Logic 2025-01 | 84.16 | Score (%) | 96.4 |
Interactive version: theaggregate.ai/model?slug=deepseek-r1 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.