DeepSeek R1 0528 — benchmark results
May 28, 2025 update of DeepSeek's open R1 reasoning model. Provider: DeepSeek. Released 2025-05-28. Access: Open.
Unified ELO 1655 ± 12, rank #356 of 1776 rated models, from 231 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Kluster Hallucination Detection - RAG Method 1 Resistance | 96.48 | Resistance (100 - Hallucination Rate %) | 100 |
| Kluster Hallucination Detection - RAG Method 3 Resistance | 96.42 | Resistance (100 - Hallucination Rate %) | 100 |
| USAMO25 | 30.06 | Score (self-reported) | 100 |
| HELM Safety XSTest | 98.8 | LM Evaluated Safety score (%) | 98.8 |
| RewardBench 2 Focus | 93.59 | Accuracy (%) | 98 |
| TuRTLe - Icarus Performance | 74.12 | Average Score (%) | 97.7 |
| TuRTLe - Icarus Synthesis | 75.33 | Average Score (%) | 97.7 |
| TuRTLe - Verilator Area | 76.37 | Average Score (%) | 97.7 |
| TuRTLe - Verilator Functionality | 76.69 | Average Score (%) | 97.7 |
| TuRTLe - Verilator Performance | 73.25 | Average Score (%) | 97.7 |
| TuRTLe - Verilator Synthesis | 74.31 | Average Score (%) | 97.7 |
| BenCzechMark | 84.82 | Average Score (%) | 97 |
Interactive version: theaggregate.ai/model?slug=deepseek-r1-0528 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.