DeepSeek R1 0528 — benchmark results

May 28, 2025 update of DeepSeek's open R1 reasoning model. Provider: DeepSeek. Released 2025-05-28. Access: Open.

Unified ELO 1655 ± 12, rank #356 of 1776 rated models, from 231 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Kluster Hallucination Detection - RAG Method 1 Resistance96.48Resistance (100 - Hallucination Rate %)100
Kluster Hallucination Detection - RAG Method 3 Resistance96.42Resistance (100 - Hallucination Rate %)100
USAMO2530.06Score (self-reported)100
HELM Safety XSTest98.8LM Evaluated Safety score (%)98.8
RewardBench 2 Focus93.59Accuracy (%)98
TuRTLe - Icarus Performance74.12Average Score (%)97.7
TuRTLe - Icarus Synthesis75.33Average Score (%)97.7
TuRTLe - Verilator Area76.37Average Score (%)97.7
TuRTLe - Verilator Functionality76.69Average Score (%)97.7
TuRTLe - Verilator Performance73.25Average Score (%)97.7
TuRTLe - Verilator Synthesis74.31Average Score (%)97.7
BenCzechMark84.82Average Score (%)97

Interactive version: theaggregate.ai/model?slug=deepseek-r1-0528 · How the rankings work · Data refreshed daily, snapshot 2026-07-22.