DeepSeek R1 Distill Llama 70B — benchmark results

DeepSeek's official R1 distillation into Llama 3.3 70B Instruct, fine-tuned on R1-generated reasoning traces (January 2025). Provider: DeepSeek. Released 2025-01-20. Access: Open.

Unified ELO 1504 ± 13, rank #782 of 1776 rated models, from 263 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
French LLM Leaderboard - Average60.02Average Score (%)100
French LLM Leaderboard - GPQA FR63.19Score (%)100
Open CoT - LSAT Analytical Reasoning22.61CoT Gain (%)100
Open CoT - LogiQA12.62CoT Gain (%)98.5
French LLM Leaderboard - BAC FR50.71Score (%)98.1
Open CoT Leaderboard15.3Average CoT Gain (%)96.2
Open FinLLM Reasoning - XBRL-Math86.67Accuracy (%)94
French LLM Leaderboard - IFEval FR66.17Score (%)90.7
Open CoT - LogiQA 214.25CoT Gain (%)89.3
Open Japanese LLM - Wikicorpus J TO E Comet Wmt2276.42Score (%)89.1
Open Japanese LLM - Wikicorpus J TO E Bert Score EN F190.98Score (%)88.5
BRIDGE Medical Leaderboard - CoT38.95Average Performance (%)87.7

Interactive version: theaggregate.ai/model?slug=deepseek-r1-distill-llama-70b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.