DeepSeek R1 Distill Qwen 32B: benchmark results
DeepSeek's official R1 reasoning distill onto Qwen2.5-32B (MIT license), SFT-trained on ~800K R1-generated reasoning traces. Provider: DeepSeek. Released 2025-01-20. Access: Open.
Unified ELO 1493 ± 1, rank #721 of 1392 rated models, from 410 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| MERA - PARus | 96.2 | Accuracy (%) | 100 |
| Open CoT - LogiQA 2 | 19.02 | CoT Gain (%) | 100 |
| Open CoT Leaderboard | 16.92 | Average CoT Gain (%) | 99.2 |
| Thai LLM - Extraction | 7.85 | Rating (0-10) | 98.6 |
| Open CoT - LSAT Analytical Reasoning | 16.52 | CoT Gain (%) | 98.5 |
| French LLM Leaderboard - GPQA FR | 54.4 | Score (%) | 98.1 |
| Open CoT - LogiQA | 10.06 | CoT Gain (%) | 94.7 |
| French LLM Leaderboard - Average | 53.39 | Average Score (%) | 94.4 |
| French LLM Leaderboard - BAC FR | 45.62 | Score (%) | 94.4 |
| Open CoT - LSAT Logical Reasoning | 20.39 | CoT Gain (%) | 93.9 |
| Open Japanese LLM - Wiki NER SET F1 | 14.16 | Score (%) | 93.1 |
| Thai LLM - Reasoning | 6.8 | Rating (0-10) | 92.3 |
Interactive version: theaggregate.ai/model?slug=deepseek-r1-distill-qwen-32b · How It Works · Data refreshed daily, snapshot 2026-09-05.