Qwen 3 32B: benchmark results
Alibaba's open Qwen 3 32B dense model with hybrid thinking (April 2025). Provider: Alibaba. Released 2025-04-28. Access: Open.
Unified ELO 1552 ± 1, rank #425 of 1392 rated models, from 478 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Open-R1 Eval Leaderboard | 73.74 | Average Accuracy (%) | 100 |
| When the Manual Lies | 98 | Attack Success Rate (ASR) (self-reported) | 100 |
| ChineseSafe Benchmark | 75.26 | Accuracy (%) | 96.3 |
| LLMZSZL Leaderboard | 66.84 | Score | 95.9 |
| MERA - ruOpenBookQA | 96 | Accuracy (%) | 95.2 |
| MERA - ruTiE | 90.91 | Accuracy (%) | 94.3 |
| MERA - LCS | 56.8 | Accuracy (%) | 93.8 |
| Turkish MMLU | 76 | Accuracy (%) | 93.8 |
| MERA - ruHateSpeech | 90.19 | Accuracy (%) | 93 |
| Active Evidence-Seeking and Diagnostic Reasoni | 48.3 | Task 2: Active Seeking Exact Accuracy (self-reported) | 92.9 |
| StakeBench | 17.6 | Agg (self-reported) | 92.9 |
| Open PL LLM - Generative | 68.1 | Average Generative Score (%) | 92.7 |
Interactive version: theaggregate.ai/model?slug=qwen-3-32b · How It Works · Data refreshed daily, snapshot 2026-09-05.