Qwen 3 32B (Thinking) — benchmark results
Qwen 3 32B evaluated with thinking enabled. Provider: Alibaba. Released 2025-04-28. Access: Open.
Unified ELO 1599 ± 17, rank #477 of 1776 rated models, from 61 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| BRIDGE Medical Leaderboard - Zero-Shot | 41.04 | Average Performance (%) | 87.7 |
| KORGym - Strategic | 0.71 | Score | 83.3 |
| AA MATH-500 | 96.07 | Accuracy (%) | 82.7 |
| BRIDGE Medical Leaderboard | 40.96 | Average Performance (%) | 72.6 |
| KORGym - Control and Interaction | 0.58 | Score | 72.2 |
| KORGym - Mathematical and Logical | 0.58 | Score | 72.2 |
| KORGym - Spatial and Geometric | 0.55 | Score | 72.2 |
| LLM2014 Logic 2025-06 | 50.72 | Median Score | 72.2 |
| BRIDGE Medical Leaderboard - Few-Shot | 47.38 | Average Performance (%) | 68.9 |
| AA AIME 2025 | 73 | Accuracy (%) | 67.6 |
| KORGym | 0.6 | Score | 66.7 |
| AA MMLU-Pro | 79.85 | Accuracy (%) | 66 |
Interactive version: theaggregate.ai/model?slug=qwen-3-32b-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.