Qwen 3 32B (Thinking): benchmark results
Qwen 3 32B evaluated with thinking enabled. Provider: Alibaba. Released 2025-04-28. Access: Open.
Unified ELO 1523 ± 1, rank #740 of 1761 rated models, from 42 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| BRIDGE Medical Leaderboard - Zero-Shot | 41.04 | Average Performance (%) | 86.1 |
| KORGym - Strategic | 0.71 | Score | 83.3 |
| CritPt | 30 | Accuracy (self-reported) | 80.8 |
| KORGym - Control and Interaction | 0.58 | Score | 72.2 |
| KORGym - Mathematical and Logical | 0.58 | Score | 72.2 |
| KORGym - Spatial and Geometric | 0.55 | Score | 72.2 |
| LLM2014 Logic 2025-06 | 50.72 | Median Score | 72.2 |
| BRIDGE Medical Leaderboard | 40.96 | Average Performance (%) | 71.3 |
| BRIDGE Medical Leaderboard - Few-Shot | 47.38 | Average Performance (%) | 67.6 |
| KORGym | 0.6 | Score | 66.7 |
| KORGym - Puzzle | 0.55 | Score | 61.1 |
| BRIDGE Medical Leaderboard - CoT | 34.46 | Average Performance (%) | 58.3 |
Interactive version: theaggregate.ai/model?slug=qwen-3-32b-thinking · How It Works · Data refreshed daily, snapshot 2026-09-05.