Qwen 3 32B (Thinking): benchmark results

Qwen 3 32B evaluated with thinking enabled. Provider: Alibaba. Released 2025-04-28. Access: Open.

Unified ELO 1523 ± 1, rank #740 of 1761 rated models, from 42 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
BRIDGE Medical Leaderboard - Zero-Shot41.04Average Performance (%)86.1
KORGym - Strategic0.71Score83.3
CritPt30Accuracy (self-reported)80.8
KORGym - Control and Interaction0.58Score72.2
KORGym - Mathematical and Logical0.58Score72.2
KORGym - Spatial and Geometric0.55Score72.2
LLM2014 Logic 2025-0650.72Median Score72.2
BRIDGE Medical Leaderboard40.96Average Performance (%)71.3
BRIDGE Medical Leaderboard - Few-Shot47.38Average Performance (%)67.6
KORGym0.6Score66.7
KORGym - Puzzle0.55Score61.1
BRIDGE Medical Leaderboard - CoT34.46Average Performance (%)58.3

Interactive version: theaggregate.ai/model?slug=qwen-3-32b-thinking · How It Works · Data refreshed daily, snapshot 2026-09-05.