Qwen 3 32B (Thinking) — benchmark results

Qwen 3 32B evaluated with thinking enabled. Provider: Alibaba. Released 2025-04-28. Access: Open.

Unified ELO 1599 ± 17, rank #477 of 1776 rated models, from 61 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
BRIDGE Medical Leaderboard - Zero-Shot41.04Average Performance (%)87.7
KORGym - Strategic0.71Score83.3
AA MATH-50096.07Accuracy (%)82.7
BRIDGE Medical Leaderboard40.96Average Performance (%)72.6
KORGym - Control and Interaction0.58Score72.2
KORGym - Mathematical and Logical0.58Score72.2
KORGym - Spatial and Geometric0.55Score72.2
LLM2014 Logic 2025-0650.72Median Score72.2
BRIDGE Medical Leaderboard - Few-Shot47.38Average Performance (%)68.9
AA AIME 202573Accuracy (%)67.6
KORGym0.6Score66.7
AA MMLU-Pro79.85Accuracy (%)66

Interactive version: theaggregate.ai/model?slug=qwen-3-32b-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.