QwQ-32B: benchmark results

Alibaba's RL-trained 32B reasoning model, matching the far larger DeepSeek-R1 on math and coding under Apache 2.0 (March 2025). Provider: Alibaba. Released 2025-03-05. Access: Open.

Unified ELO 1534 ± 1, rank #514 of 1392 rated models, from 426 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
VMLU76.13Average (%)100
VMLU - Other70.62Accuracy (%)100
VMLU - STEM81.11Accuracy (%)100
VMLU - Social Science78.49Accuracy (%)100
Open Japanese LLM - Wiki NER SET F117.7Score (%)99
ProLLM - OpenBook Q&A91.9Score (%)97.8
EuroEval English NLU - SST-570.14Sentiment classification Score (%)97.6
EuroEval Dutch NLU - CoNLL NL75.09Named entity recognition Score (%)96.5
EuroEval Dutch73.98Average Score (%)96.2
VMLU - Humanities71.78Accuracy (%)95.8
EuroEval Italian66.34Average Score (%)95.1
EuroEval Spanish61.55Average Score (%)94.2

Interactive version: theaggregate.ai/model?slug=qwq-32b · How It Works · Data refreshed daily, snapshot 2026-09-05.