QwQ-32B — benchmark results

Alibaba's RL-trained 32B reasoning model, matching the far larger DeepSeek-R1 on math and coding under Apache 2.0 (March 2025). Provider: Alibaba. Released 2025-03-05. Access: Open.

Unified ELO 1501 ± 10, rank #793 of 1776 rated models, from 409 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
VMLU76.13Average (%)100
VMLU - Other70.62Accuracy (%)100
VMLU - STEM81.11Accuracy (%)100
VMLU - Social Science78.49Accuracy (%)100
Open Japanese LLM - Wiki NER SET F117.7Score (%)99
ProLLM - OpenBook Q&A91.9Score (%)97.8
EuroEval English NLU - SST-570.14Sentiment classification Score (%)97.6
EuroEval Dutch NLU - CoNLL NL75.09Named entity recognition Score (%)96.5
EuroEval Dutch73.98Average Score (%)96.2
VMLU - Humanities71.78Accuracy (%)95.8
EuroEval Italian66.34Average Score (%)95.1
EuroEval Spanish61.55Average Score (%)94.2

Interactive version: theaggregate.ai/model?slug=qwq-32b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.