Qwen 3 0.6B (Thinking) — benchmark results

Qwen 3 0.6B evaluated with thinking enabled. Provider: Alibaba. Released 2025-04-28. Access: Open.

Unified ELO 1307 ± 26, rank #1517 of 1776 rated models, from 46 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AA Humanity's Last Exam5.7Accuracy (%)41.3
AA MATH-50075Accuracy (%)38.1
CritPt0Accuracy (self-reported)30.3
AA CritPt0Accuracy (%)27.1
UGI - Willingness (W/10)3W/10 Score22.2
AA TAU-2 Bench21.05Accuracy (%)20.8
AA AIME 202518Accuracy (%)20.6
BRIDGE Medical Leaderboard - CoT18.95Average Performance (%)13.2
AA Omniscience - Software Engineering (SWE) - R2Accuracy (%)11.9
AA Omniscience - Software Engineering (SWE) - PHP10Accuracy (%)11.5
BRIDGE Medical Leaderboard21.4Average Performance (%)11.3
BRIDGE Medical Leaderboard - Zero-Shot20.38Average Performance (%)11.3

Interactive version: theaggregate.ai/model?slug=qwen-3-0-6b-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.