Qwen 3 0.6B (Thinking): benchmark results

Qwen 3 0.6B evaluated with thinking enabled. Provider: Alibaba. Released 2025-04-28. Access: Open.

Unified ELO 1355 ± 1, rank #1528 of 1761 rated models, from 25 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AA Humanity's Last Exam5.64Accuracy (%)37.1
AA CritPt0Accuracy (%)24.4
UGI - Willingness (W/10)3W/10 Score22.7
AA TAU-2 Bench21.05Accuracy (%)21
BRIDGE Medical Leaderboard - CoT18.95Average Performance (%)13
BRIDGE Medical Leaderboard21.4Average Performance (%)11.1
BRIDGE Medical Leaderboard - Zero-Shot20.38Average Performance (%)11.1
Artificial Analysis Intelligence Index1Intelligence Index9.4
BRIDGE Medical Leaderboard - Few-Shot24.87Average Performance (%)8.3
WritingBench45.08Score (self-reported)7.4
AA Long Context Reasoning0Accuracy (%)6.1
AA IFBench23.33Accuracy (%)5.6

Interactive version: theaggregate.ai/model?slug=qwen-3-0-6b-thinking · How It Works · Data refreshed daily, snapshot 2026-09-05.