Qwen 3 Max (Thinking): benchmark results

Current thinking snapshot of Qwen 3 Max. Provider: Alibaba. Released 2026-01-27. Access: API.

Unified ELO 1608 ± 1, rank #397 of 1761 rated models, from 44 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Tau-Bench Telecom98.2Pass@1 (%)100
AI Chess Leaderboard (Reasoning)1800Elo98.8
Tau-Bench Airline69Pass@1 (%)86.7
DisasterBench66.52Exact-match Accuracy (self-reported)86.4
AA IFBench70.75Accuracy (%)85.6
LLM2014 Logic 2025-1153.6Median Score82.7
AA GPQA Diamond86.06Accuracy (%)80.6
AA Humanity's Last Exam27.99Accuracy (%)79.7
BenchTable65.5Total Score (%)79.1
Conceptual Reasoning Index - Decision Theory (DTBench)79.55Chance-Corrected Score (0-100)78.8
BenchmarkList ECI131.45Capability Index (ECI)77.2
AA Long Context Reasoning74.33Accuracy (%)76.6

Interactive version: theaggregate.ai/model?slug=qwen-3-max-thinking · How It Works · Data refreshed daily, snapshot 2026-09-05.