Qwen 3 Max (Thinking) — benchmark results

Current thinking snapshot of Qwen 3 Max. Provider: Alibaba. Released 2026-01-27. Access: API.

Unified ELO 1724 ± 34, rank #231 of 1776 rated models, from 48 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Tau-Bench Telecom98.2Pass@1 (%)100
AI Chess Leaderboard (Reasoning)1800Elo99.3
Tau-Bench Airline69Pass@1 (%)86.7
AA GPQA Diamond86.06Accuracy (%)86.3
AA Humanity's Last Exam26.18Accuracy (%)86.1
AA IFBench70.75Accuracy (%)85.6
AA Long Context Reasoning66Accuracy (%)84.1
AA Omniscience - Software Engineering (SWE) - PHP48Accuracy (%)82.7
LLM2014 Logic 2025-1153.6Median Score82.7
AA SciCode43.06Accuracy (%)82.2
AA Omniscience - Science, Engineering & Mathematics37.8Accuracy (%)81.9
AA Omniscience - Humanities & Social Sciences31.6Accuracy (%)81.1

Interactive version: theaggregate.ai/model?slug=qwen-3-max-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.