Qwen 3 Max (Thinking): benchmark results
Current thinking snapshot of Qwen 3 Max. Provider: Alibaba. Released 2026-01-27. Access: API.
Unified ELO 1608 ± 1, rank #397 of 1761 rated models, from 44 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Tau-Bench Telecom | 98.2 | Pass@1 (%) | 100 |
| AI Chess Leaderboard (Reasoning) | 1800 | Elo | 98.8 |
| Tau-Bench Airline | 69 | Pass@1 (%) | 86.7 |
| DisasterBench | 66.52 | Exact-match Accuracy (self-reported) | 86.4 |
| AA IFBench | 70.75 | Accuracy (%) | 85.6 |
| LLM2014 Logic 2025-11 | 53.6 | Median Score | 82.7 |
| AA GPQA Diamond | 86.06 | Accuracy (%) | 80.6 |
| AA Humanity's Last Exam | 27.99 | Accuracy (%) | 79.7 |
| BenchTable | 65.5 | Total Score (%) | 79.1 |
| Conceptual Reasoning Index - Decision Theory (DTBench) | 79.55 | Chance-Corrected Score (0-100) | 78.8 |
| BenchmarkList ECI | 131.45 | Capability Index (ECI) | 77.2 |
| AA Long Context Reasoning | 74.33 | Accuracy (%) | 76.6 |
Interactive version: theaggregate.ai/model?slug=qwen-3-max-thinking · How It Works · Data refreshed daily, snapshot 2026-09-05.