Qwen 3 Max (Thinking) — benchmark results
Current thinking snapshot of Qwen 3 Max. Provider: Alibaba. Released 2026-01-27. Access: API.
Unified ELO 1724 ± 34, rank #231 of 1776 rated models, from 48 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Tau-Bench Telecom | 98.2 | Pass@1 (%) | 100 |
| AI Chess Leaderboard (Reasoning) | 1800 | Elo | 99.3 |
| Tau-Bench Airline | 69 | Pass@1 (%) | 86.7 |
| AA GPQA Diamond | 86.06 | Accuracy (%) | 86.3 |
| AA Humanity's Last Exam | 26.18 | Accuracy (%) | 86.1 |
| AA IFBench | 70.75 | Accuracy (%) | 85.6 |
| AA Long Context Reasoning | 66 | Accuracy (%) | 84.1 |
| AA Omniscience - Software Engineering (SWE) - PHP | 48 | Accuracy (%) | 82.7 |
| LLM2014 Logic 2025-11 | 53.6 | Median Score | 82.7 |
| AA SciCode | 43.06 | Accuracy (%) | 82.2 |
| AA Omniscience - Science, Engineering & Mathematics | 37.8 | Accuracy (%) | 81.9 |
| AA Omniscience - Humanities & Social Sciences | 31.6 | Accuracy (%) | 81.1 |
Interactive version: theaggregate.ai/model?slug=qwen-3-max-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.