Qwen 3 Max (Preview) (Thinking): benchmark results
Preview thinking snapshot of Qwen 3 Max. Provider: Alibaba. Released 2026-01-27. Access: API.
Unified ELO 1631 ± 1, rank #440 of 3078 rated models, from 99 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA Omniscience - Software Engineering (SWE) - R | 30 | Accuracy (%) | 88.7 |
| AA Omniscience - Software Engineering (SWE) - C | 56 | Accuracy (%) | 87.5 |
| AA Omniscience - Software Engineering (SWE) - PHP | 42 | Accuracy (%) | 87.2 |
| AA Omniscience - Software Engineering (SWE) - TypeScript | 42.22 | Accuracy (%) | 86.9 |
| AA Omniscience - Software Engineering (SWE) - Kotlin | 32 | Accuracy (%) | 86.5 |
| AA Omniscience - Software Engineering (SWE) - Python | 34 | Accuracy (%) | 85.7 |
| AA Omniscience - Software Engineering (SWE) - Dart | 30 | Accuracy (%) | 85.2 |
| AA Omniscience - Software Engineering (SWE) - Julia | 28 | Accuracy (%) | 85.2 |
| AA Omniscience - Software Engineering (SWE) - JavaScript | 41.82 | Accuracy (%) | 85 |
| AA Omniscience - Software Engineering (SWE) - Java | 25 | Accuracy (%) | 84.6 |
| AA MMLU-Pro | 82.45 | Accuracy (%) | 80.5 |
| AGI-Eval Community - Interaction (Chinese) | 83.23 | Accuracy (%) | 79.3 |
Interactive version: theaggregate.ai/model?slug=qwen-3-max-preview-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.