Qwen 3 Max: benchmark results
Alibaba's Qwen 3 Max, the proprietary API-only flagship of the Qwen 3 family (September 2025). Provider: Alibaba. Released 2025-09-24. Access: API.
Unified ELO 1631 ± 1, rank #179 of 1392 rated models, from 205 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Disentangling Language Roles in Multilingual L | 84 | JointSuccess (self-reported) | 100 |
| SciEval - Astronomy | 75.64 | Astronomy (%) | 100 |
| BacktestBench | 61.84 | Overall Accuracy (OA) (self-reported) | 95.5 |
| SGI-Bench Wet Experiment | 33.62 | Wet Experiment Score | 94.4 |
| LOM Benchmark - CycleDetection | 93 | Accuracy (%) | 91.7 |
| WDCD R2 In-Document Resistance | 100 | Score (%) | 90 |
| LLM Arena RU | 1113 | Arena Elo | 89.3 |
| Every Act Has Its Price | 1004.4 | Care (self-reported) | 88.9 |
| SGI-Bench | 31.97 | SGI-Score | 88.9 |
| SciEval - Physics | 41.04 | Physics (%) | 88.9 |
| BenchTable | 72.3 | Total Score (%) | 88.7 |
| SnakeBench | 30.8 | TrueSkill Rating | 88.2 |
Interactive version: theaggregate.ai/model?slug=qwen-3-max · How It Works · Data refreshed daily, snapshot 2026-09-05.