Qwen Max — benchmark results
Alibaba's flagship Qwen Max model, represented here as a 700B-parameter closed model. Provider: Alibaba. Released 2025-04-28. Access: API.
Unified ELO 1465 ± 91, rank #932 of 1776 rated models, from 12 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| DuckDB-NSQL | 64 | Execution Accuracy (%) | 77.6 |
| SpeechMap Compliance | 66.8 | % Requests Completed | 56.1 |
| HSCodeComp - 10-digit | 3.8 | Exact Match Accuracy (%) | 46.2 |
| HSCodeComp - 8-digit | 11.23 | Exact Match Accuracy (%) | 46.2 |
| LLM Chess (Saplin) | -98.9 | ELO | 42.9 |
| SnakeBench | 20.3 | TrueSkill Rating | 41.3 |
| HSCodeComp - 2-digit | 71.52 | Exact Match Accuracy (%) | 38.5 |
| HSCodeComp - 4-digit | 48.58 | Exact Match Accuracy (%) | 30.8 |
| HSCodeComp - 6-digit | 24.21 | Exact Match Accuracy (%) | 30.8 |
| LiveCodeBench Pro | 274 | Rating (CF-style) | 1.9 |
| PHYBench | 15.33 | EED Score | 0 |
| SWE-Arena | 996 | Elo Score | 0 |
Interactive version: theaggregate.ai/model?slug=qwen-max · How the rankings work · Data refreshed daily, snapshot 2026-07-22.