Qwen 3.7 Max — benchmark results
Alibaba's flagship Qwen 3.7 model for broad high-end chat and reasoning workloads. Provider: Alibaba. Released 2026-05-20. Access: API.
Unified ELO 1836 ± 11, rank #110 of 1776 rated models, from 189 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| FutureX | 49.33 | Overall Score (latest week, %) | 100 |
| LLM Stats (HMMT Feb 26) | 97.1 | Score (%) | 100 |
| LLM Stats (MAXIFE) | 89.2 | Score (%) | 100 |
| LLM Stats (MMLU-ProX) | 87 | Score (%) | 100 |
| LLM Stats (MMLU-Redux) | 95 | Score (%) | 100 |
| LLM Stats (PolyMATH) | 86.5 | Score (%) | 100 |
| LLM Stats (SkillsBench) | 59.2 | Score (%) | 100 |
| Qwen3.7 Launch - Apex | 44.5 | Score (%) | 100 |
| Qwen3.7 Launch - GPQA Diamond | 92.4 | Score (%) | 100 |
| Qwen3.7 Launch - Global PIQA | 91.4 | Score (%) | 100 |
| Qwen3.7 Launch - HMMT 2026 Feb | 97.1 | Score (%) | 100 |
| Qwen3.7 Launch - Humanity's Last Exam | 41.4 | Score (%) | 100 |
Interactive version: theaggregate.ai/model?slug=qwen-3-7-max · How the rankings work · Data refreshed daily, snapshot 2026-07-22.