Qwen 3.7 Max: benchmark results

Alibaba's flagship Qwen 3.7 model for broad high-end chat and reasoning workloads. Provider: Alibaba. Released 2026-05-20. Access: API.

Unified ELO 1729 ± 1, rank #26 of 1392 rated models, from 247 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HMMT February 202697.1Score (self-reported)100
LLM Stats (HMMT Feb 26)97.1Score (%)100
LLM Stats (MAXIFE)89.2Score (%)100
LLM Stats (MMLU-ProX)87Score (%)100
LLM Stats (PolyMATH)86.5Score (%)100
Qwen3.7 Launch - Apex44.5Score (%)100
Qwen3.7 Launch - GPQA Diamond92.4Score (%)100
Qwen3.7 Launch - Global PIQA91.4Score (%)100
Qwen3.7 Launch - IFBench79.1Score (%)100
Qwen3.7 Launch - IMOAnswerBench90Score (%)100
Qwen3.7 Launch - MRCR-v2 128k90.4Accuracy (%)100
Qwen3.7 Launch - QwenSVG1608Elo100

Interactive version: theaggregate.ai/model?slug=qwen-3-7-max · How It Works · Data refreshed daily, snapshot 2026-09-05.