Qwen 3.7 Max — benchmark results

Alibaba's flagship Qwen 3.7 model for broad high-end chat and reasoning workloads. Provider: Alibaba. Released 2026-05-20. Access: API.

Unified ELO 1836 ± 11, rank #110 of 1776 rated models, from 189 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
FutureX49.33Overall Score (latest week, %)100
LLM Stats (HMMT Feb 26)97.1Score (%)100
LLM Stats (MAXIFE)89.2Score (%)100
LLM Stats (MMLU-ProX)87Score (%)100
LLM Stats (MMLU-Redux)95Score (%)100
LLM Stats (PolyMATH)86.5Score (%)100
LLM Stats (SkillsBench)59.2Score (%)100
Qwen3.7 Launch - Apex44.5Score (%)100
Qwen3.7 Launch - GPQA Diamond92.4Score (%)100
Qwen3.7 Launch - Global PIQA91.4Score (%)100
Qwen3.7 Launch - HMMT 2026 Feb97.1Score (%)100
Qwen3.7 Launch - Humanity's Last Exam41.4Score (%)100

Interactive version: theaggregate.ai/model?slug=qwen-3-7-max · How the rankings work · Data refreshed daily, snapshot 2026-07-22.