Qwen 3.7 Max (Thinking): benchmark results

Provider: Alibaba. Released 2026-05-20. Access: API.

Unified ELO 1677 ± 1, rank #84 of 2032 rated models, from 31 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AeroCopilotBench - Tier-1 Aviation Knowledge89.17Accuracy (%; 1,200 multiple-choice questions)100
PortBench-QA - VaR Estimation85.9Item score (%; historical-simulation value at risk from a su100
PortBench-QA81.9Mean item score (%; mean of seven QA templates, 50 test ques94.4
SecCodeBench66.81Total Score94.3
SuperCLUE General (May 2026) - Code Generation79.69Score91.3
SuperCLUE-SWE - Overall66.67Score90.6
SuperCLUE General (May 2026) - Math Reasoning82.46Score89.1
SuperCLUE General (May 2026) - Science Reasoning73.68Score87
AeroCopilotBench - Tier-2 Safety-Gated Success Rate58.9Success rate (%; goals met with no safety violation)81.8
SuperCLUE-LongContext - 256K Overall76.2Score80
Korean CSAT 2026 (Easy Mode) - Society and Culture44Points (out of 50)79.9
PortBench-QA - Max-Sharpe Allocation95.4Item score (%; long-only maximum-Sharpe weights for three or77.8

Interactive version: theaggregate.ai/model?slug=qwen-3-7-max-thinking · How It Works · Data refreshed daily, snapshot 2026-09-26.