Qwen 3.7 Max (Thinking): benchmark results
Provider: Alibaba. Released 2026-05-20. Access: API.
Unified ELO 1677 ± 1, rank #84 of 2032 rated models, from 31 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AeroCopilotBench - Tier-1 Aviation Knowledge | 89.17 | Accuracy (%; 1,200 multiple-choice questions) | 100 |
| PortBench-QA - VaR Estimation | 85.9 | Item score (%; historical-simulation value at risk from a su | 100 |
| PortBench-QA | 81.9 | Mean item score (%; mean of seven QA templates, 50 test ques | 94.4 |
| SecCodeBench | 66.81 | Total Score | 94.3 |
| SuperCLUE General (May 2026) - Code Generation | 79.69 | Score | 91.3 |
| SuperCLUE-SWE - Overall | 66.67 | Score | 90.6 |
| SuperCLUE General (May 2026) - Math Reasoning | 82.46 | Score | 89.1 |
| SuperCLUE General (May 2026) - Science Reasoning | 73.68 | Score | 87 |
| AeroCopilotBench - Tier-2 Safety-Gated Success Rate | 58.9 | Success rate (%; goals met with no safety violation) | 81.8 |
| SuperCLUE-LongContext - 256K Overall | 76.2 | Score | 80 |
| Korean CSAT 2026 (Easy Mode) - Society and Culture | 44 | Points (out of 50) | 79.9 |
| PortBench-QA - Max-Sharpe Allocation | 95.4 | Item score (%; long-only maximum-Sharpe weights for three or | 77.8 |
Interactive version: theaggregate.ai/model?slug=qwen-3-7-max-thinking · How It Works · Data refreshed daily, snapshot 2026-09-26.