Qwen 3 Max Instruct: benchmark results
Provider: Alibaba. Released 2025-09-24. Access: API.
Unified ELO 1748 ± 32, rank #242 of 2656 rated models, from 22 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AGI-Eval Community - Learning (Chinese) | 83.59 | Accuracy (%) | 80.4 |
| AGI-Eval Community - Subject Reasoning (English) | 84.95 | Accuracy (%) | 77.5 |
| AGI-Eval Community - Subject Reasoning | 86.07 | Accuracy (%) | 76.8 |
| AGI-Eval Community - Subject Reasoning (Chinese) | 88.61 | Accuracy (%) | 76.8 |
| AGI-Eval Community - Learning | 88.06 | Accuracy (%) | 75 |
| AGI-Eval Community - Mathematical Reasoning | 79.84 | Accuracy (%) | 70 |
| AGI-Eval Community - General Reasoning | 96.89 | Accuracy (%) | 68.9 |
| AGI-Eval Community - Subject Knowledge | 85.28 | Accuracy (%) | 67.1 |
| OTIS Mock AIME 2024-2025 | 73.33 | Score (%) | 63.5 |
| AGI-Eval Community - Objective Accuracy (English) | 87.32 | Accuracy (%) | 61.6 |
| AGI-Eval Community - Learning (English) | 90.99 | Accuracy (%) | 57.2 |
| AGI-Eval Community - Interaction (Chinese) | 77.2 | Accuracy (%) | 55.8 |
Interactive version: theaggregate.ai/model?slug=qwen-3-max-instruct · How It Works · Data refreshed daily, snapshot 2026-09-19.