Qwen 3 Max (Non-reasoning): benchmark results
Provider: Alibaba. Released 2025-09-24. Access: API.
Unified ELO 1632 ± 24, rank #593 of 2133 rated models, from 31 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| RecRM-Bench - Query-Item Relevance | 76.64 | Accuracy (%) of the three-level relevance score (irrelevant, | 100 |
| IndustryBench - Process Principles | 2.38 | Final (SV) score (out of 3): mean 0-3 rubric score from a Qw | 87.5 |
| AGI-Eval Community - Learning (Chinese) | 83.66 | Accuracy (%) | 82.6 |
| AGI-Eval Community - General Reasoning | 97.1 | Accuracy (%) | 81.4 |
| AGI-Eval Community - Subject Reasoning (Chinese) | 88.97 | Accuracy (%) | 79 |
| AGI-Eval Community - Subject Reasoning | 85.64 | Accuracy (%) | 74.3 |
| AGI-Eval Community - Learning | 87.93 | Accuracy (%) | 72.1 |
| IndustryBench - Quality and Metrology | 2.15 | Final (SV) score (out of 3): mean 0-3 rubric score from a Qw | 68.8 |
| AGI-Eval Community - Subject Reasoning (English) | 84.16 | Accuracy (%) | 67.4 |
| AGI-Eval Community - Subject Knowledge | 85.06 | Accuracy (%) | 63.9 |
| IndustryBench - Selection and Substitution | 2.04 | Final (SV) score (out of 3): mean 0-3 rubric score from a Qw | 62.5 |
| AGI-Eval Community - Mathematical Reasoning | 78.03 | Accuracy (%) | 61.4 |
Interactive version: theaggregate.ai/model?slug=qwen-3-max-non-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-11.