Qwen 3 Max (Non-reasoning): benchmark results

Provider: Alibaba. Released 2025-09-24. Access: API.

Unified ELO 1632 ± 24, rank #593 of 2133 rated models, from 31 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
RecRM-Bench - Query-Item Relevance76.64Accuracy (%) of the three-level relevance score (irrelevant,100
IndustryBench - Process Principles2.38Final (SV) score (out of 3): mean 0-3 rubric score from a Qw87.5
AGI-Eval Community - Learning (Chinese)83.66Accuracy (%)82.6
AGI-Eval Community - General Reasoning97.1Accuracy (%)81.4
AGI-Eval Community - Subject Reasoning (Chinese)88.97Accuracy (%)79
AGI-Eval Community - Subject Reasoning85.64Accuracy (%)74.3
AGI-Eval Community - Learning87.93Accuracy (%)72.1
IndustryBench - Quality and Metrology2.15Final (SV) score (out of 3): mean 0-3 rubric score from a Qw68.8
AGI-Eval Community - Subject Reasoning (English)84.16Accuracy (%)67.4
AGI-Eval Community - Subject Knowledge85.06Accuracy (%)63.9
IndustryBench - Selection and Substitution2.04Final (SV) score (out of 3): mean 0-3 rubric score from a Qw62.5
AGI-Eval Community - Mathematical Reasoning78.03Accuracy (%)61.4

Interactive version: theaggregate.ai/model?slug=qwen-3-max-non-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-11.