Qwen Max (2025-01-25): benchmark results

Provider: Alibaba. Released 2025-04-28. Access: API.

Unified ELO 1683 ± 36, rank #361 of 2656 rated models, from 10 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
pfgen-bench - QA Mode - Score0.84pfgen Score (mean of three)97.6
pfgen-bench - QA Mode - Helpfulness0.68Helpfulness Score97
pfgen-bench - QA Mode - Truthfulness0.97Truthfulness Score96.4
pfgen-bench - QA Mode - Fluency0.86Fluency Score95.8
SuperGPQA50.08Accuracy (%)66.2
Web-Bench15.9Pass@1 (%)59.6
MATH Level 567.18Accuracy (%)57
Epoch AI - GPQA Diamond56.12Accuracy (%)32.4
Aider polyglot coding leaderboard21.8Pass rate (%)19.7
OTIS Mock AIME 2024-2516.11Accuracy (%)19.2

Interactive version: theaggregate.ai/model?slug=qwen-max-2025-01-25 · How It Works · Data refreshed daily, snapshot 2026-09-19.