Qwen Max (2025-01-25): benchmark results
Provider: Alibaba. Released 2025-04-28. Access: API.
Unified ELO 1683 ± 36, rank #361 of 2656 rated models, from 10 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| pfgen-bench - QA Mode - Score | 0.84 | pfgen Score (mean of three) | 97.6 |
| pfgen-bench - QA Mode - Helpfulness | 0.68 | Helpfulness Score | 97 |
| pfgen-bench - QA Mode - Truthfulness | 0.97 | Truthfulness Score | 96.4 |
| pfgen-bench - QA Mode - Fluency | 0.86 | Fluency Score | 95.8 |
| SuperGPQA | 50.08 | Accuracy (%) | 66.2 |
| Web-Bench | 15.9 | Pass@1 (%) | 59.6 |
| MATH Level 5 | 67.18 | Accuracy (%) | 57 |
| Epoch AI - GPQA Diamond | 56.12 | Accuracy (%) | 32.4 |
| Aider polyglot coding leaderboard | 21.8 | Pass rate (%) | 19.7 |
| OTIS Mock AIME 2024-25 | 16.11 | Accuracy (%) | 19.2 |
Interactive version: theaggregate.ai/model?slug=qwen-max-2025-01-25 · How It Works · Data refreshed daily, snapshot 2026-09-19.