Qwen 2.5 Max — benchmark results
Alibaba's API-only MoE flagship pretrained on 20T+ tokens, launched to rival DeepSeek V3 and GPT-4o (January 2025). Provider: Alibaba. Released 2025-01-29. Access: API.
Unified ELO 1514 ± 18, rank #752 of 1776 rated models, from 23 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| PlatinumBench (MIT) | 2.05 | Avg Error Rate (%) | 60.6 |
| Chatbot Arena (Text) | 1374 | Elo | 58.1 |
| BenchTable | 51.6 | Total Score (%) | 55.5 |
| AA MMLU-Pro | 76.24 | Accuracy (%) | 53.2 |
| AI Chess Leaderboard (Reasoning) | 667 | Elo | 52.2 |
| AA SciCode | 33.68 | Accuracy (%) | 50.9 |
| AA MATH-500 | 83.47 | Accuracy (%) | 50.5 |
| AA LiveCodeBench | 35.87 | Pass@1 (%) | 43.4 |
| WebApp1K | 57.6 | Pass@1 (%) | 42.4 |
| AA GPQA Diamond | 58.69 | Accuracy (%) | 36.2 |
| Artificial Analysis Intelligence Index | 10.23 | Intelligence Index | 35.8 |
| Step Game (Lechmazur) | 1.43 | TrueSkill μ | 34.5 |
Interactive version: theaggregate.ai/model?slug=qwen-2-5-max · How the rankings work · Data refreshed daily, snapshot 2026-07-22.