Qwen 2.5 Max — benchmark results

Alibaba's API-only MoE flagship pretrained on 20T+ tokens, launched to rival DeepSeek V3 and GPT-4o (January 2025). Provider: Alibaba. Released 2025-01-29. Access: API.

Unified ELO 1514 ± 18, rank #752 of 1776 rated models, from 23 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
PlatinumBench (MIT)2.05Avg Error Rate (%)60.6
Chatbot Arena (Text)1374Elo58.1
BenchTable51.6Total Score (%)55.5
AA MMLU-Pro76.24Accuracy (%)53.2
AI Chess Leaderboard (Reasoning)667Elo52.2
AA SciCode33.68Accuracy (%)50.9
AA MATH-50083.47Accuracy (%)50.5
AA LiveCodeBench35.87Pass@1 (%)43.4
WebApp1K57.6Pass@1 (%)42.4
AA GPQA Diamond58.69Accuracy (%)36.2
Artificial Analysis Intelligence Index10.23Intelligence Index35.8
Step Game (Lechmazur)1.43TrueSkill μ34.5

Interactive version: theaggregate.ai/model?slug=qwen-2-5-max · How the rankings work · Data refreshed daily, snapshot 2026-07-22.