Qwen 2.5 Max: benchmark results

Alibaba's API-only MoE flagship pretrained on 20T+ tokens, launched to rival DeepSeek V3 and GPT-4o (January 2025). Provider: Alibaba. Released 2025-01-29. Access: API.

Unified ELO 1562 ± 1, rank #395 of 1392 rated models, from 65 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
MERA - SimpleAr100EM (%)95.7
MERA - ruTiE91.07Accuracy (%)94.7
MERA - BPS99.8Accuracy (%)93
MERA - ruOpenBookQA95.25Accuracy (%)91.1
MERA - ruDetox34.78Joint Score (%)90.4
MERA - ruHHH88.2Accuracy (%)90.4
MERA - PARus93.2Accuracy (%)89.7
MERA - RWSD74.23Accuracy (%)89.2
MERA - MaMuRAMu86.86Accuracy (%)88.5
MERA - CheGeKa44.68F1 (%)87.1
MERA - ruWorldTree98.86Accuracy (%)85.9
MERA - RCB58.68Accuracy (%)85.2

Interactive version: theaggregate.ai/model?slug=qwen-2-5-max · How It Works · Data refreshed daily, snapshot 2026-09-05.