Qwen 2.5 Max: benchmark results
Alibaba's API-only MoE flagship pretrained on 20T+ tokens, launched to rival DeepSeek V3 and GPT-4o (January 2025). Provider: Alibaba. Released 2025-01-29. Access: API.
Unified ELO 1562 ± 1, rank #395 of 1392 rated models, from 65 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| MERA - SimpleAr | 100 | EM (%) | 95.7 |
| MERA - ruTiE | 91.07 | Accuracy (%) | 94.7 |
| MERA - BPS | 99.8 | Accuracy (%) | 93 |
| MERA - ruOpenBookQA | 95.25 | Accuracy (%) | 91.1 |
| MERA - ruDetox | 34.78 | Joint Score (%) | 90.4 |
| MERA - ruHHH | 88.2 | Accuracy (%) | 90.4 |
| MERA - PARus | 93.2 | Accuracy (%) | 89.7 |
| MERA - RWSD | 74.23 | Accuracy (%) | 89.2 |
| MERA - MaMuRAMu | 86.86 | Accuracy (%) | 88.5 |
| MERA - CheGeKa | 44.68 | F1 (%) | 87.1 |
| MERA - ruWorldTree | 98.86 | Accuracy (%) | 85.9 |
| MERA - RCB | 58.68 | Accuracy (%) | 85.2 |
Interactive version: theaggregate.ai/model?slug=qwen-2-5-max · How It Works · Data refreshed daily, snapshot 2026-09-05.