Qwen 2.5 72B Instruct — benchmark results
Alibaba Qwen 2.5 72B instruction-tuned checkpoint. Provider: Alibaba. Released 2024-09-19. Access: Open.
Unified ELO 1574 ± 6, rank #550 of 1776 rated models, from 695 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| EVALITA - text-entailment | 85.07 | CPS | 100 |
| Galileo Tool Tasks - BFCL v3 Irrelevance | 99 | Accuracy (%) | 100 |
| HELM SeaHELM - Wisesight | 55.49 | Macro F1 score | 100 |
| INVESTORBENCH | 46.15 | Average stock cumulative return (self-reported) | 100 |
| LLMZSZL Leaderboard | 69.06 | Score | 100 |
| NeedleBench | 81.02 | Overall 128K score (self-reported) | 100 |
| SeaEval - Cross-Lingual Consistency - Cross-LogiQA (Zero-Shot) | 72.48 | Accuracy (%) | 100 |
| SeaEval - Cross-Lingual Consistency - Cross-MMLU (Zero-Shot) | 81.24 | Accuracy (%) | 100 |
| SeaEval - Cross-Lingual Consistency - Cross-XQuAD (Zero-Shot) | 96.83 | Accuracy (%) | 100 |
| SeaEval - Cultural Reasoning - CN-Eval (Zero-Shot) | 87.62 | Accuracy (%) | 100 |
| SeaEval - Dialogue - DREAM (Zero-Shot) | 96.28 | Accuracy (%) | 100 |
| SeaEval - FLORES Translation - Chinese-to-English (Zero-Shot) | 28.43 | BLEU (0-100) | 100 |
Interactive version: theaggregate.ai/model?slug=qwen-2-5-72b-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.