Qwen 2.5 72B Instruct — benchmark results

Alibaba Qwen 2.5 72B instruction-tuned checkpoint. Provider: Alibaba. Released 2024-09-19. Access: Open.

Unified ELO 1574 ± 6, rank #550 of 1776 rated models, from 695 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
EVALITA - text-entailment85.07CPS100
Galileo Tool Tasks - BFCL v3 Irrelevance99Accuracy (%)100
HELM SeaHELM - Wisesight55.49Macro F1 score100
INVESTORBENCH46.15Average stock cumulative return (self-reported)100
LLMZSZL Leaderboard69.06Score100
NeedleBench81.02Overall 128K score (self-reported)100
SeaEval - Cross-Lingual Consistency - Cross-LogiQA (Zero-Shot)72.48Accuracy (%)100
SeaEval - Cross-Lingual Consistency - Cross-MMLU (Zero-Shot)81.24Accuracy (%)100
SeaEval - Cross-Lingual Consistency - Cross-XQuAD (Zero-Shot)96.83Accuracy (%)100
SeaEval - Cultural Reasoning - CN-Eval (Zero-Shot)87.62Accuracy (%)100
SeaEval - Dialogue - DREAM (Zero-Shot)96.28Accuracy (%)100
SeaEval - FLORES Translation - Chinese-to-English (Zero-Shot)28.43BLEU (0-100)100

Interactive version: theaggregate.ai/model?slug=qwen-2-5-72b-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.