Qwen 2.5 72B: benchmark results

Alibaba's open 72B Qwen2.5 flagship (September 2024) with a 128K context and much stronger coding and math than Qwen2. Provider: Alibaba. Released 2024-09-19. Access: Open.

Unified ELO 1553 ± 1, rank #423 of 1392 rated models, from 255 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
NeedleBench81.02Overall 128K score (self-reported)100
Open LLM Leaderboard - MMLU-Pro55.2Score99.8
Open LLM Leaderboard - GPQA20.69Score99.4
LLMZSZL Leaderboard68.5Score99
Open PL LLM Leaderboard67.38Average Score (%)98.6
Open Arabic LLM - Aratrust Unfairness96.36Accuracy (%)98.4
Open PL LLM - Generative69.48Average Generative Score (%)97.9
Open PL LLM - Multiple Choice64.65Average Multiple-Choice Score (%)97.9
Open Arabic LLM - Alghafa Multiple Choice Sentiment Task43.66Accuracy (%)97.8
ARC Challenge (AI2)94.5Accuracy (%)97.4
Open Arabic LLM - Arabic MMLU HT High School Computer Science84Accuracy (%)97.2
Open LLM Leaderboard - BBH54.62Score96.6

Interactive version: theaggregate.ai/model?slug=qwen-2-5-72b · How It Works · Data refreshed daily, snapshot 2026-09-05.