Qwen 2 72B Instruct: benchmark results

Alibaba Qwen 2 72B instruction-tuned checkpoint. Provider: Alibaba. Released 2024-06-10. Access: Open.

Unified ELO 1550 ± 1, rank #434 of 1392 rated models, from 273 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HELM EWoK - Spatial Relations79.59EM100
HELM SeaHELM - IndoNLI86EM100
LLM Stats (TheoremQA)44.4Score (%)100
Open Chinese LLM - C-Eval Semantic94.11Accuracy (%)100
Open Chinese LLM - GSM8K78.7Accuracy (%)100
Open Chinese LLM Leaderboard74.88Average Score (%)100
OpenEval - BBQ96.5Exact Match (%)100
SeaEval - Cultural Reasoning - US-Eval (Zero-Shot)87.85Accuracy (%)100
SeaEval - Fundamental NLP Tasks - MRPC (Zero-Shot)80.64Accuracy (%)100
SeaEval - Fundamental NLP Tasks - OCNLI (Zero-Shot)78.2Accuracy (%)100
Thai LLM NLU - wisesight_thai_sentiment_seacrowd_text60.09Accuracy (%)100
Open Chinese LLM - CMMLU78.4Accuracy (%)99.4

Interactive version: theaggregate.ai/model?slug=qwen-2-72b-instruct · How It Works · Data refreshed daily, snapshot 2026-09-05.