Qwen 2 72B Instruct — benchmark results

Alibaba Qwen 2 72B instruction-tuned checkpoint. Provider: Alibaba. Released 2024-06-10. Access: Open.

Unified ELO 1553 ± 10, rank #611 of 1776 rated models, from 219 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
HELM EWoK - Spatial Relations79.59EM100
HELM SeaHELM - IndoNLI86EM100
LLM Stats (TheoremQA)44.4Score (%)100
Open Chinese LLM - C-Eval Semantic94.11Accuracy (%)100
Open Chinese LLM - GSM8K78.7Accuracy (%)100
Open Chinese LLM Leaderboard74.88Average Score (%)100
OpenEval - BBQ96.5Exact Match (%)100
SeaEval - Cultural Reasoning - US-Eval (Zero-Shot)87.85Accuracy (%)100
SeaEval - Fundamental NLP Tasks - MRPC (Zero-Shot)80.64Accuracy (%)100
SeaEval - Fundamental NLP Tasks - OCNLI (Zero-Shot)78.2Accuracy (%)100
Open Chinese LLM - CMMLU78.4Accuracy (%)99.4
Open LLM Leaderboard - BBH57.48Score98.9

Interactive version: theaggregate.ai/model?slug=qwen-2-72b-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.