Qwen 2 72B Instruct: benchmark results
Alibaba Qwen 2 72B instruction-tuned checkpoint. Provider: Alibaba. Released 2024-06-10. Access: Open.
Unified ELO 1550 ± 1, rank #434 of 1392 rated models, from 273 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM EWoK - Spatial Relations | 79.59 | EM | 100 |
| HELM SeaHELM - IndoNLI | 86 | EM | 100 |
| LLM Stats (TheoremQA) | 44.4 | Score (%) | 100 |
| Open Chinese LLM - C-Eval Semantic | 94.11 | Accuracy (%) | 100 |
| Open Chinese LLM - GSM8K | 78.7 | Accuracy (%) | 100 |
| Open Chinese LLM Leaderboard | 74.88 | Average Score (%) | 100 |
| OpenEval - BBQ | 96.5 | Exact Match (%) | 100 |
| SeaEval - Cultural Reasoning - US-Eval (Zero-Shot) | 87.85 | Accuracy (%) | 100 |
| SeaEval - Fundamental NLP Tasks - MRPC (Zero-Shot) | 80.64 | Accuracy (%) | 100 |
| SeaEval - Fundamental NLP Tasks - OCNLI (Zero-Shot) | 78.2 | Accuracy (%) | 100 |
| Thai LLM NLU - wisesight_thai_sentiment_seacrowd_text | 60.09 | Accuracy (%) | 100 |
| Open Chinese LLM - CMMLU | 78.4 | Accuracy (%) | 99.4 |
Interactive version: theaggregate.ai/model?slug=qwen-2-72b-instruct · How It Works · Data refreshed daily, snapshot 2026-09-05.