Qwen 2 72B Instruct — benchmark results
Alibaba Qwen 2 72B instruction-tuned checkpoint. Provider: Alibaba. Released 2024-06-10. Access: Open.
Unified ELO 1553 ± 10, rank #611 of 1776 rated models, from 219 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM EWoK - Spatial Relations | 79.59 | EM | 100 |
| HELM SeaHELM - IndoNLI | 86 | EM | 100 |
| LLM Stats (TheoremQA) | 44.4 | Score (%) | 100 |
| Open Chinese LLM - C-Eval Semantic | 94.11 | Accuracy (%) | 100 |
| Open Chinese LLM - GSM8K | 78.7 | Accuracy (%) | 100 |
| Open Chinese LLM Leaderboard | 74.88 | Average Score (%) | 100 |
| OpenEval - BBQ | 96.5 | Exact Match (%) | 100 |
| SeaEval - Cultural Reasoning - US-Eval (Zero-Shot) | 87.85 | Accuracy (%) | 100 |
| SeaEval - Fundamental NLP Tasks - MRPC (Zero-Shot) | 80.64 | Accuracy (%) | 100 |
| SeaEval - Fundamental NLP Tasks - OCNLI (Zero-Shot) | 78.2 | Accuracy (%) | 100 |
| Open Chinese LLM - CMMLU | 78.4 | Accuracy (%) | 99.4 |
| Open LLM Leaderboard - BBH | 57.48 | Score | 98.9 |
Interactive version: theaggregate.ai/model?slug=qwen-2-72b-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.