Qwen-72B: benchmark results
Provider: Alibaba. Released 2023-11-26. Access: Open.
Unified ELO 1588 ± 30, rank #677 of 2656 rated models, from 13 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| CFLUE | 72.8 | Accuracy | 100 |
| Open Portuguese LLM - OAB Exams | 60.32 | Accuracy (%) | 92.6 |
| Open Portuguese LLM - ENEM | 78.24 | Accuracy (%) | 91.9 |
| Open Portuguese LLM - BLUEX | 68.43 | Accuracy (%) | 91.2 |
| InfiBench | 55.34 | Score (%) | 85.7 |
| Open Portuguese LLM - ASSIN2 RTE | 92.88 | Macro F1 (%) | 84.7 |
| C-Eval | 82.8 | Average (%) | 83.8 |
| Open Portuguese LLM - FaQuAD NLI | 78.4 | Macro F1 (%) | 83.7 |
| T-Eval | 71.4 | Overall Score (%) | 80 |
| pfgen-bench - Completion Mode - Helpfulness | 0.15 | Helpfulness Score | 56.9 |
| pfgen-bench - Completion Mode - Fluency | 0.59 | Fluency Score | 55.6 |
| pfgen-bench - Completion Mode - Score | 0.49 | pfgen Score (mean of three) | 53.7 |
Interactive version: theaggregate.ai/model?slug=qwen-72b · How It Works · Data refreshed daily, snapshot 2026-09-19.