Qwen 2.5 VL 72B Instruct: benchmark results
Alibaba Qwen 2.5 VL 72B vision-language instruction-tuned checkpoint. Provider: Alibaba. Released 2025-01-26. Access: Open.
Unified ELO 1665 ± 12, rank #415 of 2656 rated models, from 136 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| GB/T-Bench - Diagnosis Recall - Normative References | 60.4 | Diagnosis recall (%; exact section, dimension and error-type | 100 |
| LLM Stats (DocVQA) | 96.4 | Score (%) | 100 |
| StickToYourRole | 83.6 | Cardinal Score | 100 |
| Open Portuguese LLM - BLUEX | 80.53 | Accuracy (%) | 98.7 |
| Open Portuguese LLM - ASSIN2 RTE | 94.73 | Macro F1 (%) | 98.6 |
| Open Portuguese LLM - ENEM | 86 | Accuracy (%) | 98.4 |
| Open Portuguese LLM - OAB Exams | 68.88 | Accuracy (%) | 98.1 |
| Open Portuguese LLM - FaQuAD NLI | 84.47 | Macro F1 (%) | 98 |
| MERA Multi - WEIRD | 70.3 | Score (%) | 94.9 |
| Open Portuguese LLM - Hate Speech | 76.96 | F1 (%) | 93.9 |
| KOFFVQA - Korean OCR | 95 | Score (%) | 93.8 |
| LatamBoard | 87.06 | LatamBoard Category Average (percentage) | 92.3 |
Interactive version: theaggregate.ai/model?slug=qwen-2-5-vl-72b-instruct · How It Works · Data refreshed daily, snapshot 2026-09-19.