Qwen 2.5 VL 32B Instruct: benchmark results
Alibaba Qwen 2.5 VL 32B vision-language instruction-tuned checkpoint. Provider: Alibaba. Released 2025-03-24. Access: Open.
Unified ELO 1500 ± 1, rank #688 of 1392 rated models, from 79 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| EVALITA - MAIA-MC | 87.26 | CPS | 100 |
| EVALITA - word-in-context | 77.56 | CPS | 100 |
| LLM Stats (InfoVQA) | 83.4 | Score (%) | 100 |
| KOFFVQA - Hallucination and Robustness | 86 | Score (%) | 94.4 |
| Open Portuguese LLM - Hate Speech | 76.96 | F1 (%) | 93.9 |
| EVALITA - text-entailment | 83.7 | CPS | 93.6 |
| EVALITA - hate-speech-detection | 75.47 | CPS | 87.2 |
| EVALITA LLM Leaderboard | 57.26 | Average CPS | 87.2 |
| FlagEval EmbodiedVerse - EgoPlan-Bench2 | 53.37 | Score | 82.6 |
| LLM Stats (DocVQA) | 94.8 | Score (%) | 81.5 |
| EVALITA - lexical-substitution | 40.67 | CPS | 80.9 |
| KOFFVQA - Korean Recognition | 60 | Score (%) | 76.5 |
Interactive version: theaggregate.ai/model?slug=qwen-2-5-vl-32b-instruct · How It Works · Data refreshed daily, snapshot 2026-09-05.