Qwen 2.5 VL 32B Instruct: benchmark results

Alibaba Qwen 2.5 VL 32B vision-language instruction-tuned checkpoint. Provider: Alibaba. Released 2025-03-24. Access: Open.

Unified ELO 1500 ± 1, rank #688 of 1392 rated models, from 79 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
EVALITA - MAIA-MC87.26CPS100
EVALITA - word-in-context77.56CPS100
LLM Stats (InfoVQA)83.4Score (%)100
KOFFVQA - Hallucination and Robustness86Score (%)94.4
Open Portuguese LLM - Hate Speech76.96F1 (%)93.9
EVALITA - text-entailment83.7CPS93.6
EVALITA - hate-speech-detection75.47CPS87.2
EVALITA LLM Leaderboard57.26Average CPS87.2
FlagEval EmbodiedVerse - EgoPlan-Bench253.37Score82.6
LLM Stats (DocVQA)94.8Score (%)81.5
EVALITA - lexical-substitution40.67CPS80.9
KOFFVQA - Korean Recognition60Score (%)76.5

Interactive version: theaggregate.ai/model?slug=qwen-2-5-vl-32b-instruct · How It Works · Data refreshed daily, snapshot 2026-09-05.