Qwen 2.5 VL 32B Instruct — benchmark results

Alibaba Qwen 2.5 VL 32B vision-language instruction-tuned checkpoint. Provider: Alibaba. Released 2025-03-24. Access: Open.

Unified ELO 1498 ± 17, rank #805 of 1776 rated models, from 65 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
EVALITA - MAIA-MC87.26CPS100
EVALITA - word-in-context77.56CPS100
LLM Stats (InfoVQA)83.4Score (%)100
KOFFVQA - Hallucination and Robustness86Score (%)94.4
Open Portuguese LLM - Hate Speech76.96F1 (%)93.9
EVALITA - text-entailment83.7CPS93.6
EVALITA - hate-speech-detection75.47CPS87.2
EVALITA LLM Leaderboard57.26Average CPS87.2
EVALITA - lexical-substitution40.67CPS80.9
LLM Stats (DocVQA)94.8Score (%)80
KOFFVQA - Korean Recognition60Score (%)76.5
KOFFVQA - Table Understanding80.33Score (%)75.3

Interactive version: theaggregate.ai/model?slug=qwen-2-5-vl-32b-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.