Qwen 3 VL 32B (Thinking): benchmark results

Alibaba Qwen 3 VL 32B vision-language model evaluated with thinking enabled. Provider: Alibaba. Released 2025-07-01. Access: Open.

Unified ELO 1551 ± 1, rank #632 of 1761 rated models, from 82 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
LLM Stats (MuirBench)80.3Score (%)100
LLM Stats (OCRBench-V2 (en))68.4Score (%)100
MedLayXPlain65.4S (self-reported)93.5
LLM Stats (ScreenSpot)95.7Score (%)93.3
LLM Stats (OCRBench-V2 (zh))62.1Score (%)90
K-MetBench78.6Accuracy (self-reported)89.7
AA AIME 202584.67Accuracy (%)82
LLM Stats (Multi-IF)78Score (%)81.8
LLM Stats (CharXiv-D)90.2Score (%)81.2
AA LiveCodeBench73.76Pass@1 (%)80
LLM Stats (MMStar)79.4Score (%)79.2
LLM Stats (WritingBench)86.2Score (%)78.6

Interactive version: theaggregate.ai/model?slug=qwen-3-vl-32b-thinking · How It Works · Data refreshed daily, snapshot 2026-09-05.