Qwen 3 VL 32B (Thinking) — benchmark results

Alibaba Qwen 3 VL 32B vision-language model evaluated with thinking enabled. Provider: Alibaba. Released 2025-07-01. Access: Open.

Unified ELO 1593 ± 21, rank #492 of 1776 rated models, from 83 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
LLM Stats (MuirBench)80.3Score (%)100
LLM Stats (OCRBench-V2 (en))68.4Score (%)100
MedLayXPlain65.4S (self-reported)93.5
LLM Stats (ScreenSpot)95.7Score (%)93.3
LLM Stats (OCRBench-V2 (zh))62.1Score (%)90
K-MetBench78.6Accuracy (self-reported)89.7
AA LiveCodeBench73.76Pass@1 (%)85.1
LLM Stats (Multi-IF)78Score (%)84.2
AA AIME 202584.67Accuracy (%)81.9
LLM Stats (CharXiv-D)90.2Score (%)80
LLM Stats (WritingBench)86.2Score (%)78.6
LLM Stats (InfoVQAtest)89.2Score (%)77.3

Interactive version: theaggregate.ai/model?slug=qwen-3-vl-32b-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.