Qwen 3 VL 4B (Thinking) — benchmark results

Alibaba Qwen 3 VL 4B vision-language model evaluated with thinking enabled. Provider: Alibaba. Released 2025-07-01. Access: Open.

Unified ELO 1495 ± 18, rank #817 of 1776 rated models, from 85 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
OCR-Robust78.63OCR1.0 Clean (self-reported)86.7
MedLayXPlain63.3S (self-reported)74.2
K-MetBench66.1Accuracy (self-reported)62.1
LLM Stats (MuirBench)75Score (%)60
LLM Stats (MMStar)73.2Score (%)47.6
LLM Stats (Multi-IF)73.6Score (%)47.4
LLM Stats (Hallusion Bench)64.1Score (%)46.7
LLM Stats (MLVU-M)75.7Score (%)42.9
AA-LCR21.3Score (self-reported)41.2
LLM Stats (MathVista-Mini)79.5Score (%)40.9
LLM Stats (ScreenSpot)92.9Score (%)40
AA Omniscience - Software Engineering (SWE) - Go14Accuracy (%)39.6

Interactive version: theaggregate.ai/model?slug=qwen-3-vl-4b-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.