Qwen 3 VL 4B Instruct — benchmark results

Alibaba's Apache-2.0 4B dense vision-language model (October 2025) with 256K context, the compact tier of the Qwen3-VL line. Provider: Alibaba. Released 2025-07-01. Access: Open.

Unified ELO 1467 ± 13, rank #921 of 1776 rated models, from 104 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
LatamBoard - ASSIN2 RTE94.08Score (%)93.9
LatamBoard - FaQuAD NLI81.37Score (%)93.9
LatamBoard - Portuguese Score89.58Score (%)93.9
LatamBoard - BLUEX67.59Score (%)90.9
LatamBoard - ENEM Challenge76.35Score (%)90.9
LatamBoard - OAB Exams53.39Score (%)90.9
LatamBoard - Spanish WNLI76.06Score (%)83.3
UGI - Willingness (W/10)7.8W/10 Score81.1
LLM Stats (ODinW)48.2Score (%)73.3
LatamBoard - Spanish XNLI47.35Score (%)66.7
TriViewBench52Overall (self-reported)66.7
LatamBoard - Spanish TeleIA66.67Score (%)62.1

Interactive version: theaggregate.ai/model?slug=qwen-3-vl-4b-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.