Qwen 3 VL 235B A22B Instruct — benchmark results

Alibaba Qwen 3 VL 235B A22B vision-language instruction-tuned checkpoint. Provider: Alibaba. Released 2025-09-22. Access: Open.

Unified ELO 1634 ± 15, rank #406 of 1776 rated models, from 109 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
EmpathyBench51.3Average Score (%)100
Konkur 1404 - Humanities52.15Accuracy (%, text-only)100
Konkur 1404 - Mathematics60Accuracy (%, text-only)100
Konkur 1404 - Overall50.54Accuracy (%, text-only)100
LLM Stats (CharadesSTA)64.8Score (%)100
LLM Stats (DocVQAtest)97.1Score (%)100
Math-VR65Overall Answer Correctness (self-reported)96.7
LLM Stats (CC-OCR)82.2Score (%)94.1
CFMME63.31Average (self-reported)93.3
Konkur 1404 - Experimental Sciences46.58Accuracy (%, text-only)92.1
Konkur 1404 - Foreign Language40.38Accuracy (%, text-only)92.1
LLM Stats (OCRBench)92Score (%)90.5

Interactive version: theaggregate.ai/model?slug=qwen-3-vl-235b-a22b-instruct · How the rankings work · Data refreshed daily, snapshot 2026-07-22.