Qwen 3 VL 235B A22B Instruct: benchmark results

Alibaba Qwen 3 VL 235B A22B vision-language instruction-tuned checkpoint. Provider: Alibaba. Released 2025-09-22. Access: Open.

Unified ELO 1606 ± 1, rank #245 of 1392 rated models, from 154 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
EmpathyBench51.3Average Score (%)100
Konkur 1404 - Humanities52.15Accuracy (%, text-only)100
Konkur 1404 - Mathematics60Accuracy (%, text-only)100
Konkur 1404 - Overall50.54Accuracy (%, text-only)100
LLM Stats (CharadesSTA)64.8Score (%)100
LLM Stats (DocVQAtest)97.1Score (%)100
Math-VR65Overall Answer Correctness (self-reported)96.6
FlagEval EmbodiedVerse - CV-Bench (test)88.78Score95.7
FlagEval EmbodiedVerse - OmniSpatial53.85Score95.7
FlagEval EmbodiedVerse - Where2Place58.35Score95.7
LLM Stats (CC-OCR)82.2Score (%)94.1
CFMME63.31Average (self-reported)93.3

Interactive version: theaggregate.ai/model?slug=qwen-3-vl-235b-a22b-instruct · How It Works · Data refreshed daily, snapshot 2026-09-05.