Qwen 3 VL 30B A3B Instruct: benchmark results
Alibaba's open 30B A3B MoE vision-language model with 256K context, OCR in 32 languages, and GUI-agent skills. Provider: Alibaba. Released 2025-07-01. Access: Open.
Unified ELO 1526 ± 1, rank #563 of 1392 rated models, from 101 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| FlagEval EmbodiedVerse - EgoPlan-Bench2 | 58.8 | Score | 100 |
| Konkur 1404 - Experimental Sciences | 50.31 | Accuracy (%, text-only) | 100 |
| LLM Stats (CharadesSTA) | 63.5 | Score (%) | 86.4 |
| LLM Stats (MLVU-M) | 81.3 | Score (%) | 85.7 |
| Konkur 1404 - Mathematics | 46.96 | Accuracy (%, text-only) | 84.2 |
| FlagEval EmbodiedVerse - CV-Bench (test) | 86.78 | Score | 82.6 |
| LLM Stats (OCRBench) | 90.3 | Score (%) | 82.6 |
| Konkur 1404 - Overall | 43.64 | Accuracy (%, text-only) | 78.9 |
| AA Omniscience - Software Engineering (SWE) - Swift | 40 | Accuracy (%) | 70.1 |
| LLM Stats (ScreenSpot) | 94.7 | Score (%) | 70 |
| FlagEval EmbodiedVerse - EmbSpatial-Bench | 76.29 | Score | 69.6 |
| FlagEval EmbodiedVerse - VSI-Bench (tiny) | 47.49 | Score | 69.6 |
Interactive version: theaggregate.ai/model?slug=qwen-3-vl-30b-a3b-instruct · How It Works · Data refreshed daily, snapshot 2026-09-05.