Qwen 3 VL 235B A22B (Thinking): benchmark results
Alibaba Qwen 3 VL 235B A22B vision-language model evaluated with thinking enabled. Provider: Alibaba. Released 2025-09-22. Access: Open.
Unified ELO 1576 ± 1, rank #529 of 1761 rated models, from 169 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| CFMME | 66.11 | Average (self-reported) | 100 |
| FlagEval EmbodiedVerse - EmbSpatial-Bench | 83.49 | Score | 100 |
| FlagEval EmbodiedVerse - RoboSpatial-Home | 69.19 | Score | 100 |
| LLM Stats (OCRBench-V2 (zh)) | 63.5 | Score (%) | 100 |
| LLM Stats (ZebraLogic) | 97.3 | Score (%) | 100 |
| Math-VR | 66.8 | Overall Answer Correctness (self-reported) | 100 |
| PCB-Bench - Task 3 BERTScore | 82.93 | BERTScore (%) | 100 |
| PCB-Bench - Task 3 SBERT | 61.68 | SBERT similarity (%) | 100 |
| K-MetBench | 84.4 | Accuracy (self-reported) | 96.6 |
| FlagEval EmbodiedVerse - All-Angles Bench | 61.12 | Score | 95.7 |
| FlagEval EmbodiedVerse - VSI-Bench (tiny) | 54.74 | Score | 95.7 |
| FlagEval EmbodiedVerse - Where2Place | 52.61 | Score | 91.3 |
Interactive version: theaggregate.ai/model?slug=qwen-3-vl-235b-a22b-thinking · How It Works · Data refreshed daily, snapshot 2026-09-05.