Qwen 3 VL 235B A22B (Thinking) — benchmark results
Alibaba Qwen 3 VL 235B A22B vision-language model evaluated with thinking enabled. Provider: Alibaba. Released 2025-09-22. Access: Open.
Unified ELO 1622 ± 14, rank #432 of 1776 rated models, from 137 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| CFMME | 66.11 | Average (self-reported) | 100 |
| LLM Stats (OCRBench-V2 (zh)) | 63.5 | Score (%) | 100 |
| LLM Stats (ZebraLogic) | 97.3 | Score (%) | 100 |
| Math-VR | 66.8 | Overall Answer Correctness (self-reported) | 100 |
| PCB-Bench - Task 3 BERTScore | 82.93 | BERTScore (%) | 100 |
| PCB-Bench - Task 3 SBERT | 61.68 | SBERT similarity (%) | 100 |
| K-MetBench | 84.4 | Accuracy (self-reported) | 96.6 |
| YapBench | 666.3 | YapIndex (lower is better) | 94.2 |
| LLM Stats (InfoVQAtest) | 89.5 | Score (%) | 90.9 |
| LLM Stats (MuirBench) | 80.1 | Score (%) | 90 |
| LLM Stats (Multi-IF) | 79.1 | Score (%) | 89.5 |
| VisuLogic | 34.4 | Overall Accuracy (%) | 87.5 |
Interactive version: theaggregate.ai/model?slug=qwen-3-vl-235b-a22b-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.