Qwen 3 VL 8B (Thinking) — benchmark results

Alibaba Qwen 3 VL 8B vision-language model evaluated with thinking enabled. Provider: Alibaba. Released 2025-07-01. Access: Open.

Unified ELO 1543 ± 13, rank #644 of 1776 rated models, from 97 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
OCR-Robust79.67OCR1.0 Clean (self-reported)93.3
MedLayXPlain64.6S (self-reported)83.9
MathVision59.6Overall Accuracy (%)73.9
CFMME53.85Average (self-reported)73.3
K-MetBench71.7Accuracy (self-reported)72.4
LLM Stats (MuirBench)76.8Score (%)70
LLM Stats (WritingBench)85.5Score (%)67.9
LLM Stats (BLINK)68.7Score (%)66.7
AA Omniscience - Software Engineering (SWE) - Go24Accuracy (%)66.1
AA Omniscience - Software Engineering (SWE) - PHP32Accuracy (%)65
CritPt0.3Accuracy (self-reported)65
AA Omniscience - Software Engineering (SWE) - Swift44Accuracy (%)63.3

Interactive version: theaggregate.ai/model?slug=qwen-3-vl-8b-thinking · How the rankings work · Data refreshed daily, snapshot 2026-07-22.