Qwen 3.8 Flash: benchmark results

Provider: Alibaba. Released 2026-08-26. Access: API.

Unified ELO 1731 ± 1, rank #25 of 1392 rated models, from 28 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
RealWorldQA88.5RealWorldQA (self-reported)99.3
LLM Stats (MathVision)95.7Score (%)95.6
LLM Stats Score49.19LLM Stats Score (conservative rating)94.6
LLM Stats (ERQA)72.3Score (%)94
AI Chess Leaderboard (Reasoning)1528Elo93.8
LLM Stats (CharXiv-R)90.6Score (%)93.5
ZeroEval GPQA Diamond91.7GPQA Diamond Score92.3
LLM Stats (ClawEval-MM)60.4Score (%)90
LLM Stats (Toolathlon)73.5Score (%)88.8
LVBench76.6Score (self-reported)84.9
ComplexConstraints43.3Score (%)81.7
LLM Stats (Agents' Last Exam)51.2Score (%)76.7

Interactive version: theaggregate.ai/model?slug=qwen-3-8-flash · How It Works · Data refreshed daily, snapshot 2026-09-05.