Qwen 3 235B A22B Instruct: benchmark results

Provider: Alibaba. Released 2025-04-29. Access: Open.

Unified ELO 1725 ± 21, rank #283 of 2656 rated models, from 27 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
EComStage85.61Task Score (%)100
EComStage - Solution Decision89.94Task Score (%)96.9
EComStage - Query Match98.75Task Score (%)93.8
EComStage - Attitude Classification87.26Task Score (%)92.2
EComStage - Query Rewrite81.04Task Score (%)87.5
EComStage - RAG-QA69.76Task Score (%)84.4
WuYuEval - Expert Module - Judge Score0.85Mean anchor-calibrated LLM-as-a-judge score (0-1)84.4
EComStage - Scenario Route84.76Task Score (%)73.4
EComStage - Intent Recognition87.78Task Score (%)71.9
WuYuEval - Expert Module - Elo1723.93Elo rating from pairwise judged comparisons71.9
COMPASS Policy - Allowed Edge90.01Policy Alignment Score (%)71.4
WuYuEval - Foundation Module - Accuracy89.63Accuracy (%)68.8

Interactive version: theaggregate.ai/model?slug=qwen-3-235b-a22b-instruct · How It Works · Data refreshed daily, snapshot 2026-09-19.