Qwen 3 235B A22B Instruct: benchmark results
Provider: Alibaba. Released 2025-04-29. Access: Open.
Unified ELO 1725 ± 21, rank #283 of 2656 rated models, from 27 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| EComStage | 85.61 | Task Score (%) | 100 |
| EComStage - Solution Decision | 89.94 | Task Score (%) | 96.9 |
| EComStage - Query Match | 98.75 | Task Score (%) | 93.8 |
| EComStage - Attitude Classification | 87.26 | Task Score (%) | 92.2 |
| EComStage - Query Rewrite | 81.04 | Task Score (%) | 87.5 |
| EComStage - RAG-QA | 69.76 | Task Score (%) | 84.4 |
| WuYuEval - Expert Module - Judge Score | 0.85 | Mean anchor-calibrated LLM-as-a-judge score (0-1) | 84.4 |
| EComStage - Scenario Route | 84.76 | Task Score (%) | 73.4 |
| EComStage - Intent Recognition | 87.78 | Task Score (%) | 71.9 |
| WuYuEval - Expert Module - Elo | 1723.93 | Elo rating from pairwise judged comparisons | 71.9 |
| COMPASS Policy - Allowed Edge | 90.01 | Policy Alignment Score (%) | 71.4 |
| WuYuEval - Foundation Module - Accuracy | 89.63 | Accuracy (%) | 68.8 |
Interactive version: theaggregate.ai/model?slug=qwen-3-235b-a22b-instruct · How It Works · Data refreshed daily, snapshot 2026-09-19.