Qwen 3 235B A22B FP8 (Thinking): benchmark results
Provider: Alibaba. Released 2025-04-29. Access: Open.
Unified ELO 1627 ± 1, rank #466 of 3078 rated models, from 29 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| BenchTable - STEM | 83.3 | Weighted Score (%) | 94.9 |
| BenchTable | 67.6 | Total Score (%) | 83.4 |
| LMGame-Bench Sokoban | 5.3 | Score | 81.2 |
| BenchTable - Reasoning | 64.9 | Weighted Score (%) | 80.2 |
| BenchTable - Tech | 70.8 | Weighted Score (%) | 71.9 |
| AGI-Eval Community - Learning (English) | 92.67 | Accuracy (%) | 71 |
| AGI-Eval Community - Learning | 87.55 | Accuracy (%) | 69.3 |
| BenchTable - Utility | 61.7 | Weighted Score (%) | 65.6 |
| AGI-Eval Community - Learning (Chinese) | 80.67 | Accuracy (%) | 63 |
| LMGame-Bench Candy Crush | 437 | Score | 58.3 |
| AGI-Eval Community - Subject Reasoning (Chinese) | 82.92 | Accuracy (%) | 52.9 |
| AGI-Eval Community - General Reasoning | 96.22 | Accuracy (%) | 51.4 |
Interactive version: theaggregate.ai/model?slug=qwen-3-235b-a22b-fp8-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.