Qwen 3.5 35B A3B (Thinking): benchmark results
Provider: Alibaba. Released 2026-02-24. Access: Open.
Unified ELO 1630 ± 1, rank #446 of 3078 rated models, from 58 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Swallow - Japanese MT-Bench - Math | 99.3 | Judge Score (normalized, %) | 97 |
| Swallow - Post-trained English - GPQA Diamond | 84 | Accuracy (%) | 93.3 |
| Swallow - Post-trained Japanese - GPQA | 77.5 | Accuracy (%) | 92.5 |
| Swallow - Post-trained English - MMLU-Pro | 85.7 | Accuracy (%) | 91.8 |
| K-MetBench | 81.8 | Accuracy (self-reported) | 91.4 |
| Swallow - English MT-Bench - STEM | 79.6 | Judge Score (normalized, %) | 91 |
| Swallow - Post-trained Japanese - MMLU-ProX | 82.9 | Accuracy (%) | 91 |
| Korean CSAT 2026 (Easy Mode) - English | 100 | Points (out of 100) | 90.5 |
| Swallow - English MT-Bench - Reasoning | 88.4 | Judge Score (normalized, %) | 89.6 |
| Swallow - Japanese MT-Bench - STEM | 76 | Judge Score (normalized, %) | 89.6 |
| Swallow - Post-trained English - HellaSwag | 94 | Accuracy (%) | 89.6 |
| Swallow - Post-trained English - MATH-500 | 98.4 | Accuracy (%) | 89.6 |
Interactive version: theaggregate.ai/model?slug=qwen-3-5-35b-a3b-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.