Qwen 3.5 35B A3B (Thinking): benchmark results

Provider: Alibaba. Released 2026-02-24. Access: Open.

Unified ELO 1630 ± 1, rank #446 of 3078 rated models, from 58 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Swallow - Japanese MT-Bench - Math99.3Judge Score (normalized, %)97
Swallow - Post-trained English - GPQA Diamond84Accuracy (%)93.3
Swallow - Post-trained Japanese - GPQA77.5Accuracy (%)92.5
Swallow - Post-trained English - MMLU-Pro85.7Accuracy (%)91.8
K-MetBench81.8Accuracy (self-reported)91.4
Swallow - English MT-Bench - STEM79.6Judge Score (normalized, %)91
Swallow - Post-trained Japanese - MMLU-ProX82.9Accuracy (%)91
Korean CSAT 2026 (Easy Mode) - English100Points (out of 100)90.5
Swallow - English MT-Bench - Reasoning88.4Judge Score (normalized, %)89.6
Swallow - Japanese MT-Bench - STEM76Judge Score (normalized, %)89.6
Swallow - Post-trained English - HellaSwag94Accuracy (%)89.6
Swallow - Post-trained English - MATH-50098.4Accuracy (%)89.6

Interactive version: theaggregate.ai/model?slug=qwen-3-5-35b-a3b-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.