Qwen 3.5 122B A10B (Thinking): benchmark results
Provider: Alibaba. Released 2026-02-24. Access: Open.
Unified ELO 1650 ± 1, rank #342 of 3078 rated models, from 129 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Nejumi 4 - BFCL - Live AST | 81.48 | Accuracy (%) | 100 |
| Swallow - Post-trained English - MMLU-Pro | 87 | Accuracy (%) | 98.5 |
| Swallow - Post-trained Japanese - GPQA | 80.1 | Accuracy (%) | 98.5 |
| Swallow - Post-trained English - GPQA Diamond | 86.2 | Accuracy (%) | 97 |
| Nejumi 4 - Toxicity - Prohibited Acts | 98.44 | Criteria met (%) | 96.3 |
| Swallow - Post-trained Japanese - MMLU-ProX | 84.1 | Accuracy (%) | 96.3 |
| Nejumi 4 - jaster (2-shot) - JaMP | 82 | Exact match (%) | 96 |
| Swallow - Japanese MT-Bench - Math | 99.2 | Judge Score (normalized, %) | 95.5 |
| Swallow - English MT-Bench - Reasoning | 89.3 | Judge Score (normalized, %) | 94 |
| Swallow - English MT-Bench - STEM | 80.8 | Judge Score (normalized, %) | 94 |
| Swallow - Post-trained English - HellaSwag | 95.1 | Accuracy (%) | 94 |
| Swallow - Japanese MT-Bench - STEM | 77.1 | Judge Score (normalized, %) | 92.5 |
Interactive version: theaggregate.ai/model?slug=qwen-3-5-122b-a10b-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.