Qwen 3.5 122B A10B (Thinking): benchmark results

Provider: Alibaba. Released 2026-02-24. Access: Open.

Unified ELO 1650 ± 1, rank #342 of 3078 rated models, from 129 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Nejumi 4 - BFCL - Live AST81.48Accuracy (%)100
Swallow - Post-trained English - MMLU-Pro87Accuracy (%)98.5
Swallow - Post-trained Japanese - GPQA80.1Accuracy (%)98.5
Swallow - Post-trained English - GPQA Diamond86.2Accuracy (%)97
Nejumi 4 - Toxicity - Prohibited Acts98.44Criteria met (%)96.3
Swallow - Post-trained Japanese - MMLU-ProX84.1Accuracy (%)96.3
Nejumi 4 - jaster (2-shot) - JaMP82Exact match (%)96
Swallow - Japanese MT-Bench - Math99.2Judge Score (normalized, %)95.5
Swallow - English MT-Bench - Reasoning89.3Judge Score (normalized, %)94
Swallow - English MT-Bench - STEM80.8Judge Score (normalized, %)94
Swallow - Post-trained English - HellaSwag95.1Accuracy (%)94
Swallow - Japanese MT-Bench - STEM77.1Judge Score (normalized, %)92.5

Interactive version: theaggregate.ai/model?slug=qwen-3-5-122b-a10b-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.