Qwen 3.5 397B A17B (Thinking): benchmark results

Provider: Alibaba. Released 2026-02-16. Access: Open.

Unified ELO 1665 ± 1, rank #252 of 3078 rated models, from 120 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Nejumi 4 - Toxicity - Prohibited Acts100Criteria met (%)100
Nejumi 4 - BFCL - Live AST76.85Accuracy (%)98.9
BenchTable - Tech85.5Weighted Score (%)94.8
BenchTable - Reasoning81.7Weighted Score (%)94.2
Nejumi 4 - jaster (2-shot) - JCoLA (in-domain)82Exact match (%)93.8
Nejumi 4 - jaster (2-shot) - MMLU-ProX (Japanese)90Exact match (%)91.9
SuperCLUE General (March 2026) - Hallucination Control84.39Score91.3
Sonar LLM Leaderboard - Java - Bug Density0.61Bugs per 1,000 lines of code (lower is better)91
Nejumi 4 - jaster (0-shot) - JNLI89Exact match (%)90.8
AGI-Eval Community - Learning90.99Accuracy (%)90.7
Nejumi 4 - jaster (2-shot) - JNLI91Exact match (%)90.1
AGI-Eval Community - Learning (English)95.38Accuracy (%)89.9

Interactive version: theaggregate.ai/model?slug=qwen-3-5-397b-a17b-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.