Qwen 3.5 397B A17B (Thinking): benchmark results
Provider: Alibaba. Released 2026-02-16. Access: Open.
Unified ELO 1665 ± 1, rank #252 of 3078 rated models, from 120 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Nejumi 4 - Toxicity - Prohibited Acts | 100 | Criteria met (%) | 100 |
| Nejumi 4 - BFCL - Live AST | 76.85 | Accuracy (%) | 98.9 |
| BenchTable - Tech | 85.5 | Weighted Score (%) | 94.8 |
| BenchTable - Reasoning | 81.7 | Weighted Score (%) | 94.2 |
| Nejumi 4 - jaster (2-shot) - JCoLA (in-domain) | 82 | Exact match (%) | 93.8 |
| Nejumi 4 - jaster (2-shot) - MMLU-ProX (Japanese) | 90 | Exact match (%) | 91.9 |
| SuperCLUE General (March 2026) - Hallucination Control | 84.39 | Score | 91.3 |
| Sonar LLM Leaderboard - Java - Bug Density | 0.61 | Bugs per 1,000 lines of code (lower is better) | 91 |
| Nejumi 4 - jaster (0-shot) - JNLI | 89 | Exact match (%) | 90.8 |
| AGI-Eval Community - Learning | 90.99 | Accuracy (%) | 90.7 |
| Nejumi 4 - jaster (2-shot) - JNLI | 91 | Exact match (%) | 90.1 |
| AGI-Eval Community - Learning (English) | 95.38 | Accuracy (%) | 89.9 |
Interactive version: theaggregate.ai/model?slug=qwen-3-5-397b-a17b-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.