Qwen 3.5 9B (Thinking): benchmark results
Provider: Alibaba. Released 2026-03-02. Access: Open.
Unified ELO 1581 ± 1, rank #738 of 3078 rated models, from 122 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Swallow - Japanese MT-Bench - Reasoning | 84.3 | Judge Score (normalized, %) | 95.5 |
| Swallow - Post-trained English - GPQA Diamond | 82.7 | Accuracy (%) | 91 |
| Nejumi 4 - BFCL - Live AST | 74.07 | Accuracy (%) | 89.3 |
| Nejumi 4 - JHumanEval | 42.46 | Sandbox pass rate (%) | 84.6 |
| FrameBench - Frame Identification - English | 81.3 | Accuracy (%; FrameNet candidate frames) | 84.2 |
| Swallow - Post-trained English - MMLU-Pro | 82.6 | Accuracy (%) | 83.6 |
| Swallow - Post-trained Japanese - GPQA | 72.5 | Accuracy (%) | 83.6 |
| Nejumi 4 - jaster (2-shot) - JaMP | 79 | Exact match (%) | 81.2 |
| Swallow - Post-trained Japanese - MMLU-ProX | 78.4 | Accuracy (%) | 80.6 |
| K-MetBench | 74.9 | Accuracy (self-reported) | 78.4 |
| Swallow - Post-trained English - HellaSwag | 90.9 | Accuracy (%) | 77.6 |
| Korean CSAT 2026 (Easy Mode) - Physics I | 41 | Points (out of 50) | 76.9 |
Interactive version: theaggregate.ai/model?slug=qwen-3-5-9b-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.