Qwen 3.5 9B — benchmark results
Alibaba's open 9B dense model of the Qwen3.5 small series: hybrid linear attention, native multimodality, 262K context. Provider: Alibaba. Released 2026-03-02. Access: Open.
Unified ELO 1552 ± 10, rank #615 of 1776 rated models, from 412 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| EuroEval Catalan NLU - Guia CAT | 77.18 | Sentiment classification Score (%) | 100 |
| LithoBench | 98.25 | MCQ$_{\mathrm{all}}$ (self-reported) | 100 |
| EuroEval Italian NLU - ScaLA IT | 62.14 | Linguistic acceptability Score (%) | 97 |
| EuroEval Spanish NLU - ScaLA ES | 52.26 | Linguistic acceptability Score (%) | 96.6 |
| EuroEval Swedish NLU - Swerec | 80.13 | Sentiment classification Score (%) | 96.1 |
| EuroEval Norwegian Common Sense Reasoning | 88.1 | Common Sense Reasoning Average Score (%) | 96 |
| Managing Procedural Memory in LLM Agents | 3 | size (self-reported) | 95.8 |
| EuroEval Hungarian Common Sense Reasoning | 64.84 | Common Sense Reasoning Average Score (%) | 95.7 |
| EuroEval Slovene Common Sense Reasoning | 61.81 | Common Sense Reasoning Average Score (%) | 95.7 |
| EuroEval Dutch NLU - DBRD | 92.67 | Sentiment classification Score (%) | 95.5 |
| EuroEval Portuguese NLU - ScaLA PT | 56.24 | Linguistic acceptability Score (%) | 95.5 |
| EuroEval Croatian Common Sense Reasoning | 66.6 | Common Sense Reasoning Average Score (%) | 95.3 |
Interactive version: theaggregate.ai/model?slug=qwen-3-5-9b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.