Qwen 3.5 9B — benchmark results

Alibaba's open 9B dense model of the Qwen3.5 small series: hybrid linear attention, native multimodality, 262K context. Provider: Alibaba. Released 2026-03-02. Access: Open.

Unified ELO 1552 ± 10, rank #615 of 1776 rated models, from 412 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
EuroEval Catalan NLU - Guia CAT77.18Sentiment classification Score (%)100
LithoBench98.25MCQ$_{\mathrm{all}}$ (self-reported)100
EuroEval Italian NLU - ScaLA IT62.14Linguistic acceptability Score (%)97
EuroEval Spanish NLU - ScaLA ES52.26Linguistic acceptability Score (%)96.6
EuroEval Swedish NLU - Swerec80.13Sentiment classification Score (%)96.1
EuroEval Norwegian Common Sense Reasoning88.1Common Sense Reasoning Average Score (%)96
Managing Procedural Memory in LLM Agents3size (self-reported)95.8
EuroEval Hungarian Common Sense Reasoning64.84Common Sense Reasoning Average Score (%)95.7
EuroEval Slovene Common Sense Reasoning61.81Common Sense Reasoning Average Score (%)95.7
EuroEval Dutch NLU - DBRD92.67Sentiment classification Score (%)95.5
EuroEval Portuguese NLU - ScaLA PT56.24Linguistic acceptability Score (%)95.5
EuroEval Croatian Common Sense Reasoning66.6Common Sense Reasoning Average Score (%)95.3

Interactive version: theaggregate.ai/model?slug=qwen-3-5-9b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.