Qwen 3.5 4B: benchmark results

Alibaba's open 4B from the Qwen3.5 small series (March 2026): natively multimodal with 262K context, built for on-device use. Provider: Alibaba. Released 2026-03-01. Access: Open.

Unified ELO 1521 ± 1, rank #592 of 1392 rated models, from 455 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
EuroEval Italian NLU - ScaLA IT57.06Linguistic acceptability Score (%)95
EuroEval Slovene Common Sense Reasoning57.96Common Sense Reasoning Average Score (%)93.5
EuroEval Catalan NLU - Guia CAT72.49Sentiment classification Score (%)93.4
EuroEval French NLU - ScaLA FR59.52Linguistic acceptability Score (%)92.8
JudgeBench Math93.75Accuracy (%)92.2
EuroEval Serbian Common Sense Reasoning53.85Common Sense Reasoning Average Score (%)92
EuroEval Portuguese NLU - ScaLA PT44.79Linguistic acceptability Score (%)91.8
EuroEval Croatian Common Sense Reasoning56.6Common Sense Reasoning Average Score (%)91.7
EuroEval Bulgarian Common Sense Reasoning56.57Common Sense Reasoning Average Score (%)91.2
EuroEval Slovak Common Sense Reasoning53.91Common Sense Reasoning Average Score (%)91
SVFSearch87.9Overall Acc. (self-reported)90.8
EuroEval Lithuanian Common Sense Reasoning57.28Common Sense Reasoning Average Score (%)90.5

Interactive version: theaggregate.ai/model?slug=qwen-3-5-4b · How It Works · Data refreshed daily, snapshot 2026-09-05.