Qwen 3.5 4B — benchmark results

Alibaba's open 4B from the Qwen3.5 small series (March 2026): natively multimodal with 262K context, built for on-device use. Provider: Alibaba. Released 2026-03-01. Access: Open.

Unified ELO 1491 ± 11, rank #832 of 1776 rated models, from 345 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
EuroEval Italian NLU - ScaLA IT57.06Linguistic acceptability Score (%)95
EuroEval Slovene Common Sense Reasoning57.96Common Sense Reasoning Average Score (%)93.5
EuroEval Catalan NLU - Guia CAT72.49Sentiment classification Score (%)93.4
EuroEval French NLU - ScaLA FR59.52Linguistic acceptability Score (%)92.8
JudgeBench Math93.75Accuracy (%)92.2
EuroEval Serbian Common Sense Reasoning53.85Common Sense Reasoning Average Score (%)92
EuroEval Portuguese NLU - ScaLA PT44.79Linguistic acceptability Score (%)91.8
EuroEval Croatian Common Sense Reasoning56.6Common Sense Reasoning Average Score (%)91.7
EuroEval Bulgarian Common Sense Reasoning56.57Common Sense Reasoning Average Score (%)91.2
EuroEval Slovak Common Sense Reasoning53.91Common Sense Reasoning Average Score (%)91
EuroEval Lithuanian Common Sense Reasoning57.28Common Sense Reasoning Average Score (%)90.5
EuroEval Catalan Common Sense Reasoning55.1Common Sense Reasoning Average Score (%)90.2

Interactive version: theaggregate.ai/model?slug=qwen-3-5-4b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.