Qwen 3.5 4B: benchmark results
Alibaba's open 4B from the Qwen3.5 small series (March 2026): natively multimodal with 262K context, built for on-device use. Provider: Alibaba. Released 2026-03-01. Access: Open.
Unified ELO 1521 ± 1, rank #592 of 1392 rated models, from 455 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| EuroEval Italian NLU - ScaLA IT | 57.06 | Linguistic acceptability Score (%) | 95 |
| EuroEval Slovene Common Sense Reasoning | 57.96 | Common Sense Reasoning Average Score (%) | 93.5 |
| EuroEval Catalan NLU - Guia CAT | 72.49 | Sentiment classification Score (%) | 93.4 |
| EuroEval French NLU - ScaLA FR | 59.52 | Linguistic acceptability Score (%) | 92.8 |
| JudgeBench Math | 93.75 | Accuracy (%) | 92.2 |
| EuroEval Serbian Common Sense Reasoning | 53.85 | Common Sense Reasoning Average Score (%) | 92 |
| EuroEval Portuguese NLU - ScaLA PT | 44.79 | Linguistic acceptability Score (%) | 91.8 |
| EuroEval Croatian Common Sense Reasoning | 56.6 | Common Sense Reasoning Average Score (%) | 91.7 |
| EuroEval Bulgarian Common Sense Reasoning | 56.57 | Common Sense Reasoning Average Score (%) | 91.2 |
| EuroEval Slovak Common Sense Reasoning | 53.91 | Common Sense Reasoning Average Score (%) | 91 |
| SVFSearch | 87.9 | Overall Acc. (self-reported) | 90.8 |
| EuroEval Lithuanian Common Sense Reasoning | 57.28 | Common Sense Reasoning Average Score (%) | 90.5 |
Interactive version: theaggregate.ai/model?slug=qwen-3-5-4b · How It Works · Data refreshed daily, snapshot 2026-09-05.