Qwen 3.5 4B — benchmark results
Alibaba's open 4B from the Qwen3.5 small series (March 2026): natively multimodal with 262K context, built for on-device use. Provider: Alibaba. Released 2026-03-01. Access: Open.
Unified ELO 1491 ± 11, rank #832 of 1776 rated models, from 345 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| EuroEval Italian NLU - ScaLA IT | 57.06 | Linguistic acceptability Score (%) | 95 |
| EuroEval Slovene Common Sense Reasoning | 57.96 | Common Sense Reasoning Average Score (%) | 93.5 |
| EuroEval Catalan NLU - Guia CAT | 72.49 | Sentiment classification Score (%) | 93.4 |
| EuroEval French NLU - ScaLA FR | 59.52 | Linguistic acceptability Score (%) | 92.8 |
| JudgeBench Math | 93.75 | Accuracy (%) | 92.2 |
| EuroEval Serbian Common Sense Reasoning | 53.85 | Common Sense Reasoning Average Score (%) | 92 |
| EuroEval Portuguese NLU - ScaLA PT | 44.79 | Linguistic acceptability Score (%) | 91.8 |
| EuroEval Croatian Common Sense Reasoning | 56.6 | Common Sense Reasoning Average Score (%) | 91.7 |
| EuroEval Bulgarian Common Sense Reasoning | 56.57 | Common Sense Reasoning Average Score (%) | 91.2 |
| EuroEval Slovak Common Sense Reasoning | 53.91 | Common Sense Reasoning Average Score (%) | 91 |
| EuroEval Lithuanian Common Sense Reasoning | 57.28 | Common Sense Reasoning Average Score (%) | 90.5 |
| EuroEval Catalan Common Sense Reasoning | 55.1 | Common Sense Reasoning Average Score (%) | 90.2 |
Interactive version: theaggregate.ai/model?slug=qwen-3-5-4b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.