Qwen 3.5 2B — benchmark results
2B on-device member of Alibaba's Apache-2.0 Qwen3.5 family (March 2026), sized for phones and laptops. Provider: Alibaba. Released 2026-03-02. Access: Open.
Unified ELO 1421 ± 13, rank #1123 of 1776 rated models, from 345 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Benchmarking Large Language Models for Graphem | 87.3 | Direct (self-reported) | 97.2 |
| EuroEval Bulgarian NLU - Cinexio | 58.3 | Sentiment classification Score (%) | 93.4 |
| EuroEval French NLU - ScaLA FR | 52.57 | Linguistic acceptability Score (%) | 88 |
| EuroEval Portuguese NLU - ScaLA PT | 35.02 | Linguistic acceptability Score (%) | 87.3 |
| EuroEval Italian NLU - ScaLA IT | 40.14 | Linguistic acceptability Score (%) | 86.6 |
| EuroEval Catalan NLU - Guia CAT | 68.8 | Sentiment classification Score (%) | 85.8 |
| EuroEval Albanian Summarization - LR SUM SQ | 33.48 | Score (%) | 83.7 |
| EuroEval Catalan NLU - MultiWikiQA CA | 72.51 | Reading comprehension Score (%) | 81.5 |
| EuroEval Spanish Summarization - Mlsum ES | 27.88 | Score (%) | 78.9 |
| EuroEval Polish Summarization - PSC | 23.61 | Score (%) | 78.2 |
| EuroEval Bosnian Summarization - LR SUM BS | 29.81 | Score (%) | 77.6 |
| StemBind | 35.9 | F Overall (self-reported) | 76.1 |
Interactive version: theaggregate.ai/model?slug=qwen-3-5-2b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.