Qwen 3.5 2B — benchmark results

2B on-device member of Alibaba's Apache-2.0 Qwen3.5 family (March 2026), sized for phones and laptops. Provider: Alibaba. Released 2026-03-02. Access: Open.

Unified ELO 1421 ± 13, rank #1123 of 1776 rated models, from 345 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Benchmarking Large Language Models for Graphem87.3Direct (self-reported)97.2
EuroEval Bulgarian NLU - Cinexio58.3Sentiment classification Score (%)93.4
EuroEval French NLU - ScaLA FR52.57Linguistic acceptability Score (%)88
EuroEval Portuguese NLU - ScaLA PT35.02Linguistic acceptability Score (%)87.3
EuroEval Italian NLU - ScaLA IT40.14Linguistic acceptability Score (%)86.6
EuroEval Catalan NLU - Guia CAT68.8Sentiment classification Score (%)85.8
EuroEval Albanian Summarization - LR SUM SQ33.48Score (%)83.7
EuroEval Catalan NLU - MultiWikiQA CA72.51Reading comprehension Score (%)81.5
EuroEval Spanish Summarization - Mlsum ES27.88Score (%)78.9
EuroEval Polish Summarization - PSC23.61Score (%)78.2
EuroEval Bosnian Summarization - LR SUM BS29.81Score (%)77.6
StemBind35.9F Overall (self-reported)76.1

Interactive version: theaggregate.ai/model?slug=qwen-3-5-2b · How the rankings work · Data refreshed daily, snapshot 2026-07-22.