Qwen 3.5 2B (Non-reasoning) — benchmark results

Qwen 3.5 2B evaluated with reasoning disabled. Provider: Alibaba. Released 2026-03-02. Access: Open.

Unified ELO 1468 ± 43, rank #915 of 1776 rated models, from 41 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AA TAU-2 Bench81.58Accuracy (%)72.3
AA Omniscience - Software Engineering (SWE) - Julia16Accuracy (%)57.8
UGI - Willingness (W/10)6W/10 Score53.7
UGI - Writing29Writing Score33.7
AA-LCR13.7Score (self-reported)32.9
AA Omniscience - Software Engineering (SWE) - R6Accuracy (%)30.9
CritPt0Accuracy (self-reported)30.3
AA Humanity's Last Exam4.87Accuracy (%)30.1
AA CritPt0Accuracy (%)27.1
AA Terminal-Bench Hard3.79Accuracy (%)25.6
AA Long Context Reasoning13.67Accuracy (%)24.9
AA Omniscience - Software Engineering (SWE) - Go10Accuracy (%)24.7

Interactive version: theaggregate.ai/model?slug=qwen-3-5-2b-non-reasoning · How the rankings work · Data refreshed daily, snapshot 2026-07-22.