Qwen 3.5 2B (Reasoning) — benchmark results
Provider: Alibaba. Released 2026-03-02. Access: Open.
Unified ELO 1368 ± 46, rank #1402 of 1841 rated models, from 37 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA TAU-2 Bench | 69.01 | Accuracy (%) | 61.9 |
| AA Omniscience | -47.48 | Score | 41.3 |
| AA Long Context Reasoning | 23.67 | Accuracy (%) | 35 |
| CritPt | 0 | Accuracy (self-reported) | 30.3 |
| AA CritPt | 0 | Accuracy (%) | 26.8 |
| AA Terminal-Bench Hard | 3.79 | Accuracy (%) | 25.6 |
| AA Omniscience - Software Engineering (SWE) - Julia | 4 | Accuracy (%) | 23.5 |
| Artificial Analysis Intelligence Index | 6.86 | Intelligence Index | 20.4 |
| AA IFBench | 30.41 | Accuracy (%) | 14.7 |
| AA MATH-500 | 42.96 | Accuracy (%) | 11.2 |
| AA GDPval | 219.82 | ELO | 10.9 |
| GDPval-AA | 220 | Elo | 10.6 |
Interactive version: theaggregate.ai/model?slug=qwen-3-5-2b-reasoning · How It Works · Data refreshed daily, snapshot 2026-07-25.