Qwen 3.5 0.8B (Non-reasoning) — benchmark results
Provider: Alibaba. Released 2026-03-02. Access: Open.
Unified ELO 1328 ± 63, rank #1475 of 1776 rated models, from 37 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA TAU-2 Bench | 65.2 | Accuracy (%) | 59.1 |
| CritPt | 0 | Accuracy (self-reported) | 30.3 |
| AA Humanity's Last Exam | 4.87 | Accuracy (%) | 30.1 |
| AA CritPt | 0 | Accuracy (%) | 27.1 |
| AA-LCR | 6.7 | Score (self-reported) | 24.5 |
| AA Omniscience - Software Engineering (SWE) - Julia | 4 | Accuracy (%) | 23.7 |
| AA Long Context Reasoning | 6.67 | Accuracy (%) | 18.6 |
| AA Omniscience - Software Engineering (SWE) - Go | 8 | Accuracy (%) | 17.5 |
| AA Omniscience - Software Engineering (SWE) - Java | 9 | Accuracy (%) | 16.2 |
| AA Omniscience - Software Engineering (SWE) - TypeScript | 7.78 | Accuracy (%) | 13.1 |
| AA Omniscience - Software Engineering (SWE) - HTML | 14 | Accuracy (%) | 10.7 |
| AA Omniscience - Software Engineering (SWE) - Dart | 8 | Accuracy (%) | 9 |
Interactive version: theaggregate.ai/model?slug=qwen-3-5-0-8b-non-reasoning · How the rankings work · Data refreshed daily, snapshot 2026-07-22.