Qwen 3.5 0.8B (Thinking): benchmark results
Provider: Alibaba. Released 2026-03-02. Access: Open.
Unified ELO 1371 ± 1, rank #2732 of 3078 rated models, from 88 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA-Omniscience Hallucination Rate | 61.35 | Hallucination Rate (%) | 74.3 |
| Tau2-Bench Telecom | 47.66 | Success Rate (%) | 53 |
| MedLayXPlain | 58.2 | S (self-reported) | 51.6 |
| AA TAU-2 Bench | 47.66 | Accuracy (%) | 51.4 |
| AA-Omniscience Index - Law | -49.7 | Omniscience Index | 43.1 |
| AA-Omniscience Index - Software Engineering (SWE) - Java | -60 | Omniscience Index | 39.1 |
| AA-Omniscience Index - Business | -50.2 | Omniscience Index | 33.3 |
| AA-Omniscience Index - Science, Engineering & Mathematics | -43.5 | Omniscience Index | 29 |
| AA-Omniscience Index - Software Engineering (SWE) - TypeScript | -63.33 | Omniscience Index | 28.8 |
| AA-Omniscience Index - Humanities & Social Sciences | -54.2 | Omniscience Index | 28.5 |
| AA-Omniscience Index - Software Engineering (SWE) - Julia | -76 | Omniscience Index | 27.6 |
| AA Omniscience | -54.52 | Score | 25.2 |
Interactive version: theaggregate.ai/model?slug=qwen-3-5-0-8b-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.