Qwen 3.5 4B (Reasoning) — benchmark results
Provider: Alibaba. Released 2026-03-01. Access: Open.
Unified ELO 1572 ± 42, rank #583 of 1841 rated models, from 35 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA TAU-2 Bench | 92.11 | Accuracy (%) | 87.4 |
| AA Long Context Reasoning | 55.67 | Accuracy (%) | 65.1 |
| AA IFBench | 50.2 | Accuracy (%) | 59.9 |
| Artificial Analysis Intelligence Index | 20.09 | Intelligence Index | 59.6 |
| AA Terminal-Bench Hard | 18.18 | Accuracy (%) | 57.3 |
| AA Humanity's Last Exam | 7.83 | Accuracy (%) | 53.9 |
| AA GPQA Diamond | 67.98 | Accuracy (%) | 48.8 |
| AA MMMU-Pro | 65.38 | Accuracy (%) | 47.4 |
| AA Omniscience - Science, Engineering & Mathematics | 23.5 | Accuracy (%) | 38.9 |
| Tau3 Banking | 8.25 | Success Rate (%) | 38.8 |
| AA Omniscience - Software Engineering (SWE) - Julia | 8 | Accuracy (%) | 36.9 |
| AA MATH-500 | 73.08 | Accuracy (%) | 35 |
Interactive version: theaggregate.ai/model?slug=qwen-3-5-4b-reasoning · How It Works · Data refreshed daily, snapshot 2026-07-25.