Qwen 3.5 27B (Reasoning) — benchmark results
Provider: Alibaba. Released 2026-02-24. Access: Open.
Unified ELO 1606 ± 21, rank #518 of 1806 rated models, from 52 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA IFBench | 75.58 | Accuracy (%) | 93.4 |
| AA TAU-2 Bench | 93.86 | Accuracy (%) | 91.1 |
| AA-LCR | 67.3 | Score (self-reported) | 88 |
| AA Long Context Reasoning | 72.33 | Accuracy (%) | 85.7 |
| AA GPQA Diamond | 85.76 | Accuracy (%) | 84.1 |
| AA Global-MMLU-Lite - Japanese | 89.33 | Accuracy (%) | 80.9 |
| AA Humanity's Last Exam | 23.91 | Accuracy (%) | 80.8 |
| Artificial Analysis Intelligence Index | 34.6 | Intelligence Index | 80.3 |
| AA Global-MMLU-Lite - Burmese | 83.83 | Accuracy (%) | 79.1 |
| AA Global-MMLU-Lite - Arabic | 87.08 | Accuracy (%) | 78 |
| AA Global-MMLU-Lite - Indonesian | 89.5 | Accuracy (%) | 77.8 |
| AA Global-MMLU-Lite - Yoruba | 71.5 | Accuracy (%) | 77.5 |
Interactive version: theaggregate.ai/model?slug=qwen-3-5-27b-reasoning · How It Works · Data refreshed daily, snapshot 2026-08-07.