Qwen 3.6 27B (Reasoning): benchmark results
Provider: Alibaba. Released 2026-04-22. Access: Open.
Unified ELO 1626 ± 1, rank #332 of 1761 rated models, from 37 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| CPTU Bench | 4.41 | Average Score (1-5) | 99 |
| AA TAU-2 Bench | 94.15 | Accuracy (%) | 92 |
| AA Long Context Reasoning | 77.33 | Accuracy (%) | 81.7 |
| AA IFBench | 67.55 | Accuracy (%) | 81 |
| AA Terminal-Bench Hard | 34.85 | Accuracy (%) | 80.3 |
| AA Omniscience - Software Engineering (SWE) - TypeScript | 33.33 | Accuracy (%) | 79.4 |
| Artificial Analysis Intelligence Index | 29 | Intelligence Index | 79.2 |
| AA Omniscience - Software Engineering (SWE) - Swift | 44 | Accuracy (%) | 78.2 |
| AA Omniscience - Software Engineering (SWE) - HTML | 40 | Accuracy (%) | 75.9 |
| AA GPQA Diamond | 84.24 | Accuracy (%) | 75.5 |
| AA Humanity's Last Exam | 23.08 | Accuracy (%) | 75.4 |
| AA-LCR | 73.33 | Accuracy (self-reported) | 74.2 |
Interactive version: theaggregate.ai/model?slug=qwen-3-6-27b-reasoning · How It Works · Data refreshed daily, snapshot 2026-09-05.