Qwen 3.5 Plus: benchmark results
Alibaba's flagship Qwen 3.5 Plus API tier with thinking and non-thinking modes. Provider: Alibaba. Released 2026-02-16. Access: API.
Unified ELO 1680 ± 1, rank #84 of 1392 rated models, from 91 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| FINESSE-Bench | 87.76 | exam-like (self-reported) | 100 |
| StemBind | 42.2 | F Overall (self-reported) | 100 |
| TOBench | 41 | Avg. (self-reported) | 100 |
| AI for Education Pedagogy - Science | 95.08 | Accuracy (%) | 99.6 |
| ClawProBench | 64.19 | Final Score (self-reported) | 96.4 |
| TempCloze | 85.74 | Mean Accuracy (%) | 95.8 |
| EgoCoT-Bench | 70.68 | Mean (self-reported) | 94.4 |
| CC-OCR V2 | 73.03 | Average (self-reported) | 92.9 |
| AI for Education Pedagogy - Maths | 90.48 | Accuracy (%) | 92 |
| AI for Education Pedagogy - Secondary | 88.36 | Accuracy (%) | 92 |
| AI for Education Pedagogy | 88.88 | Accuracy (%) | 91.2 |
| AI for Education Visual Maths - Number and Operations | 72.97 | Accuracy (%) | 91.2 |
Interactive version: theaggregate.ai/model?slug=qwen-3-5-plus · How It Works · Data refreshed daily, snapshot 2026-09-05.