Qwen 3.5 Plus: benchmark results

Alibaba's flagship Qwen 3.5 Plus API tier with thinking and non-thinking modes. Provider: Alibaba. Released 2026-02-16. Access: API.

Unified ELO 1680 ± 1, rank #84 of 1392 rated models, from 91 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
FINESSE-Bench87.76exam-like (self-reported)100
StemBind42.2F Overall (self-reported)100
TOBench41Avg. (self-reported)100
AI for Education Pedagogy - Science95.08Accuracy (%)99.6
ClawProBench64.19Final Score (self-reported)96.4
TempCloze85.74Mean Accuracy (%)95.8
EgoCoT-Bench70.68Mean (self-reported)94.4
CC-OCR V273.03Average (self-reported)92.9
AI for Education Pedagogy - Maths90.48Accuracy (%)92
AI for Education Pedagogy - Secondary88.36Accuracy (%)92
AI for Education Pedagogy88.88Accuracy (%)91.2
AI for Education Visual Maths - Number and Operations72.97Accuracy (%)91.2

Interactive version: theaggregate.ai/model?slug=qwen-3-5-plus · How It Works · Data refreshed daily, snapshot 2026-09-05.