Qwen 3.5 9B (Non-reasoning) — benchmark results

Qwen 3.5 9B evaluated with reasoning disabled. Provider: Alibaba. Released 2026-03-02. Access: Open.

Unified ELO 1519 ± 47, rank #728 of 1776 rated models, from 37 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AA TAU-2 Bench85.09Accuracy (%)77.4
AA GPQA Diamond78.59Accuracy (%)70.5
AA CritPt0.57Accuracy (%)64.9
Artificial Analysis Intelligence Index20.32Intelligence Index60.3
AA Terminal-Bench Hard18.18Accuracy (%)57.3
AA Humanity's Last Exam8.62Accuracy (%)57.2
AA MMMU-Pro66.76Accuracy (%)51.3
UGI - Writing33.52Writing Score50.2
AA Long Context Reasoning38Accuracy (%)49.4
AA Omniscience - Software Engineering (SWE) - Swift36Accuracy (%)47.7
AA Omniscience - Software Engineering (SWE) - Go16Accuracy (%)46.2
AA Omniscience - Software Engineering (SWE) - C31Accuracy (%)41.9

Interactive version: theaggregate.ai/model?slug=qwen-3-5-9b-non-reasoning · How the rankings work · Data refreshed daily, snapshot 2026-07-22.