Qwen 3.5 9B (Thinking): benchmark results

Provider: Alibaba. Released 2026-03-02. Access: Open.

Unified ELO 1581 ± 1, rank #738 of 3078 rated models, from 122 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Swallow - Japanese MT-Bench - Reasoning84.3Judge Score (normalized, %)95.5
Swallow - Post-trained English - GPQA Diamond82.7Accuracy (%)91
Nejumi 4 - BFCL - Live AST74.07Accuracy (%)89.3
Nejumi 4 - JHumanEval42.46Sandbox pass rate (%)84.6
FrameBench - Frame Identification - English81.3Accuracy (%; FrameNet candidate frames)84.2
Swallow - Post-trained English - MMLU-Pro82.6Accuracy (%)83.6
Swallow - Post-trained Japanese - GPQA72.5Accuracy (%)83.6
Nejumi 4 - jaster (2-shot) - JaMP79Exact match (%)81.2
Swallow - Post-trained Japanese - MMLU-ProX78.4Accuracy (%)80.6
K-MetBench74.9Accuracy (self-reported)78.4
Swallow - Post-trained English - HellaSwag90.9Accuracy (%)77.6
Korean CSAT 2026 (Easy Mode) - Physics I41Points (out of 50)76.9

Interactive version: theaggregate.ai/model?slug=qwen-3-5-9b-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.