Qwen 3.5 4B (Thinking): benchmark results

Provider: Alibaba. Released 2026-03-02. Access: Open.

Unified ELO 1533 ± 1, rank #1107 of 3078 rated models, from 108 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Nejumi 4 - BFCL - Live AST74.07Accuracy (%)89.3
Nejumi 4 - BFCL - Relevance Detection88.89Accuracy (%)89
Swallow - Post-trained English - GPQA Diamond76.5Accuracy (%)83.6
Swallow - Post-trained Japanese - GPQA68.8Accuracy (%)77.6
FrameBench - Frame Identification - English81Accuracy (%; FrameNet candidate frames)76.3
Nejumi 4 - Toxicity - Prohibited Acts95.83Criteria met (%)75.4
Swallow - Post-trained English - MMLU-Pro78.8Accuracy (%)71.6
Nejumi 4 - HalluLens95Hallucination resistance (%)71.3
Nejumi 4 - Toxicity - Violation Categories46.83Criteria met (%)71
Nejumi 4 - JHumanEval39.03Sandbox pass rate (%)69.9
Nejumi 4 - Toxicity - Fairness97.18Criteria met (%)69.5
Nejumi 4 - jaster (0-shot) - MAWPS98Exact match (%)68.4

Interactive version: theaggregate.ai/model?slug=qwen-3-5-4b-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.