Qwen 3.5 2B (Thinking): benchmark results
Provider: Alibaba. Released 2026-03-02. Access: Open.
Unified ELO 1414 ± 1, rank #2350 of 3078 rated models, from 45 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| MedLayXPlain | 63.7 | S (self-reported) | 77.4 |
| Avalon-ToM-Bench - Perspective-Taking | 55.88 | Accuracy (%) | 28.8 |
| Swallow - Post-trained English - GPQA Diamond | 49.5 | Accuracy (%) | 26.9 |
| FrameBench - Frame Identification - English | 68 | Accuracy (%; FrameNet candidate frames) | 26.3 |
| FrameBench - Frame Identification - Japanese | 54.3 | Accuracy (%; Japanese FrameNet candidate frames) | 24 |
| FrameBench - English | 80.3 | Accuracy (%; mean of five prompt templates) | 21.1 |
| FrameBench - Japanese | 60.7 | Accuracy (%; mean of five prompt templates) | 20 |
| Swallow - English MT-Bench - Math | 84 | Judge Score (normalized, %) | 19.4 |
| Swallow - Post-trained English - MATH-500 | 81.2 | Accuracy (%) | 17.9 |
| Swallow - Post-trained English - MMLU-Pro | 58.9 | Accuracy (%) | 17.9 |
| Avalon-ToM-Bench - Tacit Coordination | 63.5 | Accuracy (%) | 15.4 |
| Avalon-ToM-Bench - Intention Attribution | 65 | Accuracy (%) | 13.5 |
Interactive version: theaggregate.ai/model?slug=qwen-3-5-2b-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.