Qwen3-Swallow-32B-RL-v0.2 (Thinking): benchmark results
Provider: Alibaba. Access: Open.
Unified ELO 1572 ± 1, rank #555 of 2032 rated models, from 101 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Nejumi 4 - BFCL - Live AST | 75.93 | Accuracy (%) | 97.1 |
| Swallow - Japanese MT-Bench - Extraction | 76.8 | Judge Score (normalized, %) | 92.5 |
| Swallow - Japanese MT-Bench - Coding | 83.1 | Judge Score (normalized, %) | 91 |
| Nejumi 4 - jaster (2-shot) - JSICK | 84 | Exact match (%) | 89.3 |
| Nejumi 4 - BFCL - Irrelevance Detection | 91.67 | Accuracy (%) | 86.4 |
| Nejumi 4 - jaster (0-shot) - JCoLA (out-of-domain) | 88 | Exact match (%) | 86.4 |
| Swallow - Post-trained English - MATH-500 | 98.2 | Accuracy (%) | 85.8 |
| Swallow - Japanese MT-Bench - Roleplay | 71 | Judge Score (normalized, %) | 82.1 |
| Swallow - Post-trained Japanese - JHumanEval | 92.9 | Pass@1 (%) | 82.1 |
| Nejumi 4 - JBBQ (2-shot) - Accuracy | 93 | Accuracy (%) | 81.6 |
| Swallow - Post-trained English - LiveCodeBench | 63.8 | Pass@1 (%) | 81.3 |
| Swallow - Post-trained English - AIME | 82.5 | Accuracy (%) | 79.1 |
Interactive version: theaggregate.ai/model?slug=qwen3-swallow-32b-rl-v0-2-thinking · How It Works · Data refreshed daily, snapshot 2026-09-26.