Qwen3-Swallow-32B-RL-v0.2 (Thinking): benchmark results

Provider: Alibaba. Access: Open.

Unified ELO 1572 ± 1, rank #555 of 2032 rated models, from 101 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Nejumi 4 - BFCL - Live AST75.93Accuracy (%)97.1
Swallow - Japanese MT-Bench - Extraction76.8Judge Score (normalized, %)92.5
Swallow - Japanese MT-Bench - Coding83.1Judge Score (normalized, %)91
Nejumi 4 - jaster (2-shot) - JSICK84Exact match (%)89.3
Nejumi 4 - BFCL - Irrelevance Detection91.67Accuracy (%)86.4
Nejumi 4 - jaster (0-shot) - JCoLA (out-of-domain)88Exact match (%)86.4
Swallow - Post-trained English - MATH-50098.2Accuracy (%)85.8
Swallow - Japanese MT-Bench - Roleplay71Judge Score (normalized, %)82.1
Swallow - Post-trained Japanese - JHumanEval92.9Pass@1 (%)82.1
Nejumi 4 - JBBQ (2-shot) - Accuracy93Accuracy (%)81.6
Swallow - Post-trained English - LiveCodeBench63.8Pass@1 (%)81.3
Swallow - Post-trained English - AIME82.5Accuracy (%)79.1

Interactive version: theaggregate.ai/model?slug=qwen3-swallow-32b-rl-v0-2-thinking · How It Works · Data refreshed daily, snapshot 2026-09-26.