Qwen3-Swallow-8B-RL-v0.2 (Thinking): benchmark results
Provider: Alibaba. Access: Open.
Unified ELO 1527 ± 1, rank #816 of 2032 rated models, from 101 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Nejumi 4 - jaster (2-shot) - JSICK | 89 | Exact match (%) | 100 |
| Nejumi 4 - jaster (0-shot) - JSICK | 84 | Exact match (%) | 92.3 |
| Nejumi 4 - jaster (0-shot) - JCoLA (in-domain) | 79 | Exact match (%) | 83.5 |
| Swallow - Japanese MT-Bench - Math | 98.4 | Judge Score (normalized, %) | 82.8 |
| Nejumi 4 - jaster (0-shot) - JaNLI | 100 | Exact match (%) | 82.4 |
| Swallow - Post-trained Japanese - PolyMath High and Top | 45.6 | Accuracy (%) | 73.1 |
| Swallow - English MT-Bench - Extraction | 80.1 | Judge Score (normalized, %) | 72.4 |
| Nejumi 4 - MT-Bench (Japanese) - STEM | 99.5 | Judge rating (1-10, x10) | 71.3 |
| Nejumi 4 - BFCL - Relevance Detection | 77.78 | Accuracy (%) | 70.2 |
| Nejumi 4 - BFCL - Non-Live AST | 77.78 | Accuracy (%) | 64.7 |
| Swallow - Post-trained English - AIME | 73.3 | Accuracy (%) | 64.2 |
| Nejumi 4 - BFCL - Irrelevance Detection | 85 | Accuracy (%) | 61.8 |
Interactive version: theaggregate.ai/model?slug=qwen3-swallow-8b-rl-v0-2-thinking · How It Works · Data refreshed daily, snapshot 2026-09-26.