GPT-OSS-Swallow-20B-RL-v0.1 (Thinking): benchmark results
Provider: OpenAI. Released 2026-02-20. Access: API.
Unified ELO 1576 ± 1, rank #779 of 3078 rated models, from 101 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Nejumi 4 - BFCL - Non-Live AST | 81.48 | Accuracy (%) | 91.5 |
| Swallow - Post-trained English - MATH-500 | 98.4 | Accuracy (%) | 89.6 |
| Nejumi 4 - MT-Bench (Japanese) - Extraction | 97.5 | Judge rating (1-10, x10) | 88.6 |
| Swallow - English MT-Bench - Coding | 80.2 | Judge Score (normalized, %) | 86.6 |
| Swallow - Post-trained English - LiveCodeBench | 66.5 | Pass@1 (%) | 86.6 |
| Nejumi 4 - BFCL - Live AST | 73.15 | Accuracy (%) | 82.4 |
| Nejumi 4 - jaster (2-shot) - JSICK | 83 | Exact match (%) | 82.4 |
| Swallow - Post-trained English - AIME | 85.4 | Accuracy (%) | 82.1 |
| Swallow - Post-trained Japanese - PolyMath High and Top | 53.6 | Accuracy (%) | 82.1 |
| Nejumi 4 - jaster (0-shot) - JaMP | 75 | Exact match (%) | 77.9 |
| Swallow - Japanese MT-Bench - Reasoning | 79.3 | Judge Score (normalized, %) | 77.6 |
| Swallow - Post-trained English - Average | 75.2 | Average Score (%) | 76.9 |
Interactive version: theaggregate.ai/model?slug=gpt-oss-swallow-20b-rl-v0-1-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.