GPT-OSS-Swallow-120B-RL-v0.1 (Thinking): benchmark results
Provider: OpenAI. Released 2026-02-20. Access: API.
Unified ELO 1614 ± 1, rank #544 of 3078 rated models, from 101 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Swallow - Post-trained English - MATH-500 | 99 | Accuracy (%) | 97 |
| Nejumi 4 - BFCL - Live AST | 75 | Accuracy (%) | 94.9 |
| Swallow - Japanese MT-Bench - Reasoning | 83.2 | Judge Score (normalized, %) | 94 |
| Swallow - Post-trained English - LiveCodeBench | 71.1 | Pass@1 (%) | 94 |
| Swallow - English MT-Bench - Math | 99.3 | Judge Score (normalized, %) | 92.5 |
| Nejumi 4 - BFCL - Irrelevance Detection | 93.33 | Accuracy (%) | 92.3 |
| Nejumi 4 - MT-Bench (Japanese) - Roleplay | 99 | Judge rating (1-10, x10) | 91.5 |
| Swallow - Post-trained Japanese - PolyMath High and Top | 57.6 | Accuracy (%) | 90.3 |
| Nejumi 4 - Toxicity - Prohibited Acts | 97.4 | Criteria met (%) | 89.7 |
| Swallow - English MT-Bench - Coding | 80.5 | Judge Score (normalized, %) | 88.1 |
| Swallow - Japanese MT-Bench - Writing | 70.7 | Judge Score (normalized, %) | 88.1 |
| Swallow - Post-trained Japanese - JHumanEval | 93.4 | Pass@1 (%) | 88.1 |
Interactive version: theaggregate.ai/model?slug=gpt-oss-swallow-120b-rl-v0-1-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.