GPT-OSS-Swallow-120B-RL-v0.1 (Thinking): benchmark results

Provider: OpenAI. Released 2026-02-20. Access: API.

Unified ELO 1614 ± 1, rank #544 of 3078 rated models, from 101 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Swallow - Post-trained English - MATH-50099Accuracy (%)97
Nejumi 4 - BFCL - Live AST75Accuracy (%)94.9
Swallow - Japanese MT-Bench - Reasoning83.2Judge Score (normalized, %)94
Swallow - Post-trained English - LiveCodeBench71.1Pass@1 (%)94
Swallow - English MT-Bench - Math99.3Judge Score (normalized, %)92.5
Nejumi 4 - BFCL - Irrelevance Detection93.33Accuracy (%)92.3
Nejumi 4 - MT-Bench (Japanese) - Roleplay99Judge rating (1-10, x10)91.5
Swallow - Post-trained Japanese - PolyMath High and Top57.6Accuracy (%)90.3
Nejumi 4 - Toxicity - Prohibited Acts97.4Criteria met (%)89.7
Swallow - English MT-Bench - Coding80.5Judge Score (normalized, %)88.1
Swallow - Japanese MT-Bench - Writing70.7Judge Score (normalized, %)88.1
Swallow - Post-trained Japanese - JHumanEval93.4Pass@1 (%)88.1

Interactive version: theaggregate.ai/model?slug=gpt-oss-swallow-120b-rl-v0-1-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.