GPT-OSS-Swallow-20B-RL-v0.1 (Thinking): benchmark results

Provider: OpenAI. Released 2026-02-20. Access: API.

Unified ELO 1576 ± 1, rank #779 of 3078 rated models, from 101 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
Nejumi 4 - BFCL - Non-Live AST81.48Accuracy (%)91.5
Swallow - Post-trained English - MATH-50098.4Accuracy (%)89.6
Nejumi 4 - MT-Bench (Japanese) - Extraction97.5Judge rating (1-10, x10)88.6
Swallow - English MT-Bench - Coding80.2Judge Score (normalized, %)86.6
Swallow - Post-trained English - LiveCodeBench66.5Pass@1 (%)86.6
Nejumi 4 - BFCL - Live AST73.15Accuracy (%)82.4
Nejumi 4 - jaster (2-shot) - JSICK83Exact match (%)82.4
Swallow - Post-trained English - AIME85.4Accuracy (%)82.1
Swallow - Post-trained Japanese - PolyMath High and Top53.6Accuracy (%)82.1
Nejumi 4 - jaster (0-shot) - JaMP75Exact match (%)77.9
Swallow - Japanese MT-Bench - Reasoning79.3Judge Score (normalized, %)77.6
Swallow - Post-trained English - Average75.2Average Score (%)76.9

Interactive version: theaggregate.ai/model?slug=gpt-oss-swallow-20b-rl-v0-1-thinking · How It Works · Data refreshed daily, snapshot 2026-09-19.