GPT-OSS-120B (Reasoning): benchmark results
Provider: OpenAI. Released 2025-08-05. Access: Open.
Unified ELO 1580 ± 14, rank #738 of 2096 rated models, from 137 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| SEA-HELM (English) - MathArena | 45.48 | Normalized Score | 94.7 |
| SEA-HELM (English) - LiveCodeBench v6 | 68.18 | Normalized Score | 93 |
| BALSAM - Program Execution | 93.78 | Overall score (0-100, LLM-judged generation and multiple cho | 90.7 |
| SEA-HELM (Tamil) - SEA-MT-Bench (LLM Judge) | 79.98 | Normalized Score | 89.5 |
| SEA-HELM (Filipino) - SEA-Safeguard | 75.7 | Normalized Score | 87.7 |
| BALSAM - Question Answering | 73.39 | Overall score (0-100, LLM-judged generation and multiple cho | 87 |
| SEA-HELM (English) - MuSR | 77.64 | Normalized Score | 86 |
| BALSAM - Overall | 55.91 | Mean of category overall scores (0-100) | 84.6 |
| SEA-HELM (Thai) - SEA-MT-Bench (LLM Judge) | 83.21 | Normalized Score | 84.2 |
| SEA-HELM (Filipino) - NLR | 74.95 | Normalized Score | 82.5 |
| SEA-HELM (Filipino) - SEA-MT-Bench (LLM Judge) | 82.34 | Normalized Score | 82.5 |
| SEA-HELM (Filipino) - Translations | 89.96 | Normalized Score | 82.5 |
Interactive version: theaggregate.ai/model?slug=gpt-oss-120b-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-08.