GPT-OSS-120B (Reasoning): benchmark results

Provider: OpenAI. Released 2025-08-05. Access: Open.

Unified ELO 1580 ± 14, rank #738 of 2096 rated models, from 137 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
SEA-HELM (English) - MathArena45.48Normalized Score94.7
SEA-HELM (English) - LiveCodeBench v668.18Normalized Score93
BALSAM - Program Execution93.78Overall score (0-100, LLM-judged generation and multiple cho90.7
SEA-HELM (Tamil) - SEA-MT-Bench (LLM Judge)79.98Normalized Score89.5
SEA-HELM (Filipino) - SEA-Safeguard75.7Normalized Score87.7
BALSAM - Question Answering73.39Overall score (0-100, LLM-judged generation and multiple cho87
SEA-HELM (English) - MuSR77.64Normalized Score86
BALSAM - Overall55.91Mean of category overall scores (0-100)84.6
SEA-HELM (Thai) - SEA-MT-Bench (LLM Judge)83.21Normalized Score84.2
SEA-HELM (Filipino) - NLR74.95Normalized Score82.5
SEA-HELM (Filipino) - SEA-MT-Bench (LLM Judge)82.34Normalized Score82.5
SEA-HELM (Filipino) - Translations89.96Normalized Score82.5

Interactive version: theaggregate.ai/model?slug=gpt-oss-120b-reasoning · How It Works · Data refreshed daily, snapshot 2026-10-08.