GPT-OSS-120B: benchmark results
OpenAI's open-weight GPT-OSS 120B, an Apache-2.0 sparse MoE (~5B active) reasoning model (August 2025). Provider: OpenAI. Released 2025-08-05. Access: Open.
Unified ELO 1578 ± 1, rank #325 of 1392 rated models, from 658 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| HELM Safety HarmBench | 100 | LM Evaluated Safety score (%) | 100 |
| HealthBench Hard | 60 | Overall score (self-reported) | 100 |
| Onyx Open LLM Leaderboard | 90 | MMLU-Pro | 100 |
| Enkrypt AI - Insecure Code Risk | 1.33 | Risk Score | 98.1 |
| HELM Safety BBQ | 98.5 | BBQ accuracy (%) | 97.7 |
| ProLLM - LLM-as-a-Judge | 84.6 | Score (%) | 97.3 |
| HELM Safety | 98.1 | Mean score (self-reported) | 97 |
| ProLLM - Summarization | 98.2 | Score (%) | 96.6 |
| HELM AIR-Bench | 88 | Refusal Rate (%) | 95.3 |
| EuroEval Dutch Simplification - Duidelijke Taal | 55.82 | Score (%) | 93.8 |
| EuroEval Danish NLU - Dansk | 67.77 | Named entity recognition Score (%) | 93.3 |
| SEAL - MASK | 92 | Score | 92.4 |
Interactive version: theaggregate.ai/model?slug=gpt-oss-120b · How It Works · Data refreshed daily, snapshot 2026-09-05.