GPT-OSS-120B (Low) — benchmark results
GPT-OSS-120B evaluated at the low reasoning-effort setting. Provider: OpenAI. Released 2025-08-05. Access: Open.
Unified ELO 1604 ± 16, rank #472 of 1776 rated models, from 90 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Medmarks - SCTpublic | 70.41 | Score (%) | 90 |
| Medmarks - PubMedQA | 77.4 | Score (%) | 87.9 |
| Medmarks - Med-HALT Reasoning FCT | 86.7 | Score (%) | 87.1 |
| Medmarks - PubHealthBench Reviewed | 87.89 | Score (%) | 85.7 |
| Medmarks - Medbullets OP4 | 81.6 | Score (%) | 82.9 |
| Medmarks - MedCalc-Bench | 66.09 | Score (%) | 81.4 |
| AA LiveCodeBench | 70.69 | Pass@1 (%) | 80.8 |
| Medmarks - MedXpertQA Reasoning | 28.93 | Score (%) | 78.6 |
| Medmarks - MedXpertQA Understanding | 27.96 | Score (%) | 78.6 |
| Medmarks - Medbullets OP5 | 76.08 | Score (%) | 78.6 |
| AA Omniscience - Software Engineering (SWE) - Swift | 56 | Accuracy (%) | 77.6 |
| Medmarks - MedQA | 88.61 | Score (%) | 72.9 |
Interactive version: theaggregate.ai/model?slug=gpt-oss-120b-low · How the rankings work · Data refreshed daily, snapshot 2026-07-22.