GPT-OSS-120B (Medium) — benchmark results
GPT-OSS-120B evaluated at the medium reasoning-effort setting. Provider: OpenAI. Released 2025-08-05. Access: Open.
Unified ELO 1651 ± 24, rank #364 of 1776 rated models, from 40 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| Medmarks - PubMedQA | 78.13 | Score (%) | 92.9 |
| Medmarks - SCTpublic | 71.35 | Score (%) | 92.9 |
| Medmarks - Med-HALT Reasoning FCT | 87.96 | Score (%) | 91.4 |
| Medmarks - MedQA | 91.54 | Score (%) | 90 |
| Medmarks - MedCalc-Bench | 68.91 | Score (%) | 88.6 |
| Medmarks - Medbullets OP4 | 84.31 | Score (%) | 88.6 |
| Medmarks - Medbullets OP5 | 80.95 | Score (%) | 88.6 |
| Medmarks - PubHealthBench Reviewed | 88.25 | Score (%) | 87.9 |
| Medmarks - MedXpertQA Reasoning | 35.59 | Score (%) | 87.1 |
| Medmarks - MetaMedQA | 77.3 | Score (%) | 87.1 |
| SLR-Bench - Basic | 100 | Accuracy (%) | 86.4 |
| LiveOIBench | 57.55 | Avg Human Percentile | 86 |
Interactive version: theaggregate.ai/model?slug=gpt-oss-120b-medium · How the rankings work · Data refreshed daily, snapshot 2026-07-22.