GPT-OSS-120B (High) — benchmark results
GPT-OSS-120B evaluated at the high reasoning-effort setting. Provider: OpenAI. Released 2025-08-05. Access: Open.
Unified ELO 1680 ± 12, rank #308 of 1776 rated models, from 167 benchmark results.
Strongest benchmark results
| Benchmark | Score | Metric | Percentile |
|---|---|---|---|
| AA LiveCodeBench | 87.83 | Pass@1 (%) | 98.2 |
| LiveOIBench | 71.85 | Avg Human Percentile | 98.2 |
| Medmarks - MedQA | 93.27 | Score (%) | 95.7 |
| AA AIME 2025 | 93.44 | Accuracy (%) | 94.4 |
| Medmarks - MetaMedQA | 79.39 | Score (%) | 92.9 |
| SLR-Bench - Medium | 58 | Accuracy (%) | 92.7 |
| SLR-Bench - Easy | 97 | Accuracy (%) | 91.8 |
| Medmarks - LongHealth Task 1 | 89.33 | Score (%) | 91.4 |
| Medmarks - MedCalc-Bench | 71.73 | Score (%) | 91.4 |
| Medmarks - MedHallu Easy | 73.45 | Score (%) | 91.4 |
| Medmarks - Medbullets OP5 | 82.58 | Score (%) | 91.4 |
| Medmarks - Med-HALT Reasoning FCT | 87.67 | Score (%) | 90 |
Interactive version: theaggregate.ai/model?slug=gpt-oss-120b-high · How the rankings work · Data refreshed daily, snapshot 2026-07-22.