GPT-OSS-120B (High) — benchmark results

GPT-OSS-120B evaluated at the high reasoning-effort setting. Provider: OpenAI. Released 2025-08-05. Access: Open.

Unified ELO 1680 ± 12, rank #308 of 1776 rated models, from 167 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
AA LiveCodeBench87.83Pass@1 (%)98.2
LiveOIBench71.85Avg Human Percentile98.2
Medmarks - MedQA93.27Score (%)95.7
AA AIME 202593.44Accuracy (%)94.4
Medmarks - MetaMedQA79.39Score (%)92.9
SLR-Bench - Medium58Accuracy (%)92.7
SLR-Bench - Easy97Accuracy (%)91.8
Medmarks - LongHealth Task 189.33Score (%)91.4
Medmarks - MedCalc-Bench71.73Score (%)91.4
Medmarks - MedHallu Easy73.45Score (%)91.4
Medmarks - Medbullets OP582.58Score (%)91.4
Medmarks - Med-HALT Reasoning FCT87.67Score (%)90

Interactive version: theaggregate.ai/model?slug=gpt-oss-120b-high · How the rankings work · Data refreshed daily, snapshot 2026-07-22.