GPT-OSS-20B (High): benchmark results

GPT-OSS-20B evaluated at the high reasoning-effort setting. Provider: OpenAI. Released 2025-08-05. Access: Open.

Unified ELO 1495 ± 1, rank #880 of 1761 rated models, from 402 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
APEX-Agents-AA70Pass@1 (self-reported)100
EuroEval Latvian NLU - Latvian Twitter Sentiment51.26Sentiment classification Score (%)93.5
AA AIME 202589.33Accuracy (%)89.3
EuroEval Spanish NLU - ScaLA ES39.3Linguistic acceptability Score (%)89.1
EuroEval Portuguese NLU - ScaLA PT35.96Linguistic acceptability Score (%)88
EuroEval English NLU - ScaLA EN58.02Linguistic acceptability Score (%)87.7
Medmarks - MedHallu Medium65.83Score (%)87.1
AA LiveCodeBench77.67Pass@1 (%)86.9
EuroEval Italian NLU - ScaLA IT38.45Linguistic acceptability Score (%)85
EuroEval German NLU - Sb10K57.44Sentiment classification Score (%)84.6
Medmarks - SCTpublic68.35Score (%)84.3
EuroEval Slovak NLU - UNER SK60.29Named entity recognition Score (%)83.3

Interactive version: theaggregate.ai/model?slug=gpt-oss-20b-high · How It Works · Data refreshed daily, snapshot 2026-09-05.