GPT-OSS-20B (High) — benchmark results

GPT-OSS-20B evaluated at the high reasoning-effort setting. Provider: OpenAI. Released 2025-08-05. Access: Open.

Unified ELO 1521 ± 10, rank #718 of 1776 rated models, from 394 benchmark results.

Strongest benchmark results

BenchmarkScoreMetricPercentile
EuroEval Latvian NLU - Latvian Twitter Sentiment51.26Sentiment classification Score (%)93.5
AA LiveCodeBench77.67Pass@1 (%)89.3
EuroEval Spanish NLU - ScaLA ES39.3Linguistic acceptability Score (%)89.1
AA AIME 202589.33Accuracy (%)88.5
EuroEval Portuguese NLU - ScaLA PT35.96Linguistic acceptability Score (%)88
EuroEval English NLU - ScaLA EN58.02Linguistic acceptability Score (%)87.7
Medmarks - MedHallu Medium65.83Score (%)87.1
EuroEval Italian NLU - ScaLA IT38.45Linguistic acceptability Score (%)85
EuroEval German NLU - Sb10K57.44Sentiment classification Score (%)84.6
Medmarks - SCTpublic68.35Score (%)84.3
EuroEval Slovak NLU - UNER SK60.29Named entity recognition Score (%)83.3
EuroEval Slovene Knowledge64.23Knowledge Average Score (%)83.3

Interactive version: theaggregate.ai/model?slug=gpt-oss-20b-high · How the rankings work · Data refreshed daily, snapshot 2026-07-22.